IVDR Performance Evaluation Plan (PEP): How to Build an Audit-Ready Evidence Framework

Written by Catarina Sepúlveda
Published on 25.08.2026 Last updated on 09.09.2026

For an IVD manufacturer, the Performance Evaluation Plan is where the performance-evidence strategy becomes explicit. Before individual reports are assembled into the technical documentation, the PEP sets out what the device needs to demonstrate, which evidence will support each claim, how that evidence will be judged and where further work remains.

Article 56 and Annex XIII of Regulation (EU) 2017/746 place this planning within a continuous performance-evaluation process. The clinical evidence for an IVD draws on scientific validity, analytical performance and clinical performance, all considered in relation to the device’s intended purpose.

Those three evidence streams are familiar to most regulatory teams. The quality of the PEP depends largely on the connections made between them. A manufacturer may have a detailed Scientific Validity Report, an extensive analytical validation package and a sound clinical performance study, yet still leave a reviewer searching for the reasoning that connects those documents to the claims, risks and intended purpose of the device.

A well-constructed PEP makes that reasoning visible from the outset.

What is a Performance Evaluation Plan under the IVDR?

The Performance Evaluation Plan, or PEP, describes how the manufacturer intends to generate, collect, assess and maintain the evidence required to support the performance of an IVD throughout its lifecycle.

Article 56 requires manufacturers to specify and justify the level of clinical evidence needed to demonstrate conformity with the relevant General Safety and Performance Requirements. The level of evidence has to reflect the characteristics and intended purpose of the device. The same Article places performance evaluation within a continuous process based on a defined plan.

Annex XIII gives that process its practical structure. It requires manufacturers to establish and update a PEP that describes the characteristics and performance of the device, together with the processes and criteria used to generate the necessary clinical evidence.

MDCG 2022-2 develops the same lifecycle approach in its guidance on clinical evidence for IVDs and provides further context for the planning and evaluation of scientific validity, analytical performance and clinical performance.

The result should read as an evidence programme for a particular device, not as a regulatory form completed once the studies have already been designed.

The relationship between the PEP and the PER

The Performance Evaluation Plan and Performance Evaluation Report belong to the same process, but they sit at different points within it.

The PEP establishes the strategy prospectively. It records the questions that need to be answered, the evidence sources that will be used, the methods and criteria that will be applied and the activities planned to address remaining gaps.

The Performance Evaluation Report brings the resulting evidence together and evaluates what that evidence supports.

Performance Evaluation Plan (PEP)Performance Evaluation Report (PER)
Establishes the evidence strategyConsolidates and assesses the resulting evidence
Defines what needs to be demonstratedRecords what the evidence supports
Identifies evidence sources and planned activitiesEvaluates the completed body of evidence
Establishes methods, milestones and acceptance criteriaAssesses results in the context of those criteria
Identifies gaps that still require evidenceRecords whether material gaps remain
Incorporates PMPF into the forward evidence planIncorporates relevant post-market findings into the overall evaluation

Annex XIII requires the PER to incorporate the scientific validity, analytical performance and clinical performance reports and to provide an assessment of the evidence they contain. The quality of that final assessment depends heavily on the questions and criteria established earlier in the PEP.

This relationship becomes particularly important when evidence changes over time. The PEP provides the structure against which new evidence can be considered; the PER records the updated assessment.

Where scientific validity, analytical performance and clinical performance meet

The three pillars of IVDR performance evaluation have different functions.

Scientific validity establishes the association between an analyte or marker and a clinical condition or physiological state.

Analytical performance concerns the ability of the device to correctly detect or measure the analyte.

Clinical performance considers the device’s ability to yield results that correlate with a particular clinical condition or physiological or pathological process or state, taking account of the target population and intended user.

Each requires its own evidence and its own technical assessment. Within the PEP, however, they form part of the same regulatory argument.

A practical line of traceability is:

A clear traceability structure should connect the intended purpose and performance claims with the applicable GSPRs and associated risks, identify the scientific validity, analytical performance and clinical performance evidence required to support them, establish the relevant acceptance criteria, and link the resulting studies and reports to the conclusions documented in the PER.

Catarina Sepulveda, IVD Regulatory Director at MDx CRO, uses this structure when approaching PEP development:

“In practice, PEP should be structured around clear traceability: intended purpose/claim, applicable GSPR and risk, required SV/AP/CP evidence, predefined acceptance criteria, supporting study/report, conclusion in the PER. This approach makes the technical documentation easier to review and demonstrates that each performance claim is supported by appropriate and sufficient evidence.”

Catarina Sepulveda | IVD Regulatory Director at MDx CRO

Her point is useful because it shifts attention from the existence of individual reports to the reasoning that holds the performance evaluation together.

The individual disciplines still need depth. Manufacturers working on the first pillar can refer to our Scientific Validity Report under IVDR, while analytical evidence is covered in more detail in our guide to IVD Analytical Validation under IVDR. Where prospective clinical evidence is required, our IVDR Clinical Performance Study guide addresses that process separately.

The PEP sits above these activities. Its purpose is to explain why each piece of evidence is needed and how the completed evidence package will support the intended purpose of the IVD.

Reading Annex XIII as an evidence strategy

Annex XIII Part A, Section 1.1 provides a detailed list of elements for the PEP. In practice, those requirements can be organised around a series of decisions that manufacturers have to make during development.

Begin with the device that will actually be placed on the market

The intended purpose sets the boundaries of the performance evaluation.

The PEP should describe the relevant device characteristics, analyte or marker, intended use, target patient groups, intended users, indications, limitations and contraindications. Where metrological traceability is relevant, the applicable certified reference materials or reference measurement procedures also need to be identified.

These details deserve close attention because they determine the scope of the evidence that follows.

A change in target population can affect the relevance of clinical performance data. A change in specimen type may alter the analytical evidence needed. A new claim can introduce an evidence requirement that was absent from the original development programme.

Consistency across the PEP, device description, labelling, risk-management documentation and the wider IVDR technical documentation is therefore fundamental.

When that consistency is weak, evidence that is technically sound may nevertheless be difficult to apply to the device as finally described.

Map performance evidence to the GSPRs and the risk file

Annex XIII requires manufacturers to identify the General Safety and Performance Requirements in Sections 1 to 9 of Annex I that need support from scientific validity, analytical performance or clinical performance data.

A list of GSPR references may satisfy document structure, but it tells a reviewer relatively little about the manufacturer’s reasoning.

The more useful exercise is to follow the relationship from the regulatory requirement to the device claim, then to the relevant risk and finally to the evidence needed to support it.

Consider an IVD for which an incorrect result could influence a clinically significant treatment decision. The analytical and clinical performance evidence will need to support the performance assumptions used elsewhere in the technical documentation, including the risk assessment. If the risk file assumes a particular level of performance while the study acceptance criteria use a different threshold, the inconsistency becomes a regulatory issue regardless of how well either document has been written.

Performance evaluation and risk management therefore need to evolve together.

Deciding what evidence belongs in the plan

Once the intended purpose, claims and risks are defined, the PEP can describe how each evidence requirement will be addressed.

For scientific validity, the available evidence may come from the scientific literature, recognised sources of scientific information or data generated specifically for the device or analyte relationship.

The analytical performance programme will generally cover the characteristics relevant to the device and intended purpose, together with the methods, standards where applicable, statistical approaches and acceptance criteria used to evaluate them.

Clinical performance requires a similar assessment of the available evidence. In some cases, existing clinical performance data may be sufficient. In others, new data will need to be generated.

Article 56 allows reliance on other sources of clinical performance data where that approach is duly justified; otherwise, clinical performance studies form part of the evidence-generation route. The decision therefore depends on the evidence already available, its quality and relevance, and the questions that remain unresolved.

This is one of the places where an effective PEP can prevent unnecessary development work. When the evidence question is defined precisely, it becomes easier to determine whether a literature review, analytical study, retrospective dataset, prospective clinical performance study or post-market activity is capable of answering it.

Acceptance criteria give the plan its discipline

An evidence programme needs a defined point at which results can be interpreted.

Annex XIII includes milestones and potential acceptance criteria within the performance-evaluation development programme. MDCG 2022-2 also discusses the parameters and criteria used to assess whether the available evidence supports the performance and benefit-risk considerations relevant to the device.

The appropriate criteria will vary considerably between IVDs. They may relate to analytical performance characteristics, clinical performance endpoints, evidence quality, state-of-the-art benchmarks or other device-specific objectives.

Their value lies partly in timing. Criteria documented before the final results are interpreted provide a clearer record of how the manufacturer intended to judge the evidence.

When those criteria appear only in a completed study report or in the final PER, a reviewer has less visibility into whether they shaped the original evidence strategy or emerged after the results were available.

That distinction can become particularly important when a result is close to a predefined performance boundary or when several evidence sources point in different directions.

Evidence gaps belong in the plan

Few development programmes begin with a complete body of performance evidence. A useful PEP records the gaps that remain and gives each of them a route towards resolution.

The response will depend on the nature of the uncertainty.

A gap in scientific validity may require additional literature analysis or other scientific evidence. An analytical uncertainty may lead to further bench or laboratory work. A question around clinical performance may require additional clinical data. A claim that cannot be adequately supported may need to be reconsidered.

Some questions can appropriately remain open for post-market performance follow-up when the regulatory and clinical context allows it.

What matters is that the gap has an identifiable consequence for the evidence programme.

A statement such as “additional clinical evidence may be required” offers little guidance unless the PEP also records what uncertainty has been identified, why the current evidence is insufficient and what information would allow the manufacturer to resolve it.

A practical Annex XIII checklist

A reviewer should be able to move through the PEP without repeatedly consulting other documents simply to understand its structure.

PEP areaWhat the PEP should make clear
Intended purposeWhat the device is intended to detect, measure, predict, monitor or otherwise establish
Device characteristicsWhich configuration and performance characteristics are being evaluated
Analyte or markerThe analyte, marker or other measurand relevant to the device
Intended use and populationIntended users, target groups, indications, limitations and contraindications
Metrological traceabilityApplicable reference materials or reference measurement procedures
GSPR mappingWhich requirements depend on scientific validity, analytical performance or clinical performance evidence
Risk linkageWhich risks depend on particular performance assumptions or evidence
Scientific validityHow the analyte or marker association will be established and maintained
Analytical performanceWhich characteristics will be assessed and according to which methods
Clinical performanceWhich evidence will support performance in the intended population and use context
State of the artThe benchmarks, standards, clinical context and relevant scientific background used to frame the evaluation
Acceptance criteriaHow the manufacturer will judge whether evidence is adequate
Evidence gapsWhich questions remain unresolved and how they will be addressed
Development programmeStudies, activities, milestones and sequence of evidence generation
PMPFHow post-market information will maintain and extend the evidence base
Non-applicable elementsThe justification for any Annex XIII requirement considered inapplicable

Annex XIII requires justification when an element of the PEP is considered inappropriate because of the specific characteristics of the device. A well-written justification therefore carries more weight than simply marking an item “N/A”.

Team-NB’s 2025 Best Practice Guidance for IVDR technical documentation offers a useful indication of how some Notified Bodies approach the presentation of this material. Its recommendations include addressing the Annex XIII Part A 1.1 elements clearly and making the relationship between scientific validity, analytical performance, clinical performance and state of the art visible in the submission. These recommendations represent Notified Body best practice, rather than additional legal requirements under the IVDR.

Building traceability into the PEP

A traceability matrix can be useful here, particularly for devices with multiple claims, biomarkers, intended populations or performance endpoints.

The table does not need to become another large controlled document. Its value comes from bringing information that is often dispersed across the technical file into one view.

Claim / intended-purpose elementGSPR / relevant riskScientific validity evidenceAnalytical performance evidenceClinical performance evidenceAcceptance criterionEvidence gapLifecycle action
Specific claim being supportedApplicable requirement and relevant risk referenceEvidence supporting the analyte-condition associationStudy or report supporting applicable analytical characteristicsEvidence supporting performance in the intended population and settingCriterion used to judge adequacyOutstanding uncertaintyStudy, PMPF activity, review or other planned action
Additional claim or intended-purpose elementAssociated requirement and riskRelevant SV evidenceRelevant AP evidenceRelevant CP evidenceDefined criterionRemaining gap, if anyFollow-up action

The exercise tends to reveal problems that are much harder to see when the documentation is reviewed report by report.

A clinical claim may have a strong analytical package with limited corresponding clinical evidence. An acceptance criterion may differ from an assumption used in risk management. A clinical performance study may include endpoints that are scientifically interesting but only loosely connected to the claims included in the intended purpose.

These are evidence-architecture problems. They can remain hidden even when the individual reports are technically competent.

State of the art as part of the evidence argument

State of the art gives the performance evaluation its external context.

It helps establish what is currently known about the clinical condition, analyte or marker, how the diagnostic question is addressed in current practice, which alternatives are available and what level of performance is clinically meaningful.

MDCG 2022-2 identifies state of the art among the considerations relevant to the performance-evaluation strategy, alongside factors such as intended purpose, device novelty, assay technology, target population, patient risk and the availability of reference materials or methods.

Its influence should be visible in the PEP.

If the state-of-the-art assessment identifies an established performance benchmark, that information may affect the choice of acceptance criteria. If current diagnostic practice has changed, the relevance of an older comparator may need to be reconsidered. Where new clinical evidence alters the understanding of an analyte or disease association, the scientific-validity strategy may also need to change.

A literature section that sits apart from these decisions contributes less to the performance evaluation than one whose findings can be traced into the evidence plan.

Where PEPs tend to lose coherence

Many weaknesses found during technical-documentation review originate at the interfaces between documents rather than within the documents themselves.

Strong reports with weak traceability

A substantial volume of evidence does not automatically produce a clear performance evaluation.

When claims, risks and evidence sources cannot be followed through the technical file, the reviewer has to reconstruct the manufacturer’s reasoning. Questions then arise about evidence sufficiency even where several studies have already been completed.

A traceability review can sometimes resolve this without generating new data.

Evidence programmes developed in parallel

Scientific affairs, laboratory teams and clinical teams often work independently during development. That division of labour is entirely reasonable, but the resulting evidence streams need to converge before the PEP is finalised.

Differences in intended purpose wording, device configuration, populations, endpoints or performance claims can otherwise travel through separate reports unnoticed.

The PEP provides a useful point for reconciling them.

Acceptance criteria added too late

Study teams naturally focus on protocol-specific endpoints and analytical teams on individual performance characteristics. The PEP has to bring those criteria back to the regulatory claims they support.

When the evidence threshold is established after results are available, the rationale becomes more difficult to defend prospectively.

State-of-the-art work that remains isolated

A detailed literature review can still have limited regulatory value if none of its conclusions influence the evidence strategy.

The useful outputs are the ones that change or confirm a decision: a benchmark, comparator, clinical practice assumption, evidence gap, acceptance criterion or risk-management consideration.

Gaps without a defined consequence

Technical files sometimes acknowledge that evidence is limited without explaining how the manufacturer intends to deal with that limitation.

A documented gap should lead somewhere. The next step may be new evidence generation, PMPF, a narrower claim, further analysis or a justified decision that no additional activity is needed. The PEP should preserve that reasoning.

Performance evaluation after market entry

The PEP continues to have a role once the device reaches the market.

Article 56 requires performance evaluation and its documentation to be updated throughout the device lifecycle using information obtained through PMPF and the post-market surveillance system. For Class C and D devices, the PER must be updated when necessary and at least annually.

Catarina Sepulveda describes the practical consequences of this lifecycle model:

“The PEP should be treated as a living document. PMS, PMPF, complaints, new literature, changes in state-of-the-art, device changes and emerging risks should feed back into the performance evaluation and trigger reassessment where necessary. The PER then becomes the consolidated output demonstrating that the evidence remains current, coherent and sufficient throughout the device lifecycle.”

Catarina Sepulveda | IVD Regulatory Director at MDx CRO

The implications vary from device to device. New literature may change the scientific understanding of a biomarker. PMPF may provide additional evidence in a subgroup that was limited at initial conformity assessment. Complaints may expose a performance issue associated with a particular setting or specimen type. A software or assay change may alter the evidence applicable to the current configuration.

Each of these developments can affect the performance-evaluation strategy.

A lifecycle process does not require rewriting the entire PEP whenever a new publication or complaint appears. It requires a controlled mechanism for deciding whether the new information changes an assumption, evidence gap, claim, risk, acceptance criterion or planned activity.

That decision, and its rationale, should remain visible in the technical documentation.

When the evidence strategy changes

Several events deserve particular attention because they can alter the scope of the PEP:

  • changes to the intended purpose or claims;
  • modifications to the device that affect performance;
  • new scientific evidence concerning the analyte or clinical condition;
  • relevant changes in state of the art;
  • new or emerging risks;
  • PMPF findings;
  • trends from complaints or post-market surveillance;
  • evidence that an acceptance criterion has not been met;
  • new studies undertaken to address an existing evidence gap.

The appropriate response may be relatively small, such as updating a literature review or refining an existing analysis. In other cases, the change can affect study design, risk management, labelling or the broader performance-evaluation strategy.

Documenting these decisions within the lifecycle of the PEP helps preserve the history of why the evidence programme developed as it did.

Reviewing a PEP before submission

One useful final review is to read the PEP from the perspective of someone who was not involved in developing the device.

That reader should be able to understand:

  1. What the device is intended to do and which claims are being made.
  2. Which GSPRs and risks depend on performance evidence.
  3. How scientific validity supports the intended purpose.
  4. Which analytical performance characteristics have to be demonstrated.
  5. Which clinical performance evidence is required for the intended population and use.
  6. How the manufacturer intends to judge whether that evidence is adequate.
  7. Where the current evidence base remains incomplete.
  8. Which activities will address those gaps.
  9. How state of the art has influenced the evidence strategy.
  10. How PMPF and other post-market information will be incorporated after market entry.

Where those answers are scattered across the technical documentation, the PEP probably needs further work.

Where they follow a visible line from intended purpose through evidence to conclusion, the document gives the reviewer a much clearer account of the manufacturer’s performance-evaluation strategy.

The PEP as the architecture of the performance evaluation

The regulatory value of a PEP lies in the structure it gives to the evidence.

Scientific validity establishes the scientific basis for the analyte or marker. Analytical performance shows how the device performs at the analytical level. Clinical performance addresses how its results relate to the relevant clinical condition, population and intended use. Risk management, state of the art and predefined criteria provide the context in which those findings are interpreted.

The PEP brings these elements into a common plan and keeps that plan aligned as the evidence changes.

For manufacturers, this has a practical advantage as well as a regulatory one. Evidence gaps become visible earlier. Studies can be designed around questions that genuinely matter to the intended purpose. Teams working on separate parts of the technical documentation have a common reference for claims and evidence requirements. New post-market findings can be assessed against an established framework instead of being added to the file in isolation.

The result is a performance evaluation whose reasoning can be followed from the initial claim to the evidence supporting it and, eventually, to the conclusion recorded in the PER.

Manufacturers preparing a new PEP, restructuring an existing performance evaluation or responding to Notified Body questions can work with MDx CRO across the complete evidence pathway, including scientific validity, analytical and clinical performance, PMPF and integration of the resulting evidence into the IVDR technical documentation.

Learn more about our Regulatory Affairs and Technical Documentation services.

Frequently Asked Questions

What is an IVDR Performance Evaluation Plan?

The Performance Evaluation Plan describes how an IVD manufacturer will generate, collect and assess the evidence needed to support the device’s intended purpose. Under Annex XIII, it covers the relevant device characteristics, GSPRs, scientific validity, analytical and clinical performance, state of the art, methods, acceptance criteria, evidence-development activities and PMPF.

What is the difference between a PEP and a PER under the IVDR?

The PEP is prepared as the planning framework for performance evaluation. It defines the evidence strategy, methods, criteria, gaps and planned activities. The Performance Evaluation Report consolidates and assesses the evidence generated through that process and records the resulting conclusions.

Does every IVD need a clinical performance study?

The evidence route depends on the device and the clinical performance data already available. Article 56 allows manufacturers to rely on other sources of clinical performance data where that approach is duly justified. Where the available evidence cannot adequately support clinical performance, additional clinical performance data may need to be generated.

What should be included in an IVDR PEP?

Annex XIII Part A 1.1 sets out the elements of the PEP. These include the intended purpose, device characteristics, analyte or marker, intended use, relevant GSPRs, methods for assessing analytical and clinical performance, state of the art, benefit-risk parameters and criteria, the performance-evaluation development programme and PMPF planning. Any element considered inappropriate for a particular device requires justification.

How often should a Performance Evaluation Plan be updated?

The IVDR treats performance evaluation as a continuous lifecycle process and requires the performance evaluation and its documentation to be updated with relevant PMPF and post-market surveillance data. It does not prescribe a universal annual update frequency specifically for the PEP. For Class C and D IVDs, the PER must be updated when necessary and at least annually.

How are scientific validity, analytical performance and clinical performance connected in the PEP?

The three evidence streams address different aspects of device performance but contribute to the same clinical evidence package. The PEP should make their relationship to the intended purpose, claims, applicable GSPRs, risks and predefined acceptance criteria visible. This allows the manufacturer and the reviewer to follow how individual studies and reports contribute to the final performance-evaluation conclusions.

Written by:

Catarina Sepúlveda

In Vitro Diagnostics (IVD) Regulatory Affairs IVDR

Catarina Sepúlveda is an IVD Director and regulatory affairs specialist with over 10 years of experience in in vitro diagnostics (IVDs), medical devices, and life sciences. She has extensive expertise in IVDR conformity assessment, regulatory strategy, technical documentation, and quality management systems, having held leadership roles within both industry and a Notified Body. Catarina specialises… Read more…

View the Author Profile
Industry Insights & Regulatory Updates