Proposal Score Interpretation Guide for EU Bids
A score of 4.0 can look reassuring in a draft review until it is read as an evaluator will read it: not as praise, but as evidence that a criterion is convincingly addressed with identifiable residual weaknesses. In a competitive call, those weaknesses may separate a fundable proposal from one that clears every threshold but falls below the cut-off. This proposal score interpretation guide explains what the official 0-5 scale does and does not tell an EU funding applicant before submission.
The first discipline is to separate admissibility and eligibility from award criteria. A persuasive Excellence section cannot repair an inadmissible submission, an ineligible consortium, a missing annex or a breach of page limits. Equally, clearing those preliminary checks does not earn points. Points are awarded only against the published evaluation criteria and their sub-criteria for the relevant call and type of action.
Read the score against the call, not against your expectations
Horizon Europe, Digital Europe and Erasmus+ calls use related but not identical evaluation forms. The criterion labels may be familiar, but their sub-criteria, thresholds, weighting and tie-break rules can differ. A Research and Innovation Action is not assessed through precisely the same lens as an Innovation Action; a Digital Europe procurement-oriented call may put different pressure on deployment and capacity; Erasmus+ actions use their own award criteria and thresholds.
Many Horizon Europe calls apply a threshold of 3 out of 5 for each criterion and 10 out of 15 overall. That pattern is common, not universal. Some calls use different thresholds, weights or a two-stage process. The published call documentation is controlling. Treat any generic benchmark, including a previous evaluation summary report from another call, as context rather than a rule.
A total score is therefore a compressed signal. It cannot show whether the proposal has one threshold-threatening defect, whether all criteria are merely adequate, or whether a weak section sits in a heavily weighted criterion. Read the individual criterion scores first, then the total.
What each score band means in practice
The European Commission’s familiar descriptors are exacting. They should be read literally rather than as ordinary workplace feedback.
A score of 5 means the proposal successfully addresses all relevant aspects of the criterion. Any shortcomings are minor. This does not mean every sentence is perfect. It means an evaluator can locate the necessary evidence, see that it is coherent, and find no weakness material enough to affect the judgement.
A 4 means the criterion is very good, but has a small number of shortcomings. The distinction matters. A 4 often reflects a credible plan whose evidence is incomplete in a limited area: a KPI without a baseline, an exploitation route without a named owner, or a work package dependency implied rather than managed. In a crowded call, several 4s may be commercially insufficient even though the draft is plainly competent.
A 3 means the criterion is good but has a number of shortcomings. This is not a near-5 with harsher wording. It usually indicates that the proposed approach is plausible but key elements are thin, untested, internally inconsistent or not sufficiently tailored to the call. A threshold score should trigger a targeted rewrite, not relief.
A score of 2 identifies significant weaknesses. The evaluator can see elements of an answer, but important parts of the criterion are inadequately addressed. A score of 1 means serious weaknesses, while 0 means the criterion is not addressed or cannot be assessed because the information is missing or incomplete.
Where the form permits half-point scoring, a 3.5 is not a neutral middle. It records a real judgement between descriptors and should be interrogated for the evidence that prevented a 4. Read the evaluator brief and the scoring rules for the specific call before assuming that an internal score format matches the official one.
Interpret the criterion, not the heading
A common internal review error is to score chapters. Evaluators score criteria. Evidence is often distributed across the technical narrative, impact pathway, implementation tables, partner descriptions, risk register and budget justification. The question is whether those fragments add up under the relevant sub-criterion.
Excellence: is the method credible and sufficiently specified?
Excellence scores fall when ambition substitutes for method. The problem may be well framed and the objectives attractive, yet the evaluator cannot determine how an objective will be achieved, measured or validated. Unsupported claims of novelty, vague methodology, weak treatment of interdisciplinarity or open science where relevant, and an unconvincing state-of-the-art comparison all create scoring exposure.
Look for statements that begin with “will deliver”, “will enable” or “is uniquely positioned”. Each needs nearby support: a method, comparator, assumption, dataset, protocol, decision rule or measurable endpoint. Evidence elsewhere in the proposal may help, but only if an evaluator is likely to find and connect it within the time available.
Impact: can the route from outputs to outcomes withstand scrutiny?
Impact is often over-scored by drafting teams because it contains the proposal’s most persuasive language. Evaluators look for the route beneath it. Who adopts the result, what changes for them, what barriers exist, who owns exploitation, and how will progress be measured?
A credible dissemination plan is not automatically a credible exploitation plan. A list of audiences is not a market, policy or uptake pathway. Similarly, an expected impact copied from the work programme is not a contribution to that impact. The proposal needs a defensible chain from project activities and outputs to outcomes, with indicators, baselines or targets where the call expects them, and responsibilities that match consortium capability.
Implementation: can this consortium deliver what the narrative promises?
Implementation weaknesses often sit in the joins between work packages. The work plan may look complete in isolation, while task dependencies, milestones, resources and governance tell a different story. Evaluators test whether the Gantt chart, effort allocation, deliverables, risks and decision-making arrangements describe one executable project.
Warning signs include partner-month allocations that do not match technical responsibility, milestones that merely repeat deliverables, risk mitigations that are aspirations rather than contingencies, and governance bodies with no authority or escalation route. Consortium capacity is also tested here or through a dedicated sub-criterion, depending on the form. A strong organisation profile does not compensate for an unexplained role in the proposed work.
Treat disagreement as diagnostic evidence
Two evaluators can read the same passage differently without either being careless. One may infer a credible implementation detail from a technical description; another may refuse to award credit because it is not explicit. In a real panel, that disagreement is resolved through discussion and consensus, not by simply averaging private scores.
For an applicant, the disagreement is valuable. It identifies passages that rely on charitable interpretation. If a draft assessment records materially different readings of one criterion, inspect the cited evidence and ask a practical question: what would make the stricter reading impossible? Usually the answer is not more prose. It is a named responsibility, quantified assumption, cross-reference, decision point or clearer claim boundary.
BidShark’s automated assessment keeps independent specialist readings separate before adjudication and reports material disagreement rather than hiding it inside an average. That is not a human panel verdict, and it cannot predict the actual panel’s score. It is a controlled way to expose ambiguity before the document becomes fixed.
Turn a score into a revision order
Do not revise in descending order of annoyance. Revise by score risk and consequence. First address any criterion below its threshold. Next address the evidence gaps that prevent a 4 becoming a 5 in a criterion likely to determine ranking, taking account of weighting where the call applies it. Then correct cross-document inconsistencies that can damage more than one criterion.
For each finding, identify the exact passage, the relevant sub-criterion, the missing evidence and the smallest credible change. “Strengthen impact” is not an instruction. “Name the uptake owner for Output 3, specify the regulatory dependency, and align the KPI target with the baseline in Section 2.3” is an instruction that can be assigned before the deadline.
Do not attempt to argue every possible evaluator objection into the proposal. Page limits make that self-defeating. The objective is a document in which the necessary evidence is easy to locate, internally consistent and proportionate to the claim. You only get one chance to submit. Make the final reading adversarial enough that the real evaluators have less room to doubt what you mean.