Methodology
How every ranking on this site is decided.
The method behind every list on this site is fixed and published here in full before any provider is ordered, with the limits of the evidence stated.
A provider may be strong overall and still rank below a specialist on a product page.
Every service ranking flags its policy risk, with Google's own rules cited.
The scoring system has four visible outputs:
-
Eligibility status Rated or Unrated, with the check and evidence shown.
-
Provider Quality Score (PQS) universal provider quality, 0–100.
-
Service Fit Score (SFS) fit for one specific link-building service, 0–100.
-
Overall Score (OS) the page-specific total, 0–100.
Every published ranking must show PQS, SFS, OS, confidence, methodology version, score date, and evidence cutoff.
PQSSFSOSConfidence Methodology versionScore dateEvidence cutoff
“Provider Quality Score” replaces the working term “master score.” It describes what the number measures without implying that it overrides product fit.
M1.0 weights are editorial proposals pending direct input from real buyers.
Stages and controls
Stage 0
Eligibility checks
Eligibility checks set the minimum standard every ranked provider must meet. A failed check cannot be offset by price, speed, volume, or another high score.
| Check | Rated-provider requirementRated | Failure outputUnrated |
|---|---|---|
| E1Identifiable provider | A verifiable trading identity and a working way to make contact | Unrated — identity not verified |
| E2Method disclosure | Per service, states whether placements are paid, exchanged, earned, or another disclosed method; where links are paid, states whether they are qualified with a rel="sponsored" or rel="nofollow" attribute (link qualification) | Unrated — acquisition method not disclosed |
| E3No systematic prohibited method | No established systematic or uncorrected use of Google-defined link-spam practices such as automated link generation, unqualified paid links sold for ranking credit, or expired-domain abuse | Unrated — policy-risk check not met |
| E4Evidence integrity | No established systematic or uncorrected fabrication, buying reviews from customers screened to be positive, presenting reviews the provider controls as independent, or significantly misdescribing the service | Unrated — evidence-integrity check not met |
| E5Evidence sufficiency | At least one live, traceable example of delivered work for the rated service, and confidence of at least 0.50 in both the Provider Quality Score and the Service Fit Score | Unrated — insufficient evidence to rank |
How conduct evidence is resolved
The same conduct cannot fail an eligibility check and also take a score adjustment.
Established, systematic or significant, and uncorrected
The eligibility check fails; the provider is Unrated.
Established, isolated, and corrected with evidence
Score the current observable state in the relevant PQS or SFS criterion and publish a dated correction note. No separate penalty is applied.
Contested or uncertain
Reduce evidence confidence on the affected criterion, or hold the provider Unrated pending resolution if the uncertainty concerns an eligibility check.
M1.0 deliberately omits a separate integrity penalty multiplier (a factor that would multiply the whole score down). Its proposed values had no buyer evidence and could score the same incident twice: once in a criterion and again after combination.
Selling paid placements does not itself fail an eligibility check. Google states paid links are not a policy violation when qualified with rel="sponsored" or rel="nofollow"; of the two, sponsored is preferred. What we rate is disclosure, whether paid links carry one of them, and whether the service is described accurately.
Scenario B — a failed check cannot be offset
Synthetic providers test the model; they are not market claims.| Provider | Check | PQS input | SFS input | Published result |
|---|---|---|---|---|
| Risky Volume Seller | E3 failed: established systematic unqualified paid links sold for ranking credit | 95 | 97 | Unrated — policy-risk check not met |
Result: price, speed, volume, or high scores cannot restore eligibility.
Five checks come before any score, and a provider that fails one is Unrated: it leaves the ranking, with the failed check and its evidence shown.
Stage 1
Provider Quality Score
PQS scores the provider as an organisation, regardless of which service page is being ranked.
| ID | Universal criterion | What is scored | Weight |
|---|---|---|---|
| Q1 | Evidence of delivered work | Whether live delivered work can be traced to the provider, is current and relevant, and can be independently inspected | |
| Q2 | Quality standard and publisher vetting | Whether a written standard for accepting or rejecting publishers exists and observed work follows it; product-specific thresholds stay in SFS | |
| Q3 | Policy compliance and method disclosure | Disclosure of how links are acquired, whether paid links are qualified with rel="sponsored" or rel="nofollow", and consistency between stated and observed methods | |
| Q4 | Delivery reliability and remedy | Evidence of on-time, complete delivery, replacement or remedy terms, and how removed placements are handled | |
| Q5 | Commercial transparency | The price or how it is worked out, inclusions, exclusions, contract terms, and ownership of deliverables are stated before purchase | |
| Q6 | Reporting and outcome discipline | Reporting separates delivered work from outcomes and states measurement limits | |
| Q7 | Client evidence and references | References and case evidence can be traced to real clients and verified; warmth or volume of testimonials is not scored | |
| Q8 | Organisational stability and capacity | Evidence of continuity, relevant personnel, capacity, and ability to support the stated service scope | |
| Q9 | Service model and responsiveness | Named ownership, a clear contact route, a way to escalate problems, and stated response expectations | |
| Total |
The fixed criterion scale
Every criterion has a written scoring guide, its rubric, before any provider is scored. The five fixed points on the scale are its anchors:
Eligibility is handled by the checks, not by calling 0 “minimum acceptable.” Missing evidence never becomes 0.
These nine criteria score the provider as an organisation, not the service you buy from it. The criterion weights are fixed for the methodology version, and every scoring guide is written before any provider is scored.
Stage 2
Service Fit Score
SFS is worked out separately for each service category.
Product-rubric derivation
For each service page:
- Define the buyer’s job in one sentence.
- List the service-specific attributes that vary among competent providers.
- Delete anything already scored in PQS.
- Identify the relevant Google policy that bears on the service, or state that none applies.
- Keep 5–8 criteria.
- Write observable 0/25/50/75/100 anchors before scoring providers.
- Name the PQS criterion closest to it and state what remains that is genuinely service-specific.
- Set weights by comparing each criterion’s full written 0→100 range across documented representative buyer scenarios, not across the providers currently listed.
- Fix the anchors and weights for that service methodology version.
The boundary with PQS
Permitted service-specific dimensions
- quality or value relative to the product being bought;
- relevance of evidence to that service;
- how each channel works;
- service-specific outcome definitions;
- buyer context such as sector, geography, language, scale, or ownership requirements.
Universal-only dimensions that must not be rescored in SFS
- whether a vetting standard exists;
- disclosure of how links are acquired and whether paid links carry rel="sponsored" or rel="nofollow";
- general delivery reliability and remedies;
- price and contract disclosure;
- general reporting quality;
- whether references can be verified;
- organisational stability and responsiveness.
A price or value criterion may not exceed 10% of SFS.
Every SFS criterion must carry a boundary note naming the PQS criterion closest to it and stating what remains that is genuinely service-specific. If the reviewer must consult that PQS score to assign the criterion, the criterion is merged or deleted.
Fit is separate from quality: this score covers only what changes with the service you're buying, and anything already measured in PQS stays there.
Stage 3
Overall Score
M1.0 gives equal weight to both scores and combines them with a geometric mean. Scores are calculated at full precision and displayed to one decimal place.
Equivalent general form
Why 50/50
- Universal quality and service fit are both necessary.
- Two independently run large language model (LLM) probes disagreed materially on the relative weight of provider quality and service fit, so model opinion cannot settle the split.
- No study of buyer preferences settles it.
- Equal weighting is the starting point that assumes the least, and it is easy to explain.
Why not simple addition
- A weak side pulls the total down; a strong side cannot fully make up for it.
- It allows a strong specialist to outrank a stronger general provider on the specialist’s service page.
- It improves when either score improves, stays within 0–100, and does not depend on which other providers are listed.
If PQS or SFS is 0, OS is 0. Because missing evidence shrinks toward 50, a zero can happen only when evidence produces zeros across a whole score. That result triggers a manual review of the checks and scoring guides before publication.
Scenario A — product fit can change order
Synthetic providers test the model; they are not market claims.Two synthetic providers, two pages. The generalist wins the general page on quality, and the specialist wins the digital-PR page on fit.
General page
Ranked by OS descending
Verified Generalist
PQS 88SFS 80
83.9
OS
Product Specialist
PQS 72SFS 74
73.0
OS
Digital-PR page
Ranked by OS descending
Product Specialist
PQS 72SFS 92
81.4
OS
Verified Generalist
PQS 88SFS 52
67.6
OS
Result: the higher-quality generalist wins the general page; the specialist wins the digital-PR page. PQS remains visible, so the fit reversal does not erase the universal-quality difference.
The total is the geometric mean of the two scores, so a weak side pulls the total down rather than hiding behind a strong one. The 50/50 split is an editorial starting point, disclosed as such on every ranking.
Control
Evidence confidence and missing data
Each criterion receives a best estimate of its score, written sᵢ, and a certainty factor for the evidence behind it, written λᵢ.
Evidence certainty
| Evidence certainty | λ | Operational meaning |
|---|---|---|
| High | 1.00 | Current, directly traceable to the provider, open to independent inspection, sufficient sample |
| Moderate | 0.85 | Mostly direct evidence with one known, limited weakness |
| Low | 0.60 | Thin, indirect, provider-selected, stale, or partly conflicting evidence |
| Very low | 0.35 | Serious uncertainty; usable only with an explicit caveat |
| None | 0.00 | No usable evidence |
Missing or uncertain evidence shrinks the estimate toward neutral 50.
Confidence in the Provider Quality Score, written C(PQS), is the weighted average of criterion certainty. The same calculation with the service weights gives confidence in the Service Fit Score, written C(SFS). Published confidence for a product ranking is the weaker of the two.
Overall confidence
Insufficient evidence — Unrated under E5
Below 0.50
Low
0.50–0.699
Moderate
0.70–0.849
High
0.85–1.00
Confidence is shown beside the score and never multiplied into it. Q1 scores the provider’s evidence practice; confidence measures the researcher’s certainty about each measurement. They measure different things.
Scenario C — missing evidence is not a bad score
Synthetic providers test the model; they are not market claims.| Provider | PQS estimate | SFS estimate | C(PQS) | C(SFS) |
|---|---|---|---|---|
| Sparse-Evidence Provider | 78 | 82 | 0.43 | 0.46 |
Published result Unrated — insufficient evidence to rank
Result: the provider is not called poor. It is excluded because the estimates are not sufficiently evidenced.
Confidence tells you how well evidenced a score is. Missing or uncertain evidence pulls the estimate toward neutral 50, and the label is shown beside the score, never multiplied into it.
Control
Ranking, ties, and independent scores
- Rank eligible providers by OS, highest first.
- A gap below 0.5 points is a tie.
- Tied providers receive the same rank and are displayed alphabetically.
- No undisclosed tie-breaker is permitted.
- Adding, removing, suspending, or rescoring one provider changes no other provider’s criterion scores, PQS, SFS, confidence, or total.
- Anchors, weights, certainty values, eligibility checks, and the formula stay fixed within a major methodology version.
Scenario D — how a tie is decided
Synthetic providers test the model; they are not market claims.| Provider | PQS | SFS | OS |
|---|---|---|---|
| Provider Delta | 80 | 75 | 77.46 |
| Provider Echo | 76 | 79 | 77.49 |
Gap
0.03
Tie threshold
0.5 points
Result: the 0.03-point gap is below the 0.5 tie threshold. Both receive the same rank and are displayed alphabetically.
Scenario E — adding a provider changes no other score
Adding a synthetic provider with PQS 100 and SFS 100 changes none of the inputs or outputs above. The new provider ranks above eligible entries; all existing scores and their order relative to one another remain unchanged.
Synthetic providers test the model; they are not market claims.
Providers are ordered by Overall Score, and a gap under 0.5 points is a tie: shared rank, alphabetical order, no hidden tie-breaker. Adding or removing a provider changes nobody else's score.
Control
Evidence and governance
The architecture is evidence-backed; the numeric weights are not buyer-survey findings.
Verified source support
- Google defines link spam, paid-link qualification, automated link generation, site-reputation abuse, and expired-domain abuse.
- The UK government’s multi-criteria decision analysis (MCDA) manual supports fixed scoring scales, adding weighted criterion scores into a total, setting weights by comparing each criterion’s full range, explicit acceptability checks, multiplying where a weak side must not be offset, sensitivity analysis, and removing double counting.
- The US Federal Acquisition Regulation (FAR, Subpart 15.3) supports disclosing evaluation factors in advance, weighing price against quality, checking evidence of past performance for currency, relevance, source and context, and treating a provider neutrally when no past-performance evidence exists.
- US Federal Trade Commission (FTC) rules support clear disclosure and honest-review controls when their factual and jurisdictional conditions apply.
- Regulation (EU) 2019/1150 is used voluntarily as a transparency model: disclose main ranking parameters and why they matter without exposing a gameable implementation.
Evidence gap
No primary study was found measuring how buyers of link-building or marketing agencies weight provider-selection criteria. Therefore every weight, threshold, certainty value, review schedule, and tie threshold in M1.0 is an editorial starting point. They must be disclosed as such and tested against direct input from real buyers before being described as buyer-validated.
Version, score date, and evidence controls
Methodology version
M1.0 candidate
Score snapshot
S<M-version>.<YYYY-MM-DD>
Score date
<YYYY-MM-DD>
Evidence cutoff
<YYYY-MM-DD>
Tie
A gap below 0.5 points
Status
locked for methodology-page concepts; not yet calibrated on real providers
Versioning
- Methodology: M<major>.<minor>.
- Major change: a change to the criteria, weights, anchors, eligibility checks, certainty values, the formula that combines the scores, or the tie rule.
- Minor change: clarification proven not to change any score.
- Score snapshot: S<M-version>.<YYYY-MM-DD>.
- A major change requires every affected provider to be rescored before the new ranking publishes.
- Keep previous methodology versions and score snapshots archived, and publish a change list that separates provider changes from methodology changes.
Review schedule
- Quarterly: re-check sampled placements, current commercial terms, and eligibility-check evidence.
- Annually: a full provider rescore under the current fixed weights.
- Event-triggered: an ownership change, a complaint backed by evidence, new eligibility evidence, or a Google policy change.
- Evidence older than 18 months gets a lower confidence rating for age; evidence older than 36 months cannot satisfy Q1 or a service evidence requirement.
Corrections and appeals
- Publish a free way to submit evidence and appeal.
- Acknowledge within 5 working days; provide a substantive response within 20 working days.
- Record the evidence, decision, reviewer, and date.
- Successful appeals produce a corrected score and dated public correction note.
- Contested eligibility evidence remains pending, not assumed true.
Conflicts and commercial relationships
- Disclose any significant provider relationship at the ranking, subject to facts and applicable law.
- Keep a register of relationships and record when a reviewer steps aside because of one.
- Paid inclusion, if ever introduced, must not change scores or order and must be visually distinguishable.
- Obtain legal review before making jurisdiction-specific compliance claims. This specification is not legal advice.
Required sensitivity checks before real publication
Synthetic validation proves the formula behaves as designed; it does not prove the weights match what buyers value. Before ranking real providers:
- Vary every PQS and SFS weight within its documented range and report whether the top three places change.
- Sweep the weight given to provider quality from 0.30 to 0.70, with service fit taking the remainder, and report the points where any two providers swap order.
- Compare the geometric mean with simple addition and explain any provider that moves three or more places.
- Remove evidence from strong and weak criteria; the adjusted scores must move smoothly toward 50, with no jumps.
- Test each eligibility check with systematic, isolated-and-corrected, and contested versions of the same evidence; each version must land in exactly one state.
- Have two reviewers independently score at least 20% of provider–criterion pairs; publish how often they agree, both as a simple rate and adjusted for chance agreement (Cohen’s kappa).
- Re-run the PQS/SFS boundary test after any criterion wording change.
- Simulate the cheapest way to gain 10 points on each criterion; the action must improve buyer evidence or outcomes, not merely presentation.
Publication requirements
Every ranking page must visibly provide:
- the provider’s eligibility status;
- PQS, SFS, OS, and confidence;
- criterion definitions, anchors, fixed weights, and reasons for relative importance;
- evidence cutoff, score date, and methodology version;
- eligibility, correction, and conflict disclosures that affect interpretation;
- primary sources for policy claims;
- the fact that M1.0 weights are editorial proposals pending direct input from real buyers.
Structured data (machine-readable page markup) may represent only facts and scores visible on the page. No methodology content may claim that a language model uses, prefers, or rewards these fields; the independent model probes test presentation only and are not buyer evidence or scoring authority.
This is not legal advice; applicability depends on facts and jurisdiction.
No primary study was found measuring how buyers of link-building services weight provider-selection criteria. Every M1.0 weight is an editorial starting point, published as such and not buyer-validated.
The weights, anchors, checks and formula on this page stay fixed within a methodology version. Changes arrive as new versions, with affected providers rescored and the differences published.