Deterministic matching vs fuzzy logic in sanctions screening: the evidentiary burden

Deterministic matching vs fuzzy logic in sanctions screening: the evidentiary burden

Standard semantic and fuzzy matching approaches in PEP and sanctions screening generate false positive rates that create compliance vulnerabilities, not reduce them. Designing an auditable screening data pipeline requires understanding why the matching logic matters as much as the list coverage, and how to build a defensible record for high-risk country list updates.

13 min read

This article is for informational purposes only and does not constitute legal advice. Consult a qualified legal professional for advice specific to your situation.

  • The false positive problem in sanctions screening is not primarily a data quality problem. It is a matching logic problem. The EU's consolidated sanctions list, the UK OFSI list, and OFAC's SDN list all contain reasonably well-structured data. The false positive rates that plague operational screening programmes arise from fuzzy and semantic matching approaches that generate approximate matches across non-identical strings without a defined evidentiary threshold for what constitutes a hit requiring escalation. The result is a high volume of alerts that compliance teams must triage manually, which is both operationally expensive and structurally unreliable: alert fatigue produces missed matches alongside cleared false positives.
  • Deterministic matching is not inherently superior to fuzzy matching. The question is whether the matching algorithm's output is auditable. A fuzzy matching system that applies defined, documented, and consistent rules for when an approximate match becomes an alert, and that retains a record of the algorithm parameters applied to each match decision, can be auditable. A deterministic matching system that uses rigid exact-match logic produces high precision but excessive false negatives, missing genuinely sanctioned individuals whose names are transliterated differently from the listed form. The evidentiary burden requires a system whose matching logic can be explained and defended to a supervisor, not a system that minimises alerts by either extreme.
  • High-risk country list updates are where operationally mature screening programmes fail most visibly. The EU's consolidated sanctions list is updated irregularly in response to Council decisions, while the EU's AML high-risk third country list is updated by Commission delegated regulation. Both types of update must propagate through the screening system before they take legal effect. A screening configuration that is updated after the effective date has a gap during which transactions are being screened against a stale list. For sanctions, that gap is a potential sanctions evasion failure. For AML high-risk countries, it is an enhanced due diligence gap. The two list types have different legal consequences and require different escalation procedures; conflating them in a single screening workflow creates compliance gaps in both.
  • The AML Regulation and the AMLA supervisory framework are increasing the evidentiary standard for screening decisions. AMLA's coordination of national AML supervisors and the shared database of material supervisory findings mean that a screening programme found inadequate in one member state is increasingly visible to NCAs across the bloc. The defensibility of the screening output, meaning the ability to explain the matching logic, the alert triage procedure, the decision record for each cleared alert, and the update process for each list change, is becoming a primary supervisory focus, not a secondary one.

Why screening generates the compliance problem it is meant to solve

Sanctions screening and PEP screening are designed to be compliance controls. Their function is to identify customers or counterparties who are subject to legal prohibition or who require enhanced scrutiny before a firm engages with them. The legal requirement is unambiguous: EU credit institutions and financial firms may not process transactions involving designated persons or entities, and must apply enhanced due diligence to politically exposed persons and their associates.

The operational problem is that screening systems, as typically implemented, generate alerts at a rate that the compliance function cannot triage accurately. A payment firm processing fifty thousand transactions per day with a screening system generating a two percent false positive rate produces one thousand alerts per day. At that volume, human review is rarely thorough review. Alert fatigue, the cognitive and operational consequence of triaging more alerts than the team can meaningfully assess, is a well-documented phenomenon in sanctions compliance, and its consequence is not that compliance teams are working less hard. It is that the review process becomes pattern-matching against alert characteristics rather than substantive assessment of whether a given alert represents a genuine match.

The irony is that a screening programme generating a high volume of false positives is not a conservative programme in any meaningful sense. A compliance function that clears nine hundred and ninety alerts a day to get to ten genuine matches is not reviewing those nine hundred and ninety alerts with the same rigour it would apply to a triage queue of fifty. The false positives are not creating safety margin. They are consuming the operational capacity that genuine matches require.

This is the screening paradox: the matching logic that generates alerts is supposed to reduce compliance risk, but when it generates alerts at a rate incompatible with thorough review, it increases the systemic risk of genuine matches being missed.

Matching logic: why the algorithm choice defines the compliance output

Screening systems match customer and transaction data against list entries using one of three broad matching approaches, or some combination of them. The choice of approach, and the calibration of the parameters within each approach, determines the volume and character of the alert output.

Exact matching

Exact matching compares strings character by character. An alert is generated only when the customer name field contains an identical string to a listed name or alias. This approach produces very few false positives: a customer named "Ali Ahmed" will not generate an alert against a listed entry for "Ali Ahmad" unless both spellings appear in the list's alias entries.

The compliance problem with exact matching is false negatives. Sanctioned individuals and entities appear on lists with names transliterated from Arabic, Russian, Chinese, or other scripts into the Latin alphabet using varying conventions. The same individual may be listed under multiple transliterations, and the transliteration used in the customer record may not match any of the listed forms. Exact matching against a customer whose name in the firm's system is "Muammar Gaddafi" will miss a list entry for "Moammar Qaddafi" unless the list contains both forms as aliases.

For AML screening, where the matched parties are PEPs rather than designated persons, the false negative risk of exact matching is particularly material. PEP lists are compiled from public records across multiple languages and transliteration conventions, and a PEP whose name in the firm's onboarding system reflects the customer's own spelling of their name in a Latin-alphabet passport may not match the form used in the reference list.

Fuzzy matching

Fuzzy matching uses string similarity algorithms to generate alerts when a customer name is sufficiently similar to a listed name, even if not identical. The commonly used algorithms include Levenshtein distance (which counts the minimum number of character edits required to transform one string into another), Jaro-Winkler similarity (which weights transpositions at the beginning of strings more heavily than at the end), and phonetic algorithms such as Soundex and Metaphone (which match strings that sound similar when spoken in English).

Each algorithm has a similarity threshold parameter: a score above which a comparison generates an alert and below which it does not. The threshold is the primary calibration lever, and it is where most false positive problems originate. A Levenshtein distance threshold set to flag any pair of strings differing by two or fewer characters will flag "Ali Ahmed" against "Ali Ahmad" (one character difference), against "Ali Ahmet" (one character difference), and against "Al Ahmed" (one character difference). If the list contains several hundred common Middle Eastern names, a customer base with significant Middle Eastern representation will generate alerts at a rate that is operationally unmanageable at any reasonable threshold.

The output of a fuzzy matching system is not a yes or no determination. It is a similarity score. The compliance team must then decide, for each alert, whether the score indicates a genuine match requiring further investigation or a coincidental string similarity that can be cleared. That decision, made under time pressure on a large volume of daily alerts, is where the systemic reliability of the screening programme is actually tested.

Semantic and embedding-based matching

Newer screening systems use vector embeddings of name strings, trained on large corpora of name data, to compute semantic similarity between customer names and list entries. These systems can capture relationships that character-level fuzzy matching misses, including cross-language transliteration equivalences and culturally specific name variations.

Semantic matching can reduce false positive rates relative to character-level fuzzy matching for certain name populations, particularly where transliteration variation is the dominant source of false positives. But the compliance challenge with embedding-based matching is auditability. When a semantic matching system generates an alert, the reason for the alert is a vector similarity score derived from an opaque model. The compliance officer reviewing the alert cannot examine the algorithm's reasoning in the way they can examine a Levenshtein distance calculation. The decision to clear the alert then rests on the reviewer's own judgment about whether the names look similar, rather than on a documented algorithmic rationale.

This auditability gap is not a theoretical concern. Supervisory examinations of screening programmes increasingly focus on the firm's ability to explain its alert triage decisions. A system that generates alerts with clear, documented algorithmic rationale produces a more defensible record than one that generates alerts based on opaque similarity scores, even if the second system's false positive rate is lower.

The evidentiary burden: what a defensible screening record requires

The evidentiary standard for a screening decision is not simply that the decision was made. It is that the decision can be explained and defended to a supervisor reviewing it, potentially months or years later.

A defensible screening record for a cleared alert requires:

The match details as at the time of screening. The customer data that was screened, the list entry against which the alert was generated, and the similarity score or match rationale produced by the algorithm. This record cannot be reconstructed after the fact; it must be captured at the point of screening and retained.

The version of the list screened against. A screening result is only meaningful in the context of the list version it was generated from. If the EU consolidated sanctions list was updated between the transaction and the review of the alert, the list in force at the time of screening is the legally relevant version. Screening systems that do not retain list version information make it impossible to reconstruct what the system was screening against at any given point.

The triage decision and its rationale. Who reviewed the alert, what they concluded, and why. The documentation standard for cleared alerts should distinguish between alerts cleared because the customer name does not match the listed name on any reasonable analysis, alerts cleared because the customer has been through enhanced verification and assessed as not the listed person, and alerts cleared because the algorithm generated a match on a common name with insufficient discriminating characteristics to support a conclusion either way. These are different evidentiary situations and should be documented differently.

The escalation record where applicable. Where an alert was escalated to a senior compliance officer or to an external legal assessment, the escalation record, the response received, and the ultimate decision should be retained as part of the alert file.

High-risk country list updates: the propagation window

The most systematically neglected element of AML screening programme maintenance is the update propagation process for high-risk country classifications. This is distinct from sanctions list updates, though operationally similar in structure.

Sanctions list updates, particularly to the EU's consolidated financial sanctions list, occur without a fixed schedule and in response to EU Council decisions adopting or amending sanctions programmes. Major sanctions packages involving Russia, Iran, Belarus, Myanmar, and other jurisdictions have been adopted at intervals that reflect political and diplomatic developments. Each update generates a new version of the consolidated list, and that version must be loaded into the screening system before the regulation adopting it enters into force. Sanctions regulations adopted by the Council typically enter into force on the day of publication in the Official Journal. A screening system that loads list updates on a daily or weekly batch schedule has a window during which transactions are screened against a list that does not reflect the current legal position.

For the AML high-risk third country list, the update mechanism is set out in Articles 29 and 30 of the AML Regulation (Regulation (EU) 2024/1624). When the Commission identifies a third country with significant strategic deficiencies or compliance weaknesses in its AML/CFT regime, it is required to adopt a delegated act within 20 calendar days of making that determination. Once adopted, the delegated act is subject to the European Parliament and Council's objection period, one month extendable by one month, before it enters into force. This means the compliance window between the Commission's initial determination and the point at which the delegated act becomes binding can span several months, but the firm's obligation to monitor for and respond to those publications runs from the moment they are adopted and notified, not from a later date. The operationally relevant trigger is adoption and publication, not the end of the objection period.

This creates a different, and in some respects more demanding, monitoring requirement than sanctions screening. The firm must track Commission delegated acts in this area from the point of publication, load the updated country classification before the act takes effect, and ensure the EDD trigger logic reflects the new classification from that date.

The gap between list publication and system update is where operationally mature screening programmes distinguish themselves from operationally adequate ones. A screening programme designed for rapid update propagation has:

A monitoring function that detects Official Journal publications of sanctions regulations and AML high-risk country delegated regulations as soon as they are published, rather than relying on a vendor's scheduled update cycle.

A loading process that can ingest a new list version and make it active in the screening system within hours of detection, not within days.

A testing step that confirms the new list version is correctly loaded before it is made active, including spot-checks on known additions and removals.

A documented record of when each list version was loaded and activated, creating an auditable log of the system's list currency at any point in time.

Vendors who supply sanctions list data as part of a screening service should be assessed against these criteria. A vendor that delivers list updates on a once-daily schedule, regardless of the publication timing of new sanctions regulations, is structurally unable to support a screening programme that maintains list currency to the standard the legal obligations require.

The PEP screening specificity problem and the AML Regulation resolution

PEP screening has a different false positive profile from sanctions screening. Sanctions lists contain a defined population of designated persons, and a match against a sanctions list means the customer is potentially that specific person. PEP lists define categories of positions, not specific individuals, and matching against a PEP list means the customer's name matches a name in a database of people who hold or have held positions that generate PEP status.

The false positive problem in PEP screening is driven by two factors: the breadth of PEP-generating positions and the frequency of common names among the population of people who hold those positions.

Under the directive framework, each member state defined the scope of PEP-generating functions independently, and the PEP databases used by commercial screening vendors aggregated those national definitions in ways that varied in coverage and recency. A customer named "Maria Santos" flagged as a PEP match might be a former member of a Portuguese municipal council, a Mozambican parliamentary deputy, a Brazilian state secretary, or an entirely different person who happens to share a name with one of the above. Each of those scenarios requires different handling, but the alert itself provides no information about which scenario applies.

The AML Regulation addresses the PEP definition divergence through a structured list mechanism set out in Article 43. Each member state is required to issue and maintain a list of the exact functions that qualify as prominent public functions under its national law, notify that list to the Commission, and keep it updated. The Commission assembles these national lists, together with a list of prominent public functions at EU institutional level, into a single consolidated list published in the Official Journal. AMLA makes that list publicly available on its website and issues guidelines on how to assess the risk level associated with different categories of PEP (Article 42(2)). The practical consequence for screening programmes is that commercial PEP database vendors will need to align their EU coverage to this consolidated list, and firms will need to verify that their vendor's update process tracks the list as it is amended rather than relying on proprietary definitions.

This mechanism does not solve the common name false positive problem. A uniform definition of which positions generate PEP status still produces alerts when a customer name matches a PEP database entry for a different person with the same name. What it does is reduce the over-broad and inconsistent scoping of PEP-generating positions that was a systematic feature of directive-based national implementations, which should reduce the population of PEP database entries for positions that were included in some national implementations but will not fall within the consolidated EU-level list.

Designing for auditability: the architecture requirements

An auditable screening data pipeline has structural features that distinguish it from a screening configuration optimised primarily for alert volume reduction.

Immutable alert records. Each alert generated by the screening system must be written to an immutable record at the point of generation, capturing the input data, the list version, the match details, and the algorithm parameters. This record cannot be modified by the triage process; the triage outcome is appended as a separate record. The immutability of the alert record is what makes the evidentiary burden supportable: when a supervisor asks to see the decision record for a specific cleared alert from eight months ago, the record of what the system generated and what the reviewer decided must be exactly what it was at the time.

List version control with activation timestamps. Every list version loaded into the screening system must be recorded with a unique identifier, a publication date from the source, and an activation timestamp in the screening system. The screening system must be able to report, for any transaction screened at any point, which list version was active at the time of screening. Without this capability, the system cannot support the reconstruction of the screening environment at a given historical point, which is the foundation of any supervisory investigation into a missed match.

Algorithm parameter versioning. The matching algorithm parameters, including similarity thresholds, the set of algorithms in use, and any entity-type-specific parameter variations, must be versioned alongside list data. A change in matching parameters changes the system's alert behaviour, and that change must be traceable so that historical alert rates can be explained in the context of the parameters in effect at the time.

Triage decision documentation with reviewer identity. Alert triage decisions must be attributed to a specific reviewer, timestamped, and documented with a substantive rationale rather than a category code. The rationale does not need to be lengthy, but it must be sufficient to support a subsequent reviewer, including a supervisory examiner, in understanding why the alert was cleared or escalated. A triage record that reads "cleared: no match" is not defensible if the alert was for a name similarity score of 0.87 with a common name population. A triage record that reads "cleared: customer name is common in origin jurisdiction; no other discriminating characteristics match listed person; date of birth, nationality, and account activity inconsistent with listed person's profile" is.

What the AMLA supervisory framework changes for screening programmes

The AML Regulation and the AMLA supervisory architecture are raising the evidentiary standard for screening programmes in two specific ways that compliance teams should factor into their design decisions.

The central AML/CFT database that AMLA maintains under Article 11 of Regulation (EU) 2024/1620, to which national NCAs contribute supervisory findings, means that screening programme deficiencies identified in one member state supervisory examination are increasingly visible to NCAs across the bloc. A screening programme found inadequate in a German BaFin examination is not invisible to the Central Bank of Ireland or the AFM when those supervisors assess similar firms. The information asymmetry that allowed substandard screening programmes to persist in some jurisdictions without cross-border visibility is structurally narrowing.

AMLA's own direct supervisory mandate, which commences with the first selection of significant obliged entities in 2028, is still developing its methodology, but the mandate documents and consultation papers published to date indicate that screening programme auditability, specifically the ability to demonstrate the matching logic, the list currency at any point in time, and the decision record for alert triage, is a primary supervisory focus rather than a secondary one.

For firms outside AMLA's direct supervision population, the relevant signal is that national NCAs are being assessed by AMLA's peer review function on the rigour of their own supervisory activity. NCAs that do not examine screening programme auditability in depth are more exposed to AMLA peer review criticism, which creates pressure on those NCAs to raise their examination standards. The direction of travel is toward screening programme examinations that require a firm to walk through the alert triage record, explain the matching algorithm parameters, and demonstrate the list update log, not merely assert that a screening programme exists.

Forseti monitors EU sanctions list updates, AML high-risk third country delegated regulations, and AMLA supervisory developments as they are published, anchored to verified official sources. Start for free.

For the AML Regulation's customer due diligence data schema requirements that feed into the screening identity data, see structuring KYC profiles for the AMLA single rulebook: Level 2 RTS requirements. For the high-risk country list update process and its maintenance burden, see EU AML high-risk third countries: the maintenance burden compliance teams underestimate. For how the AML Regulation changes the substantive compliance baseline, see the EU AML single rulebook: what uniform enforcement means for fintechs.

📋 Track EU financial regulation continuously

Forseti monitors EU financial regulation and delivers personalised alerts anchored to verified official sources.

14-day free trial. No credit card required.