Shield Glossary

Lexicon

A lexicon in communications surveillance is a structured set of words, phrases, linguistic patterns, and related rules that identifies communications that may indicate specific regulatory, conduct, or compliance risks. Lexicons map known risk behaviors to observable language so surveillance systems can select communications for review. They remain widely used because they are transparent and auditable. Most large financial institutions now supplement them with natural language processing, behavioral analytics, and AI that evaluate context and meaning instead of relying only on predefined keywords.

What Is a Lexicon in Communications Surveillance?

A surveillance lexicon is the rule set that decides which employee communications a compliance analyst reviews. Each entry links a known risk behavior, such as sharing inside information or coordinating on price, to the language an employee might use while doing it.

A mature lexicon is much more than a keyword list. It combines several types of rules:

ComponentWhat it doesExample
TermsSingle words associated with a risk“guaranteed,” “insider”
PhrasesMulti-word expressions“keep this between us”
Proximity rulesFlag terms that appear near each other“price” within five words of “move”
Wildcards and stemmingCapture variations of a root word“manipulat*” matches manipulate, manipulated, manipulation
ExclusionsSuppress known benign matchesDisclaimers, newsletters, email signatures
Risk mappingTie each rule to a risk scenarioInsider dealing, collusion, conduct risk

Lexicons run across every channel an institution captures: email, Bloomberg chat, Microsoft Teams, WhatsApp, SMS, and transcribed voice calls. When a message matches a rule, the surveillance system generates an alert for an analyst to review.

What Risks Do Surveillance Lexicons Detect?

Surveillance lexicons are usually organized by risk category, and each category ties back to a specific regulatory obligation. The table below shows common categories and where they come from.

Risk CategoryLanguage a Lexicon targetsRegulatory Context
Insider dealing and unlawful disclosureReferences to non-public information, deal code names, requests for secrecyMAR (EU and UK)
Market manipulationCoordinated pricing, holding or pushing a level, benchmark referencesMAR; Dodd-Frank anti-manipulation provisions
Information barrier breachesCross-desk references to restricted clients or transactionsMAR; FCA SYSC 10
Off-channel communicationRequests to move a conversation to personal WhatsApp, Signal or a mobile phoneSEC Rule 17a-4; MiFID II record-keeping
Non-financial misconductHarassment, bullying, or discriminatory languageFCA conduct rules, including CP25/18
Mis-selling and customer harmPromises of returns or guarantees to clientsFINRA Rule 2210

Institutions typically start from a vendor-supplied lexicon and then add terms specific to their own business lines, products, jurisdictions, and internal deal names.

Why Do Financial Institutions Still Use Lexicons?

Financial institutions still use lexicons because every lexicon alert can be traced to a specific, documented rule. When a regulator or internal auditor asks why a message was flagged, the answer is concrete: this phrase matched this rule, which maps to this risk.

That traceability matters under supervisory regimes such as FINRA Rule 3110, which requires firms to establish, document, and maintain procedures for reviewing correspondence. Lexicon logic is easy to document, test, and evidence.

Lexicons also give compliance teams direct control. When a new risk emerges, such as a new deal code name, a new product, or a new piece of desk slang, a surveillance team can add a rule the same day without retraining a model. For known risk language, lexicons provide precise, predictable coverage.

What Are the Limitations of Lexicon-Based Surveillance?

The core limitation is that lexicons match words, not meaning. A lexicon cannot tell the difference between “fixing a deal” and “fixing up an old car.” Both match. Only one is a risk.

That context gap creates problems that grow with every new channel:

  • Alert volume. Broad rules generate large numbers of false positives, and analysts spend their time clearing noise instead of investigating genuine risk. Teams end up reviewing August’s alerts in October.
  • Evasion. Employees who know what the lexicon looks for change their language by using misspellings, emojis, abbreviations, and code words.
  • Channel fit. Short-form messages on Teams and WhatsApp carry far less context than email. Voice transcripts introduce transcription errors that exact-match rules miss.
  • Language coverage. Multilingual institutions need lexicons in every language their employees use, and direct translation rarely captures local idiom.
  • Maintenance burden. Every rule added to catch a new risk adds noise. Every rule removed to reduce noise risks opening a coverage gap.

The result is a double exposure. Real threats get buried in alert queues, and the institution struggles to demonstrate to regulators that its surveillance actually works.

How Do Lexicons Work Alongside AI and Behavioral Analytics?

Modern communications surveillance layers lexicons with contextual detection rather than replacing them. Each layer covers what the others miss:

  • Semantic analysis evaluates the meaning of a full message or conversation, not isolated words.
  • Behavioral analytics compares communication against an employee’s normal patterns, including who they talk to, when, and on which channel.
  • Lexicon detection provides precise, auditable coverage for known risk language and institution-specific terms.

Shield calls this combination multilayered AI coverage. The challenge is keeping the transparency that made lexicons defensible in the first place. AI models that compliance teams cannot explain to a regulator create a new kind of exposure.

Shield Surveillance runs custom lexicons alongside 100+ out-of-the-box risk scenarios covering market abuse, MAR, Dodd-Frank, and conduct risk. Alert Transparency makes every alert explainable by showing the scenario, the rule, and the relevancy score. 

How Should Compliance Teams Evaluate a Lexicon Program?

Compliance teams should evaluate a lexicon program on coverage, precision, and defensibility. These questions are a practical starting point:

  1. Is every rule mapped to a named risk scenario and regulation?
  2. When was each rule last tested for hit rate and escalation rate?
  3. Does the lexicon cover every captured channel, including voice transcripts, WhatsApp, and non-English communications?
  4. Are exclusions documented and reviewed so benign suppressions do not hide real risk?
  5. Can the team explain any alert to a regulator, including alerts generated by AI models?
  6. Does the lexicon include semantic and behavioral detection that evaluates context?

If several answers are “no,” the lexicon is likely generating noise while leaving gaps.

Related Terms

  • False positive
  • Alert rate
  • Behavioral analytics
  • Alert Transparency
  • eComms surveillance
  • Digital Communications Governance and Archiving (DCGA)