Speakers

Rajeev Dave
Managing Director Surveillance, Natixis

Michael Cowell
Senior Manager Risk, Regulatory, and Forensics, Deloitte

Eren Erman
Global Head of Compliance Technology, TP ICAP

Alex de Lucena
Director of Product Strategy, Shield
Voice communications are expanding rapidly across trading floors, collaboration platforms, and mobile calls—bringing new regulatory scrutiny and operational complexity. Yet many institutions still rely on random sampling, brittle transcription, or siloed tools that fragment oversight and leave critical risks undetected.
Join industry experts as they discuss how leading institutions are rethinking voice monitoring for stronger compliance.
You’ll learn:
Hear real-world insights on closing monitoring gaps and transforming voice from a compliance burden into a trusted source of risk intelligence.

Rajeev Dave
Managing Director Surveillance, Natixis

Michael Cowell
Senior Manager Risk, Regulatory, and Forensics, Deloitte

Eren Erman
Global Head of Compliance Technology, TP ICAP

Alex de Lucena
Director of Product Strategy, Shield
Shield Insiders In-Focus
Speakers
Yael:Hi, everyone, welcome to our webinar on voice monitoring. Thank you for joining, and welcome to Rethinking Voice Monitoring. My name is Yael, and I’m the Director of Product Marketing at Shield. Today we hope to provide an insightful session looking at how leading institutions are rethinking their voice monitoring. Our webinar will cover a brief look at the regulatory landscape and how it differs across jurisdictions; the key pillars for building an effective voice program; how AI is influencing the landscape; and finally, how to strategically leverage voice in your compliance approach.
Some quick housekeeping: the webinar is being recorded and will be shared to view on demand afterward. All attendees are hidden and muted, but you can ask questions at any point by typing them into the chat, and we’re available throughout to answer. We’ll try to leave time at the end for live questions; if we don’t get to yours, we’ll follow up by email. We’re not sharing any promotional content, and all speakers are providing opinions personal to them. Without further ado, I’ll hand over to Alex, our moderator for today and Shield’s Director of Surveillance and Governance Strategy, to introduce our panelists.
Alex:Thanks, Yael, and hello, everybody. Really happy to be here — I love these discussions. I find voice especially interesting because every time we have this discussion the details change a bit; it’s one of the areas where we continue to see advancement. Before we jump into the agenda, I always like people to introduce themselves. Raj, Mike, Eren, in that order.
Rajeev:Thanks, Alex. My name is Rajeev Dave. I’m the Global Head of Surveillance with Natixis, a French investment bank, part of the BPCE Group. I’ve been in the surveillance space since approximately 2008 — I started in bank holding company Reg W surveillance and moved into eComms, voice, trade, and insider dealing surveillance in 2010, and have been in the space ever since.
Mike:Hi, everyone. I’m Michael Cowell, a Senior Manager with Deloitte Financial Advisory Services. My core at Deloitte was historically digital forensics and eDiscovery, specifically focused on mobile device and cell phone forensics. Around 2018, I started to see how that overlapped with e-communications capture and monitoring from mobile devices, as phones were being used more for general business. Since then, I’ve worked with clients across the board, FSI and beyond, to help them develop compliance programs, select tools, and work through implementation.
Eren:Hi, I’m Eren Erman. I’m American, though you can’t tell from my accent — I’ve been living in London for seventeen years. My background: I spent six years at PwC in London building our surveillance and regulation practice, in energy and commodities and financial services. I then moved to TP ICAP, where I led the compliance technology function for six years. I don’t represent TP ICAP — everything I say today is my own personal view. Currently, I’m the CEO and founder of RegTrail, a reg intelligence platform focused on horizon scanning for the energy and commodity trading space.
Alex:Amazing. And I’m Alex de Lucena, Director of Surveillance and Governance Strategy here at Shield. A lot of these voice discussions cover a well-worn path — we always start with the journey from nothing to random samples to transcription, we parse word error rates, and there’s this “making the best of it” vibe that always seems to lag eComms but is better than nothing. I think that misses something crucial about the moment we’re in: today’s GenAI transcription tools show significantly lower error rates, they’re multilingual, and they can be optimized for finance. A lot of the gaps we’ve long pointed to across surveillance have been solved by the technology we’re seeing today, yet we still see different rates of adoption and certain elements holding firms back. So today I want to focus on the technology, the operating models, and the risk strategy — and understand how we pull all these elements together.
Alex:Mike, you’re in the business of advising people. Let’s start with a snapshot of today. We understand the regulations that have driven voice to date, but it’s diffuse, with divergence top of mind across global regulators. What’s guiding capture and, by extension, monitoring of voice across regulators today?
Mike:First, as a consultant, I need to say I’m not a lawyer and my statements aren’t to be construed as legal advice. With US regulation, the CFTC and Dodd-Frank have specific mandates for swap dealers and for trade reconstruction. But a lot of other eComms regulation in the US is vague on purpose — regulators expect firms to make good risk-based decisions to meet the framework, without giving specifics at each step. What we often see is large multinational firms capturing and surveilling their voice, especially if they have a swap dealer, because the technology is already in place. But once you move away from those organizations — to more US-focused or smaller firms getting their feet wet — because it’s not detailed out, voice often gets captured but not necessarily monitored. As GenAI makes this space easier and regulators get that data more often, it’s my belief they’ll then expect that data in all circumstances, flushing out the vagueness. We’re not at that point yet, but I see it close on the horizon.
Alex:In your personal view, is there wiggle room around the interpretability of applying monitoring to voice? Or is it more that we’re following organic adoption, where firms just happen to be at different points on the same journey?
Mike:There is an interpretation factor. The other thing US regulators lean on is that once a consent order is issued to one firm, they expect all firms to read it and apply it. So as more voice comes up in those consent orders, the gray area will become no longer gray, because those orders drive future adoption.
Alex:Eren, over to you. What does the regulatory landscape look like on your side of the pond?
Eren:I’ll keep it short. Primarily it’s the Market Abuse Regulation, it’s REMIT for wholesale energy products, and it’s MiFID. MiFID and REMIT require voice recording for certain derivative transactions; MAR is your broader risk-assessment landscape. Voice recording under MiFID and REMIT — people are doing that. But to build on Mike’s point, the transcription and surveillance is not prescribed in regulation. That said, now that technology can transcribe voice, third-party vendors are engaging with regulators and showing them their kit — the regulators are actually using it for their own surveillance, so their expectations are indirectly shifting. Everybody has their own risk assessment; they have to decide whether voice is a risk. If you’re transacting using voice, that would fall under it, and your defense for not having transcription or surveillance would have to be compelling if you’re a regulated firm.
Alex:What would an argument for not having one look like?
Eren:You’d say voice isn’t used to transact — you’re only using it for market color, for example. Unregulated firms may say they don’t use voice to transact; it’s policed out. So while you might use voice for certain components of the pre-trade conversation, you wouldn’t use it to actually transact, and that’s where you draw the line. I’d also go a step further: a lot of investment firms are already doing voice proactively. The exchanges are involved too — in some enforcement decisions they request not just trade information but the communications, including voice and eComms. So firms are having to produce that, and the expectation is shifting to “you’re recording voice and you have the transcription — we want to see that too.”
Alex:One thing that’s safe to say is the difference between Europe — especially the UK — and the US: under MiFID you tend to see groups of people being voice-recorded, whereas under Dodd-Frank you tend to see individuals recorded. That’s one of the complexities of navigating these regulatory differences.
Alex:Raj, we generally want our programs to be flat from an operational perspective — easier to maintain. How do you balance these different regulatory obligations that say “record this person but not that one,” giving you whole populations to record and monitor while leaving others out?
Rajeev:Great question. I should mention I recently relocated to Paris — I’ve been here a year, after New York — so I’m trying to find that middle ground between the regulatory regimes in Europe and the US. As someone building a global program, I completely agree we want something as consistently applied as possible across jurisdictions. But I’m learning quickly that’s not always possible and is outside your control. In Europe, privacy is extremely important. In France, there are strong union rules prohibiting certain types of voice recording — you need permission from certain union groups, and there are opt-in clauses in employment agreements. So I’ve been spending a lot more time with our privacy office, technology partners, and legal teams.
There’s also DORA, which has prescriptive rules on third-party outsourcing far more intense than the US. A simple example: the right-to-audit clause. In the US, internal audit generally relies on a third-party Big Four firm to review a vendor. Under DORA, in my interpretation, there’s an obligation on internal audit groups to do their own reviews, not simply rely on third parties. So these are all things to consider when negotiating a contract or applying controls to a select group.
We try to harmonize and find opportunities to apply the intent of the rule in a way that makes sense — it’s risk-based. We spend a lot of time on risk assessments of teams and individuals based on their activities, on who really needs to be in scope. There are challenges with record keeping and retention — it’s a big cost to maintain records for seven, ten, thirty-five years, depending on the swap products. And it starts with the basics: every firm has a policy around global communications, usually combined for voice and eComms, and it starts with how you define a business communication. I’ve been at firms where a calendar invite is considered a business communication. And where does it start with voice? If you call your spouse or a personal contact, it’s under the recording privilege — do you need to keep and retain it, or can you destroy it? All of these questions have to be thought through.
Alex:I’ll stick a pin in one thing: with some of the GenAI tools we have now, there’s the ability to understand context more broadly and potentially suppress personal and private communications that have always been a speed bump — someone telling their partner they want pizza and they’ll hit the gym before home. But let’s move from taking in all these regulatory inputs to turning it into a program.
Alex:Raj, what are the most important pillars for setting up a voice program?
Rajeev:Most importantly, you need to understand your businesses and what regulatory framework you’re under — broker-dealer, swap dealer, security-based swap dealer, FCM, how you’re set up across jurisdictional entities. It starts with the regulatory regime, because you want to meet and manage your regulatory risk. Once you have a harmonized view of who the population should be for voice recording and surveillance, the next objective is to develop internal solutions or find technological ones. In our case, we have hundreds of people in the monitored community, and the data is in the terabytes at a minimum — so how do we get our hands around that volume?
I always use the chain analogy: surveillance is at the end of the chain. If the records aren’t captured adequately to begin with — if we don’t have all the corporate devices recorded and maintained, mobile carriers providing the requisite information, all the channels people speak on, Teams ingested into our vault system — we’re only as good as the data available upstream. So it all falls into a chain of governance. A lot of people fail because they think, “Let’s just put a new tool in to solve the problem at the end of the chain,” but they fail to realize they have gaps upstream — and no matter what tool you put in, you’ll still have those gaps, the false positives, the data-quality issues. And I’d emphasize: I have never been part of an exam where it was not expected that you had a voice surveillance program. If you’re recording and storing data, there’s an expectation you’re monitoring it. I don’t think you even have a defense.
Alex:We’ll have another very short webinar where you come back and tell us how that defense went. Mike, I see you nodding — can you fill out the structure?
Mike:Raj touched on understanding your channels and where the data actually comes from. Whenever I’m given a soapbox, I love to get on it on this topic, because it’s the hardest piece for clients to fully understand — especially with a cell phone, there isn’t one magic solution to capture every channel. People think of a phone as a single endpoint: put something on it and you capture everything. But you can call me on my phone directly, call my Teams number routed to my phone, or call me on Teams, which doesn’t even go over the phone. I tell people to think about a phone as a box of channels, and you have to solve for every single app individually. The pitfall is thinking, “We bought this great product, installed it, we’re getting text messages and native phone calls” — and then an investigation comes around and you find everyone’s been calling on WhatsApp because it’s easier.
Alex:Or you’re not capturing voice notes, or you’re capturing them but not transcribing them.
Eren:I’m coming at it from a slightly different role, but if you’re just starting: voice is a cross-functional transformation in an organization. Creating a steering committee with executive sponsorship is critical. The people involved would be compliance, IT, voice infrastructure, and HR. I mention HR because when traders or brokers operate across jurisdictions or are multi-hatting in different regulatory regimes, that HR data — from your enterprise HR system — is fundamental to defining the lineage of who’s monitored and what gets turned on. The recording capabilities and the metadata behind them are fundamental too. You can’t just assume you can turn on the switch and get really good data — you have to map out each channel, the data you need, make sure it’s recording correctly, and build controls around it, even before you find a third party or build in-house for surveillance or transcription. You have to have everything lined up first to then ingest into whatever surveillance system.
Rajeev:Can I add real quick? From my experience — this is maybe the third or fourth bank I’ve worked at in surveillance — I have never found an institution that could tell me who the accountable executive is for something like voice. It’s a shared ownership, to Eren’s point, and what happens is people assume IT is the owner, but the reality is this is business data, business ownership. Most times people rely on IT and assume IT understands the business requirements to develop a holistic program — and they just don’t, without your participation, your interpretation of the rule and expectations. So Eren’s point is spot on that it’s cross-collaboration, but at the end of the day you need to find someone who’s the accountable executive, and it’s extremely difficult to find one person who will say, “I own this program end to end.”
Alex:That’s a great point. When I was in a surveillance role, the voice team managing desk phones was an old-school team that understood analog technology and all its kinks — they had nothing to do with mobile capture. Once the SEC and CFTC fines came out around WhatsApp and SMS capture, that was a whole other team spun up just to meet that need, and then your compliance-tech teams scaffold across it. Another thing that comes up is recording quality — a little more moot with GenAI transcription, but with legacy platforms using older transcription methods, if the recording is compressed at too high a rate, it’s hard to get a good transcription. The reason compression rates are high is that it costs less, and you assume you’re just storing the communications — it isn’t until you have to transcribe them that you hit the wall that the audio quality is too bad. So who approves that cost? The tech guys overseeing the landlines want nothing to do with it; their day-to-day is fixed. Now it has to go up to that nebulous executive sponsor.
Alex:Eren, one of your good points earlier was that a lot of education comes with voice — an opportunity to bring stakeholders together. What does that education piece look like?
Eren:First, getting a current-state view of what’s happening in the organization and being able to articulate the exposure and risk. You’d be surprised — when you lift the hood, you go into a corner and there’s a desk doing something you just find out about. So that roadshow of finding out the inventory takes time, and that in itself is education. The second point is helping people understand what voice monitoring is as a capability, why it’s important, and the business case to invest — that’s more at the exec level than with compliance or IT, who already know it’s a risk. And the third is helping them understand the value it brings beyond compliance — there’s a business-partnering element, a conduct-and-culture element, front-office opportunities with analytics or other use cases. That all culminates in a business case showing it’s not just “click a button, start transcribing,” but that a lot of opportunities come with it.
Alex:And we often teach people a lot about something we’re all familiar with — eComms is so established, but with voice, people have consumer experience to reference. Even my mother watches a message I leave her get transcribed. So there’s an education piece explaining the tech, but the consumer experience lets you close the gap if you make it relatable.
Alex:Mike, let’s transition into what surveillance across voice looks like in production. Traditional elements have been random samples, transcription, lexicons. What makes a good program?
Mike:On a good program, I want to echo the cross-functional team point. The most successful firms understand you can’t just have compliance and legal lobbing things over the fence to IT and expect IT to magically figure it out. It’s having a team where business representatives discuss what they need, IT opines on what’s technically possible, and legal, HR, and compliance are in the room to talk about what’s required, so they find a happy medium. If it’s just thrown to IT and IT doesn’t understand the rationale, you end up in odd situations — or everything is captured well, but you have thirty different tools each doing a micro piece, with data-format and compression issues, and you’re spending a lot of money on many tools when there might be a way to tie it together.
Alex:How executive should this committee be?
Mike:Relatively high up — heads of business or their representative, and someone in legal senior enough to make decisions on the spot about interpreting the regulation, not someone who has to go determine and come back. And with IT, this is expensive and balloons in cost quickly as you bring on more channels and meet clients where they are — WhatsApp, WeChat in China. Getting that knowledge and spend together requires a pretty high seat. You don’t necessarily need C-suite on the panel, but you’re getting close to the decision-makers, because as quickly as technology moves, you need these decisions and the spend actioned rapidly.
Alex:Raj, is there a role or vertical within financial institutions that would be the ideal sponsor?
Rajeev:My experience is our legal team, but also the CISO, the chief data officer — someone at that level who understands the requirements and challenges is an important ally. And I think the world is changing. We were just having this conversation around “show me the return on my investment.” That’s not something in compliance I’ve ever had to answer — for the last fifteen years, you’d pull up the Wall Street Journal, show an article, say “this could be us,” and spend money. The world is shifting. Regulatory changes are coming; it’s becoming somewhat less regulated in certain jurisdictions. So people are challenging the investments we make on compliance tools much more, forcing us to think about the return: what are the other use cases, can I converge different groups doing similar functions onto one tool, can you show improvements with AI on headcount, can you flatline the team rather than keep hiring to review false positives? The old adage of “it’s in the Wall Street Journal” doesn’t fly anymore. Historically, we’ve also been measured against the number of SAR — suspicious activity report — filings, and those always come from the trade side, not really from eComms or voice programs. So people are looking for that metric as a return: show me you’re finding the smoking gun, something a simpler lexicon-based solution couldn’t have found.
Alex:I’d turn that on its head slightly. What I’m hearing is less “can we put in a lexicon to solve this?” and more “if we’re going to do this, can AI do it in a way that brings the efficiencies we want and keeps costs flat?” That’s the change I’ve seen — and it’s pretty much jurisdiction-agnostic, even though the US is diverging toward a looser, more business-friendly approach, Europe is hardening, and the UK is somewhere in the middle. The message across each is: if we’re going to do this and there’s AI, will it make us more efficient and keep costs flat?
Rajeev:I agree, and it’s FOMO, Alex. For the last five years people read about AI at the periphery, knew it could help in certain areas — it was a fear of missing out. Now, especially with tier-one banks, it’s less regulatory and more “show me my investment will yield value.” I work at more of a tier-two investment bank, which is interesting, because tier-one banks operate in almost every jurisdiction — at one bank I counted three hundred regulators we were under. Just saying “you won’t be in the newspaper” or “you’ll avoid a fine” isn’t as resonant anymore.
Alex:Let’s fast-forward to the AI capabilities. Mike, can you talk about the delta between where programs are and where the AI is, and how we bridge that gap?
Mike:If we go back one step to machine-learning models and natural language processing, you often had to pick a language, sometimes a specialty, and do training yourself or pick a vendor who trained their model in the language you believe people will speak, on the topics that make sense. As soon as your representative on the phone stepped out of that box, your transcription could rapidly go downhill. Change language — maybe you have a call in both English and Spanish, but the call was set to English at the start, so the Spanish model might not pick it up.
Alex:Or you’re transcribing it twice and you have two calls.
Mike:Exactly — just to generate those two transcriptions. And that’s before someone with a thick accent, or someone in Scotland talking to someone in Japan. I hate to pick on Switzerland, but it’s the easiest example: the dialects and specific languages spoken in Switzerland are really only spoken there, so there wasn’t a huge market of vendors focusing on those languages, and voice recording isn’t a high priority there. GenAI takes us away from that — the foundational models are trained on so much, they’ve seen every language by scraping the internet, so that concern gets sidelined; you can switch languages.
Alex:Underline that: GenAI capabilities are inherently multilingual.
Mike:Multilingual, and they know all the acronyms, whereas traditionally you’d have to train on your specific product acronyms. You can switch language five or six times in a single conversation and get a transcript all in English — it translates on the fly. The other big piece, switching into my eDiscovery role, is that previous iterations of AI didn’t have the ability to summarize the transcription. There’s a famous example of traders talking about Chinese food prices, which doesn’t make sense because they’re hiding that they’re talking about stock prices. Your transcription might say they’re discussing the price of Chinese food, but the summary might flag, “Why was that discussed at eight in the morning, and why was one meal three hundred dollars? That doesn’t make sense — they’re talking about something else.” The summary captures what the literal transcription does not. That’s a direct translation to voice-call monitoring: not just understanding word for word, but why it’s being said and what it means in the larger context.
Alex:You could take summarization a step further — there are instruction-based LLMs where you can bake in a lot to make it understand calls through the lens of a compliance officer, to identify the particular risk, because with a long call a summarization can miss the risk. So it can be configured for the purview specific to certain risks, whether someone’s in compliance, surveillance, or HR.
Alex:Eren, anything you’d factor in around the performance of the AI?
Eren:I’ll take it back to the grassroots of onboarding a voice transcription vendor. Everyone assumes AI and GenAI will deliver almost perfect transcription. While accuracy rates are high, especially in mainstream languages, you can’t assume that. My takeaway: bake time into your voice transcription to test the provider for accuracy. “Ground truth” takes time — that’s the term, where you record and transcribe to 100% accuracy yourself, then send that exact voice note to the vendor, they transcribe it, and you mark their homework. Depending on what you trade, how you trade it, and how it’s spoken, the transcription might not be accurate, and tuning is required. Make sure you have the right vendor that will work with you on tuning your models. If they aren’t accurate, your whole operating model of getting efficiencies with AI goes down the rabbit hole — you increase risk because you have more alerts to review, and it’s discoverable in court. So don’t go live until you’ve done the homework.
Alex:You have to do the testing, and there’s always some tuning. One other thing to insist on — this is me in the vendor seat — is that vendors should have their own benchmarking for their transcription against both ground truth and other transcription providers. A lot of times institutions ask, “Why can’t we just use AWS?” Putting aside the advantages of having everything in one platform, eComms next to voice, what’s in your vendor sauce that’s specific to financial institutions? Those are important questions to ask any voice vendor.
Mike:One thing, Alex — sorry to jump in. Right now we have about six mainstream foundational models, plus anyone can build their own LLM on top. So GenAI is not necessarily always equal to GenAI. You don’t just want to see what it does on the front end — understand which back end is being used. Is it a mainstream model or proprietary? Going back to my eDiscovery hat, one of our GenAI tools runs four different foundational models in the background, because one is really good at images, one at audio, one at interpreting text. Every foundational model has a pro and a con — it does one thing really well and something else really badly. I often see people advertising a model for voice and think, “I don’t know if that’s the model I’d use.” So having that understanding, on top of “is it accurate,” helps you drive what happens in the future.
Alex:And I’d add revisiting those decisions down the line — we see customers with controls in place that ask their vendors, essentially, “Is your solution still good?” It falls on us to then validate that.
Alex:Raj, one interesting perspective you raised was reframing voice as business intelligence — seeing it as more than a surveillance obligation, as an asset.
Rajeev:Happy to. If you think of voice as a strategic asset, there are multiple use cases. The way I think about it: what other groups in my institution rely on voice to perform their day-to-day? We have large retail banking operations doing quality-of-service reviews and metrics on customer interactions, payment centers, customer contact. We have a large legal team looking at records and transcriptions — historically we had to call a third-party law firm because we couldn’t rely on our own internal systems for transcription. We have teams listening to voice records looking for voice trading activity. So there are a whole bunch of use cases where the generative AI technology could be deployed. It’s imperative for us as surveillance officers — this goes back to the ROI discussion — to explore those use cases and see whether other stakeholders could benefit, to get buy-in and the funding you need. The more the merrier: get more teams leveraged onto single sources to pull data from. There’s a great opportunity for financial services firms to use this data beyond compliance and surveillance — even voice biometrics, voiceprints. We have such a hard time identifying who’s speaking to whom, what phone number this is.
Alex:“Voice is my password,” and then identifying.
Alex:It sounds like we could talk for another hour — it feels like we just got started. A through line here is the community we bring together to make the case for change, and Raj’s point that there are opportunities beyond just surveillance that can pad that case and put firms on a future footing as the tech across voice keeps advancing. We’re out of time. We’ve been answering questions as they came into the chat, and we’ll email anyone whose question we didn’t get to. I want to thank Rajeev Dave, Michael Cowell, and Eren Erman — this has been a really great discussion. Thank you, everyone, for joining Rethinking Voice Monitoring. Have a great day.
Rajeev:Thank you. Thanks, everybody.
Eren:Thanks, Alex, for hosting. Goodbye.
Mike:Thanks, everyone.
It depends on jurisdiction, but the expectation is shifting toward monitoring. In the US, CFTC and Dodd-Frank rules are specific for swap dealers and trade reconstruction, while much else is deliberately vague — so many firms capture voice but don’t yet surveil it. In the EU and UK, MiFID and REMIT require recording for certain derivative transactions. The panel’s blunt take: if you’re recording and storing voice, there’s now an expectation you’re monitoring it.
Under Dodd-Frank, the US tends to record individuals and is prescriptive mainly for swap dealers, leaving much else to risk-based interpretation. In the UK and EU, MiFID, REMIT, and the Market Abuse Regulation tend to require whole groups to be recorded. Europe also layers on privacy: French union rules can prohibit recording certain people, and DORA imposes stricter third-party audit obligations than the US.
Aim for a flat, harmonized program — it’s easier to maintain — but accept that some differences are outside your control. Privacy law, French union consent rules, opt-in employment clauses, DORA’s audit requirements, and rules on who you can and can’t record all force local variation. The practical path is risk-based scoping: assess teams and individuals by activity to decide who truly needs to be in scope.
Because surveillance sits at the end of the chain. If records aren’t captured properly upstream — corporate devices, mobile carriers, every channel, Teams into the vault — no downstream tool can fix it; you’ll still have gaps, false positives, and data-quality issues. As one panelist put it, a phone is a “box of channels”: solving for one app (say, native calls) while everyone quietly moves to WhatsApp leaves a blind spot.
This is notoriously hard — panelists said they’ve never found an institution that could name a single accountable executive for voice. It’s genuinely shared ownership across compliance, IT, voice infrastructure, and HR, ideally coordinated by a steering committee with executive sponsorship. The key warning: don’t assume IT owns it. This is business data, and IT can’t infer the regulatory requirements without compliance and the business at the table.
Substantially. Older machine-learning models forced you to pick a language and specialty, and accuracy collapsed the moment a call drifted outside that box. GenAI models are inherently multilingual, handle accents and industry acronyms, and can transcribe a call that switches languages several times into a single English transcript. They also summarize — surfacing hidden meaning, like traders discussing “Chinese food prices” at 8 a.m. that actually mask stock prices — and can be tuned to a compliance lens.
No — accuracy is high in mainstream languages, but you shouldn’t assume it. The panel’s advice is to build in time for “ground truth” testing: transcribe a sample to 100% accuracy yourself, then mark the vendor’s homework. Expect tuning for how your firm trades and speaks, ask vendors for benchmarking against ground truth and peers, and understand which foundational model sits underneath, since each has real strengths and weaknesses.
It can be a strategic asset, and reframing it that way strengthens the business case. The same captured voice and transcription can serve quality-of-service reviews in retail banking, legal record transcription, and even voice biometrics for identifying speakers. Converging those use cases onto a single source spreads the cost, wins internal buy-in, and helps answer the ROI questions compliance teams increasingly face.