Speakers
Cheryl McKinnon
Forrester | Principal Analyst
Ofir Shabtai
Shield | CTO & Co-Founder
Aaron Gardner
Shield | eDiscovery Expert
The role of the archive is changing. As communication channels expand, data volumes grow, and teams adopt AI-driven monitoring and investigations, archiving—once treated primarily as a compliance storage requirement—is becoming a far more strategic asset: the trusted data foundation that powers modern compliance.
Yet many financial institutions remain constrained by environments designed for a different era, with archives built for retention rather than a unified source of truth, real-time access, or AI-enabled workflows.
Join guest speaker Cheryl McKinnon, Principal Analyst at Forrester Research, along with industry experts from Shield, for a practical discussion on:
Get real-world perspectives on how forward-looking institutions are transforming their archive into a robust foundation for modern communications governance.
Cheryl McKinnon
Forrester | Principal Analyst
Ofir Shabtai
Shield | CTO & Co-Founder
Aaron Gardner
Shield | eDiscovery Expert
Shield Insiders In-Focus
Speakers
Yael:Welcome to our webinar on Reimagining Archiving for AI-Driven Compliance, and thank you for joining us. My name is Yael, and I’m the Director of Product Marketing at Shield. Today we hope to provide an insightful session looking into why and how leading institutions are reassessing their archiving — away from just compliance storage and toward being a strategic business asset.
Let’s meet the panelists. We’re here with Cheryl McKinnon, Principal Analyst at Forrester. Cheryl has spent over a decade advising technology leaders on enterprise content management, content platforms, information archiving, and records and retention strategies; prior to Forrester, she led her own consulting firm across private and public sector organizations. Joining her is Ofir Shabtai, Shield’s co-founder and CTO, with over twenty-five years of experience in product and software development, focused on the intersection of AI and building large-scale, real-time systems. And last but not least, Aaron Gardner, one of Shield’s compliance and eDiscovery experts, who brings more than twenty years in eDiscovery, including serving as a lead advisor at a top-three global provider, leading litigation support at a top law firm, and founding his own boutique consultancy.
Cheryl will take us through the archiving market as it stands and what the latest shifts mean. Aaron will go through what’s needed to support investigations and stand up to regulatory scrutiny. Ofir will cover what modern archiving looks like from an architectural perspective. And we’ll wrap up with your questions. A quick housekeeping note: the webinar will be available on demand, all attendees are hidden and muted, you can ask questions live in the chat, and all discussions are personal opinions. Over to Cheryl.
Cheryl:Thank you, and a big thank you for the invitation. I’ve been covering the archiving market at Forrester Research for about thirteen years. Archiving is a mature market — compliance and storage management have been requirements for many years. We think of an archiving platform as one that helps you migrate digital information from one or many source systems into a repository and retain it for a specific period. But while it’s a mature market, the last one to two years have been among the most innovative we’ve seen in a long time. The main trend, as in many technology markets, is that AI has given fresh life and opportunity to archiving, providing new ways of working with the information organizations capture and preserve. The challenge continues to be the many data sources and the proliferation of communication types, so we need to keep pace on how we capture that content. The disruptor is that large vendors are engaging in M&A, so there may be market disruption on the horizon. But adoption is strong — our data tells us 80% of organizations have some form of archiving in place.
The key use cases haven’t changed tremendously, though we’re seeing emerging ones. The fundamental use case is overall information lifecycle management — particularly in regulated industries where you must capture particular content, retain it, and use it for investigations and legal discovery. There’s a range of related use cases: in financial services, supervision and surveillance of communication is key to compliance, and with fragmentation of content types, organizations may need to capture structured data or documents in addition to chats and email.
Cheryl:Some of the challenges I hear regularly, particularly in large regulated industries, are that conversations are now multimodal. We’re not just sending emails or chats — we’re using audio and video, we have transcripts and automatically generated call summaries, and governance can become quite complex. We see governance gaps across different meeting types, with different approaches for retaining email versus video or audio. Especially five or six years ago when we were all remote, governance often lagged adoption. Legacy archive platforms, in market for twenty or often thirty years, may struggle with the complexity of new media types — particularly video, which is very large; an hour of recording with multiple people can easily be over a hundred megabytes. These meeting records are electronically stored information (ESI), which means they could be discoverable. We may also have different retention for a textual transcript versus the audio or video, and we now need to protect intellectual property, financial data, and roadmap information contained within video conversations.
App proliferation has been a compliance concern for years — over the last two to three years, there’s been about $2 billion in US fines incurred by financial services organizations that couldn’t accurately or consistently capture communications on some of these apps, or inappropriately deleted them. Organizations want their regulated employees to meet customers where they are, but need the right governance wrappers around that, which means archiving platforms need a large array of connectors and integrations, and the ability to trace a conversation as it moves from email to a phone call or chat app for a complete picture.
AI has had a tremendous impact. Natural language to interrogate documents and conversations is becoming mainstream — many users use ChatGPT in daily consumer life, and those habits are coming into work. Traditional keyword and metadata search isn’t disappearing, but AI can enable search to combine natural language, sentiment, intention, and alternative terminology while keeping the value of metadata and dates. It also lets us engage with rich media or very long documents like contracts, summarizing to the key points, and even query structured data in natural language. Our recent software survey found the top two things people do with AI today are researching and finding information, and synthesizing information — both of which lend themselves well to investigations and discovery.
The next thing is the rise of AI agents. I’ve been an advocate for a decade that an enterprise archive sits on a lot of untapped enterprise knowledge. The rise of agents means we can develop assistants to perform tasks — governance agents to optimize storage and retrieval of inactive information, supervision agents that look at the broad scope of a conversation rather than single keywords, legal-hold agents, and classification agents to identify and secure sensitive data. One of our recent AI surveys shows 60% of organizations are exploring or piloting agents in at least one enterprise application, with about 24% in actual production use.
Cheryl:Thinking beyond the archiving assumptions of the last twenty-plus years, buyers today want a broad set of connectors and integrations for an increasingly fragmented set of workplace applications, and in many cases extending governance to the inputs and outputs of AI — the prompts people use and the content AI generates. I’m seeing a trend toward pricing models that support more self-service, opening the archive to more users to ask questions, and more knowledge-discovery use cases alongside the traditional investigative ones. Customers ask about modern governance and security — detecting and flagging sensitive data, supporting data-protection and privacy regulations — and the value of modern cloud platforms, where the vendor takes responsibility for iterative feature delivery, security patches, and continuous improvement. Finally, the fundamentals remain key: a strong foundation of information lifecycle management, with a layered approach to retention tailored to the value of each communication or data type — short-term, medium- and long-term, and even the lifetime of the institution. With that, I’ll turn it over to Aaron.
Aaron:Thank you, Cheryl — a great market overview. One thing I’ve had the benefit of doing in my career is setting up managed services for very large organizations — tens of thousands, hundreds of thousands of people having their communications archived — around the human element of interacting with the archive and delivering on discovery requests. Archives, particularly at larger institutions, have a very long life cycle, not unlike road or electricity infrastructure. Over time, many were originally installed without a clear disposition strategy or thought about what we’d need ten years from now, and as a result, some archives at bigger institutions have petabytes of data. It can be difficult to switch off an established archive with ten, sometimes going on twenty, years of data in it. So in the day-to-day life of someone responding to discovery requests, they’re often working with technology that’s at least ten, sometimes twenty, years old.
Combine that with the evolution in data sources: there’s increasing pressure to produce things like collaboration, Teams chats, and SMS in more native formats so the context isn’t lost. That’s led to a lot of manual processes wrapped around the technology. A simple example: ten data sources go into our main archive and we do discovery there, but we’ve got three data sources we can’t do that with, so we go to the original source and run a separate manual process. And in a regulatory response, there’s huge emphasis on chain of custody — proving you got the data you were supposed to get, ran the search you were supposed to run, got the right results, and delivered them. The standard in US-style discovery is close to perfection, with little leeway for errors. The upshot is it tends to take a very long time — the technology wasn’t designed to search such huge volumes, searches take a long time and can come back incomplete — and it doesn’t scale well; some organizations have dozens of people supporting an active litigation or regulatory portfolio. And those processes are brittle — with so much manual effort, things can get missed.
A lot of this comes back to basic blocking and tackling: being sure you got the data you were supposed to get in the first place — a six-month gap in a feed is a big problem. Completeness and accuracy of search results — without picking on any platform, there are platforms where I’ve seen searches simply not return results from particular datasets. Long-term identity management, as people leave the firm and change names — I know organizations that built entire processes just to manage identities from twenty years ago so analysts can look up all the different IDs to run searches. Getting the data out quickly is critical, especially in regulatory responses against a court deadline. And having the ability to report and evidence all activity is critical. So the question becomes: with twenty-five-plus years of experience, and the benefit of learning from not having a coherent disposition strategy at the outset, how can we bring the latest technology? The shift to the cloud is a game changer, because you can continue to evolve on the platform, and a fully modernized stack can help the analysts in the trenches — and the in-house and outside counsel relying on them — do this far more efficiently, reliably, and at scale. With that, I’ll turn it to Ofir.
Ofir:Thank you, Aaron. I’m Ofir, CTO and co-founder of Shield. I’m going to focus on the modern archive and, of course, AI — and why moving to a modern archive is not only a technological transformation but opens opportunities for redefining and reimagining business workflows. In the market, I’m still seeing organizations running on an archive built fifteen to twenty years ago, with processes bound by the capabilities and limitations of that technology stack. Legacy systems all started as an email archive, but over the years more data sources came in — Slack, Teams, SMS, WhatsApp, WeChat, an endless amount of new channels — and all of them were converted back to email, because that’s what the legacy archive knows how to work with. This created fragmented data: some archives deal only with electronic communications, another for files, another for voice, not combined together, with inconsistent formats and a lot of metadata relevant to compliance use cases lost in the conversion to EML. Add the pressure around data completeness and integrity in recent years, limited visibility, and a painful user experience where end users conduct manual, time-consuming investigations record by record.
When you move from a legacy to a modern archive, you’re going through three simultaneous shifts of the last fifteen years: big data — handling the ever-growing amount of channels and volumes; the cloud, which provides scalability and elasticity; and AI, which brings extreme efficiency. The most fundamental aspect of the modern archive is looking at it as a trusted data foundation for the whole organization, compared to just long-term storage with retrieval. The infrastructure moves from passive, static, on-prem, and siloed to a unified place in the cloud that supports multi-regional data residency — one archive storing US data in the US, EU data in the EU, APAC in APAC, without multiple deployments. The data moves away from flattened EML to a unified, rich, and future-proof data model. From there, actionability is completely different — it supports investigations, surveillance, AI pipelines, and advanced data analysis. And the stakeholders change: it’s no longer only IT and records management, but compliance, legal, surveillance, supervision, and other stakeholders.
Ofir:Now I’d like to focus on the AI revolution — what it means going from a simple storage system to an agentic archive. When I talk about an agentic archive, I mean extreme efficiency, not marginal — and that’s an opportunity to reimagine your business workflows, internally and with external stakeholders. A few examples in the core of archiving work where an AI agent provides extreme efficiency: a data-integrity agent that continuously monitors data completeness, identifies gaps, and recommends remediation — very tedious, time-consuming work for IT teams today. A discovery agent that automates record collection, deep dive, and review of electronic communications for a discovery request, saving days of work. And an export agent that provides export packages with advanced, specific use cases you normally wouldn’t have in standard export capabilities.
Let me drill into one use case: search. The current method is keyword-based — syntax, not semantics or context. Today the flow starts with a discovery request to find conversations on a topic — say, an employee named John who is bullying another employee. You anticipate how the topic might appear across email, chat, and voice transcriptions; maybe John is bullying in Spanish; and English in the US is different from English in the Nordics, especially in informal chat. You identify all the relevant keywords and permutations, apply wildcards and fuzzy matching, get too many irrelevant results, then search within the search, refine the query, apply complex operators to reduce noise — very tedious, with questionable quality of results. At Shield, we introduced Shiela, an AI assistant for discovery cases, where you just ask your query in natural language — “show me all the records where John is bullying another person,” or “show me any signs of profit from a specific stock.” Behind the scenes, AI gathers the information, reviews the communications, filters out what’s not relevant, and shows you what’s relevant by relevancy rate. We’re already seeing massive gains in efficiency.
My recommendation: think about the archive as your data foundation, combined with the extreme efficiency of AI agents, to redefine your business workflows — and find a vendor who can become your partner in this journey, because you’ll face unknowns, like data and formats from twenty years ago you didn’t know existed, and the AI approval process is still evolving. Thank you.
Yael:Thank you, Ofir, Aaron, and Cheryl. Let’s move to the questions you submitted. From a head of market surveillance at a global bank: many vendors mention AI in their presentations — how do you see AI playing a factor in archiving? Cheryl, get us started.
Cheryl:AI has the opportunity to transform how we engage with information across the whole life cycle. Even at ingestion, AI can help enrich metadata, extract insights, and do validation checking to make sure we’ve got everything from the source application. From a user point of view, the ability to use natural language to query and interrogate the archive, and the opportunity for agents to automate tasks and look more broadly than what humans alone could look for — especially with investigations and supervision.
Aaron:The agentic development is potentially huge. As modern archives enable these capabilities, we’ll probably see QA agents in the short term, sooner than agents responding to discovery requests. That said, there’s a great place for something like Shiela in an internal-investigation context, where you’re trying to get your arms around what’s happening — “what is this complaint about John about?” There’s often a lot of time pressure in an internal investigation — maybe not the John one, but an allegation that goes to the board level — so a Shiela-type agent provides tremendous value in understanding what’s in the communications and what the risk is.
Ofir:The agentic opportunity is also a business opportunity to reimagine your roles, where agents work for you. There’s a difference between “I’m doing my job and I get some AI help with some tasks” and reimagining the role so that an agent does the tedious work and leaves you free to deal with the real challenges.
Aaron:My point is: don’t wait for vendors to pitch AI to you. Reimagine what you’d like to get from AI, then put it as a requirement for the vendor — because you know your life and business workflows better than we do.
Yael:I love that — pitch to the vendor, don’t let them pitch to you.
Yael:From a senior compliance officer at a tier-one bank: as employees use tools like ChatGPT or Copilot, how are firms approaching the archive and supervision of those interactions today? Aaron?
Aaron:This is definitely a hot topic. I first heard about it in customer conversations about a year ago as business adoption ramped up. One of the first use cases was data access and policy compliance: are people trying to use ChatGPT to discover information within the organization they wouldn’t otherwise have access to? In the old world of file shares and intranets, you couldn’t just find anything you wanted if you weren’t supposed to — there’s a risk that if a GPT crawls all these data sources, those barriers aren’t maintained. That’s evolved, and now a lot of what I hear is whether these are communications subject to SEC 17a-4-type regulations. That’s an ongoing debate — the prevailing view I’m hearing is that as long as it’s just me and the GPT, it’s not a communication, but as soon as I share that chat with someone else, it becomes a communication that needs to be captured. As with any novel data source, there’s a natural desire to resist having to do more, but some aspects will be unavoidable.
Yael:That’s an interesting point — at the point where it’s just you and the AI, it’s essentially thoughts to yourself written on paper, and whether we identify the AI as a sentient other or just you thinking to yourself is a key definition.
Aaron:Right — and you can play this out further. If it’s actually an agent doing things, is that still just me talking to myself? Because now it’s doing things in the business that an employee would do. So it’s far from a cut-and-dried conclusion, and it’ll be interesting to see how it plays out.
Ofir:We’re seeing early adopters take a risk-based approach — considering the Copilot interaction as a record and archiving it as an additional data source. We’ll see how it evolves.
Aaron:Ultimately, if it’s electronically stored information, it’s potentially discoverable — even a note to yourself, if it’s your rationale behind a business decision. As a requesting party, I’d be after this stuff all the time. Think of the criminal context — there are well-publicized cases where a chat device like Alexa is subpoenaed because it’s evidence in a crime.
Yael:From what you’re seeing, how comfortable are regulators today with AI playing a role in investigations, and where do they still expect manual validation or supporting evidence?
Aaron:In my experience, regulators tend to be pretty outcome-driven — if you provide the information they ask for and there are no obvious gaps, they’re generally satisfied. It’s not that they’re unconcerned with the underlying processes, but they lean outcome-driven. Where it gets more interesting is civil litigation, because the requesting party has an incentive to dig into the discovery and pick at it. So if you’re representing as a responding party that you used AI to find the information, I’d fully expect to be questioned on that.
Ofir:The solution is probably to add a human in the loop, or AI that provides additional reports you can later audit — the behavior of the AI and the decision-making behind the scenes — to gain confidence and approve it internally. Sometimes the challenge is approving the use of AI internally with our internal auditors, even before it goes to the regulator.
Aaron:Building on that — the standard of care in US civil litigation is often that we negotiate boolean search terms, send them over, and you give the results back. Right or wrong, that’s a technology everyone understands. When that turns into “I’ll give you search terms and you use an LLM to turn them into a natural-language query,” that’s where it gets interesting, and there’ll be a lot of internal discussion. If I’m the eDiscovery counsel, I’d be relatively comfortable using an agent or LLM just to figure out what’s in a dataset. But responding to a discovery request, out of the gate I’d probably prefer to still use keywords, because that’s known and I won’t have to argue about it. As data volumes grow, maybe I adopt AI, but I’m going to need validation that it’s the true and correct dataset.
Yael:From a head of compliance at an international bank: how do you balance strict retention policies with the need to keep data usable for investigations and AI models?
Cheryl:This will be an interesting area of debate as we embrace the age of AI. There’s a tension between “keep everything for a long period” and “have a compressed life cycle to dispose of things so you don’t face large-scale, multiyear discovery.” AI brings a new opportunity to use that corporate memory — so for organizations with an aggressive deletion cycle, it means they have holes in their corporate memory. We’re going to need the point of view from records and information management as well as knowledge managers and AI experts on what the right balance is, and it may vary sector to sector.
Aaron:The basics, at least in a US regulatory and discovery context, haven’t changed — you retain what you’re obligated to retain, preserve information for potential litigation, and beyond that keep what’s useful to the business. The huge opportunity now is classification agents. The problem operationalizing classification is why many organizations have thrown up their hands and said “we’ll keep everything forever,” because they couldn’t operationalize the decisions — better to risk overproducing than a discovery sanction for deleting something. Setting up a program where people manually classify records has never worked, but this is a great job for agents. The challenge will be getting comfortable they’re doing it correctly, but it’s a massive opportunity.
Ofir:Another limitation today is scale and cost, which can be a showstopper at large volumes. We don’t hear a lot about it lately because everything is AI, but in parallel there’s continuous innovation in querying big data and applying AI on big data in the cloud, so if we adopt new technologies, we can apply AI on massive amounts of data at a reasonable cost.
Yael:We’re coming to time. Thank you to our wonderful panelists — this has been a great, informative discussion, and it goes to prove archiving can be a lively and innovative topic. To sum up one main takeaway: archiving is no longer just about retaining data. It’s about being able to trust it, use it, and act on it — and the foundation for all of that is the quality and completeness of your data. Without it, firms can’t move fast enough, can’t investigate effectively, and it won’t stand up to regulatory scrutiny. That’s why this shift matters: turning the archive from passive storage into something that actively supports investigations, decisions, and real business workflows. Thank you once again to our speakers and to everyone who joined today.
Archiving is mature — Forrester’s data shows around 80% of organizations already have some form in place — but the last one to two years have been the most innovative in a long time. Multimodal conversations (audio, video, transcripts), roughly $2 billion in fines for capture failures, and the arrival of AI are pushing firms to treat the archive not as compliance storage but as a strategic, trusted data asset.
Legacy systems started as email archives, so as new channels arrived — Slack, Teams, SMS, WhatsApp, WeChat — everything got converted back to email (EML). That flattening fragments the data across separate systems for comms, files, and voice, and strips out metadata that matters for compliance. A modern archive keeps a unified, rich, future-proof data model in the cloud, with multi-region data residency in a single deployment.
It changes the whole life cycle. At ingestion, AI can enrich metadata and validate completeness. For users, natural-language querying replaces brittle keyword search, and it can summarize rich media or long contracts. And agents can automate tasks and look more broadly than humans alone. Forrester found about 60% of organizations are exploring or piloting agents, with roughly 24% already running one in production.
It’s an archive where AI agents do the tedious, time-consuming work, delivering extreme rather than marginal efficiency. Examples from the panel: a data-integrity agent that continuously monitors completeness, flags gaps, and recommends remediation; a discovery agent that automates record collection and review; and an export agent that builds tailored packages. The point isn’t just assisting your job — it’s reimagining roles so agents handle the grunt work.
Keyword search means anticipating every way a topic might appear — across email, chat, and voice transcripts, in multiple languages and regional slang — then applying wildcards, wading through noise, and searching within the search. With natural language, you just ask: “show me records where John is bullying another person.” Shield’s Shiela assistant gathers the communications, filters out the irrelevant, and returns results ranked by relevancy, cutting a tedious process dramatically.
It’s an active debate. The prevailing view is that a private exchange between you and the model isn’t a communication — but the moment you share that chat with someone else, it likely becomes one that must be captured. And if it’s electronically stored information, even a note-to-self rationale is potentially discoverable. Some early adopters are already taking a risk-based approach and archiving these interactions as a record.
Regulators tend to be outcome-driven — if you provide what’s asked and no gaps emerge, they’re generally satisfied. Civil litigation is where it gets tricky, because the requesting party has an incentive to scrutinize AI-based methods. The practical guidance: keep a human in the loop, generate auditable reports of the AI’s decision-making, and, for now, keywords remain the safer default when responding to a discovery request.
Retain what regulations require, what litigation demands, and what’s genuinely useful — but recognize that aggressive deletion leaves holes in your “corporate memory” that AI could otherwise mine. Manual classification programs have historically never worked, which is why many firms defaulted to keeping everything. Classification agents finally make defensible, tiered retention operational at scale — with cost and volume still a real design consideration.