A/B Testing
A method of comparing two versions of a system, message or workflow to evaluate which performs better against defined criteria. It may be used in legal-tech product design but requires care where legal outcomes are affected.
eDiscovery Certification Council
The complete eDCC glossary of 1,200 terms and phrases spanning eDiscovery, electronic evidence, eForensics, artificial intelligence, privacy, information governance, litigation support and legal technology.
1,200 terms and phrasesUse the glossary as a quick-reference dictionary or as a gateway into the wider eDCC Knowledge Hub. Definitions are concise, vendor-neutral and written to connect technical, legal, evidential and governance concepts.
A method of comparing two versions of a system, message or workflow to evaluate which performs better against defined criteria. It may be used in legal-tech product design but requires care where legal outcomes are affected.
The summarisation of key facts, fields or attributes from longer documents or records into a structured form.
An organisational policy defining permitted and prohibited use of information systems, devices, networks and data. In eDiscovery and investigations it may help establish expected user behaviour and relevant governance controls.
Technical and administrative measures that restrict access to systems, data or functions to authorised users, roles or processes.
A record showing access to a system, file, application or data repository, often including user, time, action and source information. Access logs can be important evidence in investigations and authenticity analysis.
A request to obtain information or access rights. In privacy practice, it may refer to a data-subject access request; in systems it may refer to permission to data or functionality.
The obligation to take responsibility for decisions and controls and to be able to demonstrate compliance, governance and appropriate oversight.
The data-protection principle requiring organisations to take responsibility for compliance and be able to demonstrate it.
The data-protection principle requiring personal data to be accurate and, where necessary, kept up to date, with reasonable steps taken to correct or erase inaccurate data.
The act of obtaining data from a source. In digital forensics, acquisition is performed using controlled methods designed to preserve integrity.
Data readily available to users through live systems and applications, as distinct from archived, backup or deleted data.
A machine-learning approach in which the model selects or prioritises examples for human review so that human decisions can improve subsequent predictions. Continuous Active Learning is a prominent eDiscovery application.
A technology-assisted review approach in which the system selects additional training documents expected to most improve the model, rather than relying solely on random or user-selected samples.
A formal determination that a country, territory, sector or international organisation provides an adequate level of protection for personal data, allowing specified international transfers without additional transfer safeguards.
Review focused on completeness, formatting, scope, documentation and compliance with required procedures.
The legal acceptability of evidence for consideration by a court or tribunal under the applicable evidential rules.
Analytical methods beyond basic keyword search, including clustering, communication analysis, anomaly detection, predictive modelling, semantic search and machine learning.
The study of attacks against machine-learning systems and methods for resisting or mitigating those attacks, including evasion, poisoning and related manipulation.
A prompt intentionally crafted to manipulate, bypass or degrade an AI system, including attempts to override instructions or expose protected information.
A conclusion a court may be permitted to draw against a party because relevant evidence was destroyed, withheld or not properly preserved, depending on the applicable law and circumstances.
Coordination of multiple AI agents, tools or workflow steps to complete a broader task.
AI systems designed to plan and execute multi-step tasks using tools, memory, workflows or other agents with a degree of autonomy. In eDiscovery, agentic systems may coordinate tasks such as search, summarisation, classification and reporting but require governance and auditability.
A software-based AI component that observes context, makes decisions and takes actions toward a defined objective, often using external tools or data sources.
Independent or internal processes used to provide confidence that AI systems are governed, tested, monitored and fit for intended use.
A structured examination of an AI system, its data, controls, performance, governance and risks against defined criteria.
A dataset, task or evaluation protocol used to compare the performance of AI models or systems.
An AI assistant embedded in a workflow or application to support users with tasks such as search, summarisation, drafting, analysis or decision support.
Systematic testing of an AI model or workflow against defined criteria such as accuracy, factuality, robustness, safety, consistency and task usefulness.
The extent to which users can understand or obtain meaningful information about how an AI system produced an output or decision.
The policies, roles, controls and oversight used to ensure AI is developed or used lawfully, safely, ethically, securely and in accordance with organisational objectives.
An AI-generated statement or output that appears plausible but is unsupported, incorrect or fabricated. Hallucinations are a significant validation risk in legal and evidential work.
An event in which an AI system causes or contributes to material harm, policy breach, security failure, inaccurate outcome or other significant operational problem.
A maintained register of AI systems, models, use cases, owners, data dependencies, risk levels and governance status within an organisation.
The knowledge and skills needed to understand AI capabilities, limitations, risks and appropriate use in a professional context.
A computational system trained or configured to generate predictions, classifications, content or other outputs from input data.
Structured testing in which evaluators deliberately probe an AI system for weaknesses, unsafe behaviour, misuse pathways or control failures.
A structured process for identifying, assessing, controlling, monitoring and documenting risks associated with AI systems and their use.
The discipline concerned with reducing harmful, unreliable or unintended outcomes from AI systems through design, testing, controls and monitoring.
A machine-based system that generates outputs such as predictions, recommendations, classifications, content or decisions from inputs for explicit or implicit objectives.
Information provided to users or affected persons explaining that AI is being used and, where appropriate, its purpose, role, limitations and relevant rights or safeguards.
The testing and evaluation of an AI system or workflow to determine whether its outputs are sufficiently accurate, reliable, consistent and fit for the intended legal or operational purpose.
A defined sequence of AI-assisted and human activities used to complete a task or process.
A document review workflow in which artificial intelligence helps classify, prioritise, summarise, analyse or quality-check documents while human reviewers retain oversight.
A defined set of computational rules or procedures used to solve a problem, process data or produce an output.
Systematic differences in AI or algorithmic outputs that may unfairly disadvantage particular people, groups or circumstances because of data, design, assumptions or deployment conditions.
The degree to which an AI system behaves consistently with intended objectives, instructions, constraints, laws and human values.
A file-system storage unit used to hold file data; understanding allocation can assist forensic recovery and analysis.
A revised disclosure statement or related procedural document reflecting corrections or changes to a party's disclosure position, where required by the applicable regime.
Methods used to examine data for patterns, relationships, trends or other useful information. In eDiscovery, analytics may include email threading, clustering, communication analysis, TAR and near-duplicate detection.
Statistical and machine-learning techniques used to organise ESI, surface patterns, prioritise documents, reduce review volume and generate case insights.
The addition of notes, labels, tags or structured information to data or documents. In machine learning, annotation commonly creates labelled examples.
Analytical techniques used to identify records, communications, transactions or behaviours that differ materially from expected patterns.
Processing personal data so that individuals are no longer identifiable by means reasonably likely to be used. Properly anonymised information is generally no longer personal data under data-protection law.
Techniques intended to hide, alter, destroy or mislead forensic evidence or obstruct forensic analysis.
A defined interface that allows software systems to exchange data or invoke functionality. APIs are increasingly important for cloud collection and legal-tech automation.
Collection of cloud or application data through an authorised API, often preserving structured fields and metadata unavailable through simple screenshots or manual exports.
A defined network address through which an application programming interface exposes specific data or functionality.
Data created by a specific application, such as chat databases, caches, logs, preferences or local storage.
A record generated by software documenting events, errors, actions, transactions or user activity.
Metadata generated by a software application, such as document properties, author fields, message identifiers, revision details or application-specific attributes.
A file format selected for long-term preservation because it is durable, open, well documented or widely supported.
A repository or process for retaining information over time, commonly for business, legal, regulatory or historical purposes.
A secondary mailbox or storage location used to retain older or additional email and related items, often subject to eDiscovery searches.
A data object or trace created by system or user activity. In digital forensics, artefacts can include logs, registry entries, browser records and application data.
A chronological reconstruction created from multiple digital artefacts such as logs, file timestamps, browser history and messages.
A hypothetical form of AI capable of broad, flexible intellectual performance across many domains at or above human level. It should not be confused with today's task-specific AI systems.
A broad field concerned with computer systems capable of performing tasks associated with human intelligence, such as language understanding, classification, reasoning, perception or decision support.
A file or object associated with an email, message or other parent item. In eDiscovery, attachments are often treated as part of a document family.
A legal protection that may prevent disclosure of confidential communications between a lawyer and client made for the purpose of obtaining or providing legal advice, subject to applicable law and exceptions.
The forensic examination and analysis of audio recordings to address authenticity, enhancement, speaker, event or evidential questions.
A chronological record of system or user activity created to support accountability, security, compliance and investigation.
A documented history showing actions, decisions, transfers or changes affecting data, evidence or a workflow. A strong audit trail supports repeatability and defensibility.
The process of establishing that evidence, data, a user or a system is what it is claimed to be. For digital evidence, authentication may rely on metadata, hashes, witnesses, logs or other corroborating information.
The process of establishing that digital evidence is what it purports to be and has not been altered, typically supported by hash values, metadata and chain-of-custody records.
Metadata identifying or indicating the creator, editor or account associated with a file, document or message.
Automated assignment of documents, data or records to categories based on rules, metadata, machine learning or AI.
Automated identification and masking of sensitive information based on patterns, rules or AI, subject to quality control.
A process in which a decision is made wholly or partly by automated means, potentially triggering specific transparency or rights obligations under applicable law.
Rights or safeguards that may apply when individuals are subject to solely automated decisions producing legal or similarly significant effects, depending on the applicable privacy regime.
Use of rules, analytics or AI to identify documents that may contain privileged communications or work product for further review.
Use of technology to classify, prioritise, summarise or otherwise assist review of documents with limited or targeted human intervention.
A copy of data created to support recovery after deletion, corruption, failure or disaster. Backup systems may become relevant to discovery where responsive data is not reasonably available elsewhere.
A planned schedule for creating, retaining, reusing or replacing backup media or copies.
A defined group of backup copies created according to a particular schedule, system or recovery plan.
Magnetic tape used for backup or archival storage. Restoring and searching backup tapes may be costly and can raise accessibility and proportionality issues.
The period during which backup operations are scheduled to run.
A text-analysis representation that treats a document as a collection of words or tokens, usually without preserving word order. It is used in some search and machine-learning methods.
Intentional limitation of network transfer speed, sometimes used during remote collections to avoid disrupting business systems.
A reference state, dataset or performance level used for comparison in testing, monitoring or investigation.
A defined group of documents assigned or processed together, for example a reviewer batch in document review.
A review-management method in which documents are divided into groups and allocated to reviewers for coding, quality control or specialist review.
A unique sequential identifier applied to pages or documents, commonly used to track, cite and control produced material.
The process of assigning unique sequential identifiers to documents or pages in a production or legal record set.
See Bates Numbering.
Planning intended to maintain or restore essential organisational functions during disruption. Information systems and recovery arrangements identified in BCP may be relevant data sources.
Analysis of user, system or communication behaviour to identify patterns, anomalies or relationships relevant to security, investigations or legal review.
A principle concerning use of original or reliable evidence where required by applicable evidential rules. Its precise operation is jurisdiction-specific.
A method or approach generally regarded as effective or prudent based on experience, guidance, evidence or professional consensus, but not necessarily legally mandatory.
Objective coding of basic document attributes such as date, author, recipient and document type.
A system using two states, typically 0 and 1, in which digital information is represented and processed.
A database object used to store large binary data such as images, documents, audio or other files.
Internal data-protection rules approved for certain international transfers of personal data within a multinational corporate group or group of enterprises.
The smallest conventional unit of digital information, representing a binary value of 0 or 1.
A copy intended to reproduce every bit from a source medium or defined data area. In forensics, integrity is normally verified using cryptographic hashing.
A forensic image capturing the sequence of bits from a storage source, commonly including allocated and unallocated space within the acquired scope.
A model whose internal decision process is difficult for users to understand or explain, even if inputs and outputs are observable.
Evidence derived from or related to blockchain systems, such as transactions, wallet activity, smart contracts or distributed ledger records.
A probabilistic data structure used to test whether an item may be a member of a set, with possible false positives but no false negatives under normal implementation assumptions.
Logical connectors (AND, OR, NOT, proximity) used to construct precise keyword searches.
A search technique using operators such as AND, OR and NOT to combine or exclude search conditions.
The legally prescribed period for notifying a regulator or affected individuals after a qualifying personal-data breach, depending on jurisdiction.
The set of individuals or records potentially affected by a data breach and requiring analysis or notification.
Technology used to review compromised data, identify personal information and determine notification populations following a cyber incident.
An organisational arrangement permitting personnel to use personally owned devices for work. BYOD complicates preservation, privacy and collection.
A data item generated by web-browser use, such as history, cache, cookies, downloads, saved credentials or session information.
Stored browser data used to populate forms automatically, potentially containing personal or account information.
Locally stored copies of web content used to improve browsing performance and potentially containing relevant historical evidence.
A small browser-stored data item used for session, preference, authentication or tracking purposes and potentially relevant to forensic analysis.
Records showing files downloaded through a browser, often including URL, destination path and timestamp.
A record of websites or resources accessed through a browser, potentially useful in investigations and chronology reconstruction.
A period of interaction between a user and web application, often associated with cookies, tokens and session identifiers.
The effort, cost, time, technical difficulty or operational impact associated with preservation, collection, review or production. Burden is often considered alongside relevance and proportionality.
The predecessor regime that led to Practice Direction 57AD in England and Wales, introducing disclosure models and structured disclosure planning in the Business and Property Courts.
Information created, received or maintained in the course of organisational operations.
A governed collection of agreed business terms and definitions used to promote consistent understanding of data, concepts and responsibilities across an organisation.
Technologies and methods used to analyse organisational data and present insights through reports, dashboards and metrics.
Information created or maintained as part of an organisation's activities and retained for operational, legal, regulatory or evidential purposes.
An evidential principle in some jurisdictions permitting qualifying business records to be admitted despite hearsay rules, subject to the applicable legal requirements.
Policies and discovery issues arising when employees use personal devices for work and those devices may contain relevant ESI.
A unit of digital information commonly consisting of eight bits.
The Coalition for Content Provenance and Authenticity specification family for attaching cryptographically verifiable provenance information to digital media.
Temporary stored data used to improve performance or access. Cached content can contain evidential information not readily visible through the live application.
See Continuous Active Learning.
California privacy legislation granting specified rights concerning personal information and imposing obligations on covered businesses, as amended by later legislation.
A record generated by telecommunications systems showing information about calls or messages, such as numbers, timestamps, duration and routing data.
A file recovered from raw storage by recognising file structure or signatures rather than relying on normal file-system references.
A user role responsible for configuring and managing an eDiscovery case, permissions, searches, holds or review sets.
The early evaluation of facts, evidence, risk, scope and likely costs associated with a legal matter or investigation.
A structured timeline of events, evidence and references used to understand or present a legal matter.
Formal completion of a legal or investigative matter, including finalising productions, releasing holds, archiving records and documenting outcomes.
A unique value used to associate evidence, documents, workflows and reports with a particular matter or investigation.
Judicial decisions that interpret and apply legal principles and may have binding or persuasive authority depending on the jurisdiction and court hierarchy.
The organisation and control of tasks, deadlines, evidence, participants, documents and decisions within a legal matter.
A court hearing or procedural meeting at which the court and parties address how a case should proceed, potentially including disclosure or discovery issues.
A structured collection of notes, issues, evidence references and analysis maintained for a legal matter or investigation.
The legal and factual plan for managing a matter, including issues, evidence, procedural steps, risks and objectives.
The group of legal, technical, investigative, client and support professionals responsible for a matter.
The documented record of the possession, transfer, handling and control of evidence from collection through analysis, storage and presentation.
A document recording evidence identifiers, possession, transfers, dates, times and signatures or acknowledgements.
Messages and associated metadata from chat or collaboration platforms, including direct messages, channels, threads, reactions, edits and attachments.
An exported representation of chat or collaboration data, potentially containing messages, threads, reactions, attachments and metadata.
A sequence of related chat messages grouped by conversation, reply structure or topic.
A value calculated from data to help detect changes or transmission errors. Cryptographic hashes are generally stronger integrity mechanisms than simple checksums.
An item associated with a parent document, such as an email attachment, embedded object or file within an archive.
An older term for an ethical or information barrier; many organisations now prefer terms such as ethical wall or information barrier.
Dividing large documents or datasets into smaller units for indexing, retrieval or AI processing.
The assignment of data or documents to defined categories based on content, metadata, rules, human judgement or machine learning.
An agreement allowing parties to recover material disclosed inadvertently, commonly privileged information, subject to applicable rules and terms.
A contractual or court-approved term setting out how inadvertently produced privileged information can be returned or protected.
A jurisdiction-specific form of legal professional privilege protecting qualifying confidential lawyer-client communications.
Another expression for attorney-client or legal professional privilege, depending on jurisdiction.
Difference between the actual time and a device or system clock, which can affect timeline analysis.
A record of administrative, user or system actions generated by a cloud service and potentially valuable for investigations and eDiscovery.
Collection of data directly from cloud-hosted systems, often using application exports, APIs, administrative tools or specialist collection technology.
Delivery of computing resources such as storage, applications or processing over networked infrastructure, commonly provided as shared or scalable services.
A tool or integration used to acquire or synchronise data from cloud services into another system such as an archive, SIEM or eDiscovery platform.
A documented plan for retrieving, migrating or deleting organisational data when ending use of a cloud service.
The preservation, collection and analysis of evidence from cloud services and infrastructure, where traditional physical-device acquisition may not be possible.
A cloud-hosted backup of mobile-device data that may be accessible for forensic or eDiscovery collection.
A policy or configuration defining how long cloud-hosted content is retained, deleted or preserved.
An organisation providing cloud infrastructure, platform or software services.
A point-in-time representation of cloud storage, virtual machines or systems used for backup, recovery or forensic preservation.
Data created and maintained primarily within cloud applications or services rather than as conventional local files.
Authoritative practical guidance issued by a regulator, standards body or professional organisation to explain expected conduct or recommended procedures.
The process of assigning tags, values or categories to documents during review, such as responsiveness, issue, confidentiality or privilege.
The process of applying tags or classifications to documents (for example relevance, privilege, issues) either manually by reviewers or automatically by systems.
A reviewer's recorded determination about a document, such as responsive, privileged, confidential or issue-related.
The set of fields, tags and controls presented to reviewers for recording document-review decisions.
The set of review fields, tags and controls displayed to reviewers for document coding.
Information created in collaboration platforms such as Microsoft Teams, Slack or similar services, including chats, channels, reactions, files and linked content.
The acquisition of potentially relevant electronically stored information from identified sources for processing, review, investigation or production.
An estimate of the quantity of content expected from a search or collection before actual ingestion or export.
A record documenting what data was collected, from where, by whom, when, by what method and with what validation information.
The technical and procedural approach used to acquire data, such as forensic imaging, logical export, API collection or targeted file collection.
The boundaries of a collection exercise, such as custodians, systems, date ranges, file types, locations and data categories.
A specific source, account, device, folder, custodian or dataset selected for collection.
A standardised structure or schema used to represent data consistently across systems or processes.
Analysis of message patterns, participants, frequency, timing and relationships to identify important people, events or networks.
A visual or analytical representation of communications among people, accounts or organisations.
Analysis of sender-recipient relationships, frequency, centrality and communication patterns to identify important participants and relationships.
A structured description of the knowledge, skills and behaviours expected for roles or levels of professional practice.
A repository designed to retain communications or records in accordance with legal, regulatory or organisational obligations.
A search conducted across organisational repositories for legal, regulatory, audit, investigation or information-governance purposes.
A dataset exposed, accessed, altered or exfiltrated during a security incident.
The scientific examination, analysis or evaluation of digital evidence from computers and storage systems for investigative or legal purposes.
A broad term for technology used to support document review, including search, analytics, clustering and machine-learning-assisted classification.
An earlier term largely synonymous with Technology-Assisted Review (TAR); the use of computer algorithms to facilitate or prioritise document review.
An analytics technique that groups documents by conceptual similarity, often without requiring predefined categories.
Grouping or retrieving documents based on semantic similarity or themes rather than exact keyword matches.
Search based on semantic similarity or meaning rather than exact keyword matches alone.
A user interface for constructing structured search queries using fields, keywords and logical operators without manually writing the complete query syntax.
A statistical range calculated from sample data that is intended to contain an unknown population value with a stated level of confidence under the assumptions of the method.
The principle that information is accessible only to authorised persons or processes and protected from unauthorised disclosure.
Another term for a confidentiality ring used in some jurisdictions.
A label applied to documents or information to indicate treatment under a confidentiality agreement, protective order or internal policy.
A restricted group of authorised individuals permitted to access highly confidential material under an agreement or court order.
A table comparing predicted classes with actual classes, showing true positives, true negatives, false positives and false negatives.
One possible lawful basis for processing personal data where an individual freely gives a specific, informed and unambiguous indication of agreement, subject to applicable data-protection law.
A court order reflecting terms agreed by the parties and approved by the court, potentially including disclosure or discovery arrangements.
A file that can contain other files or data objects, such as ZIP, PST, OST or archive formats.
A searchable index representing the textual content and metadata of data held in a system.
Information that helps establish where content originated, how it was created and whether it was modified.
A search across stored information using content, metadata or other indexed properties to identify items matching defined criteria.
The amount of input information an AI model can consider within a single interaction or processing request, commonly measured in tokens.
A TAR workflow in which reviewer decisions are repeatedly used to update or influence the system's prioritisation of documents during the review.
Ongoing review of systems, controls, risks or performance rather than one-time assessment.
Use of software or AI to extract, classify and analyse contractual terms, obligations, risks or patterns.
The process and technology used to create, negotiate, approve, execute, store, monitor and renew contracts.
A comparison group not exposed to a tested intervention or treatment, used to help evaluate causal effects. In eDiscovery analytics, the term may also be used informally for comparison datasets.
A set of documents with known coding used to measure the performance of a TAR or predictive-coding model.
A curated set of approved terms used consistently for classification, indexing or retrieval.
Under data-protection law, the person or organisation that determines the purposes and means of processing personal data.
A personal-data transfer in which one controller discloses data to another independent controller, each determining its own purposes and means.
A duplicate kept for reference but not designated as the official record copy.
Changing data from one format to another, such as rendering a native file to TIFF or PDF.
A small data item stored by a web browser or service to maintain state, preferences, authentication or tracking information.
A mechanism for obtaining or recording user choices regarding cookies or similar technologies where consent is legally required.
Key documents central to the issues in a case and identified for early disclosure under certain procedural regimes.
A body or collection of documents or text used for search, analysis, review or machine learning.
A request by an individual to correct inaccurate or incomplete personal information under applicable privacy law.
Damage or alteration that makes data incomplete, inconsistent, unreadable or unreliable.
A mathematical measure commonly used to compare the direction of vectors and estimate semantic similarity between embeddings.
Disclosure required by a court order specifying the scope, timing, method or format of documents to be disclosed.
Personal data relating to criminal convictions and offences, subject to specific restrictions under UK/EU data-protection law.
A legal safeguard or basis used to permit international transfers of personal data, such as adequacy, standard contractual clauses or binding corporate rules.
Discovery or disclosure involving data, parties, systems or legal obligations across multiple jurisdictions.
A statistical method for estimating model performance by training and testing across different subsets of the available data.
Rendering encrypted data unreadable by securely destroying the relevant encryption keys.
A hash produced by a cryptographic hash function designed to make it computationally infeasible to reconstruct the input or find practical collisions.
Reducing a data population before detailed review using defensible techniques such as date filtering, file-type filtering, deduplication, DeNISTing or search.
A source primarily associated with a specific individual custodian, such as their mailbox, laptop or OneDrive account.
A person or organisational role associated with potentially relevant information and whose data may be subject to preservation, search, collection or review.
A record linking a custodian to potentially relevant devices, accounts, applications, repositories and data sources.
A structured discussion used to identify a person's relevant activities, terminology, devices, applications, repositories and potential sources of information.
A standardised set of questions used to gather information from custodians about data sources, preservation obligations and relevant subject matter.
Monitoring custodians through identification, legal hold, interview, collection and release stages.
A deduplication method that removes duplicate copies within each custodian's data while preserving a copy for every custodian who possessed the document.
Use of eDiscovery technologies and workflows to identify, process and review data in connection with cyber incidents, breach response and investigations.
An event affecting the confidentiality, integrity or availability of systems or data, such as unauthorised access, malware or data exfiltration.
The coordinated process of detecting, containing, investigating, remediating and learning from cybersecurity incidents.
An archive retained for preservation or compliance but not routinely accessed or indexed for business use.
Information an organisation retains but does not actively use or understand, often creating cost, security, privacy and discovery risk.
Recorded facts, information, signals or representations capable of being stored, processed or communicated.
UK legislation affecting data use, access and aspects of the UK data-protection framework; related guidance should be checked for current implementation and transitional effects.
A request to obtain access to data, whether under privacy rights, litigation, internal governance or system permissions.
The examination of data using statistical, computational or visual methods to identify patterns, relationships, anomalies or insights.
A security incident resulting in accidental or unlawful destruction, loss, alteration, unauthorised disclosure of, or access to protected information, depending on the applicable legal definition.
A notification to a regulator, affected individuals or other parties following a qualifying personal-data or security breach, as required by applicable law.
Review of compromised or potentially compromised data to identify affected people, personal information, confidential material, privilege or regulatory exposure.
Rapid initial assessment of a suspected breach to determine scope, affected data, individuals, systems, risk and notification obligations.
A structured inventory describing data assets, locations, owners, definitions, lineage and other metadata to improve discovery and governance.
The assignment of data to categories based on sensitivity, business value, legal status or handling requirements.
A controlled environment in which multiple parties can analyse data under technical and governance restrictions designed to limit exposure of raw information.
A common expression for a controller under data-protection law: the entity determining why and how personal data is processed.
A person or function responsible for the operational care, maintenance or protection of data. The term differs from an eDiscovery custodian, who is associated with potentially relevant information.
The process of locating, identifying and understanding data across systems, repositories and formats.
The authorised removal or destruction of data using methods appropriate to sensitivity, legal requirements and storage media.
Unauthorised transfer of data from a system, network or organisation.
The extraction and transfer of data from a source system into another format or destination.
The framework of decision rights, standards, accountability and controls used to manage data quality, ownership, access, security and use.
A documented list of data assets, repositories, systems, categories and associated governance information.
A repository designed to store large volumes of structured, semi-structured and unstructured data, often in native or near-native form.
The stages through which data passes from creation or acquisition to use, storage, retention, archival and deletion.
Information describing where data originated, how it moved and how it was transformed across systems and processes.
Legal or policy requirements that certain data be stored, processed or retained within specified geographic boundaries.
Controls designed to identify and prevent unauthorised or inappropriate disclosure, transfer or use of sensitive data.
A documented inventory or representation of data sources, systems, owners, locations, flows, formats and retention characteristics.
The process of identifying and documenting where information is created, stored, transferred, controlled and retained.
Obscuring or substituting sensitive data while preserving structure or usability for testing, analysis or controlled access.
The transfer of data from one system, format or storage environment to another while preserving required content and metadata.
The principle of limiting personal data to what is adequate, relevant and necessary for the intended purpose.
A person or function accountable for decisions about a data asset, such as access, quality, retention or approved use.
A privacy right in some regimes enabling individuals to receive qualifying personal data in a structured, commonly used, machine-readable format and, where applicable, transmit it elsewhere.
An export prepared to satisfy a data-portability request in a structured, commonly used and machine-readable format.
A contract governing processing of personal data, commonly used between a controller and processor to address legally required terms and safeguards.
An entity that processes personal data on behalf of a controller.
The legal, organisational and technical framework for protecting personal data and regulating how it is collected, used, shared, retained and secured.
The UK statute that supplements and works alongside the UK GDPR and contains additional data-protection provisions.
A legal and governance approach requiring privacy and data-protection principles to be embedded into systems and processing from the outset and by default.
The effect a processing activity may have on individuals' privacy, rights and freedoms.
A structured assessment used to identify and reduce privacy risks from processing likely to result in high risk to individuals.
A designated privacy professional required in specified circumstances under data-protection law to advise, monitor compliance and act as a contact point for regulators and individuals.
The degree to which data is accurate, complete, consistent, timely and suitable for its intended use.
The physical or legal location in which data is stored or processed.
The continued storage of information for defined business, legal, regulatory or evidential periods.
The defined period for which a category of information should be kept before review, archival or disposal.
A policy defining how long categories of information should be retained and when they should be archived or deleted.
An agreement governing the purposes, responsibilities, safeguards and conditions under which organisations share data.
A system, device, repository, application, account or other location from which potentially relevant information may be obtained.
The concept that data is subject to the laws and governance requirements of the jurisdictions in which it is stored, processed or controlled.
A role responsible for operational governance and quality of specified data, often acting under policies set by data owners.
An identified or identifiable individual to whom personal data relates.
A request by an individual to exercise the right of access to personal data held about them, subject to applicable law and exemptions.
A legal right exercisable by an individual concerning personal data, such as access, rectification, erasure, restriction, objection or portability, depending on the applicable law.
Searches intended to locate personal data relating to a specified individual across relevant systems.
Steps taken to verify the identity and authority of a person exercising privacy rights before disclosing or altering personal data.
Movement or disclosure of data between systems, organisations, locations or jurisdictions.
An assessment of legal and practical risks associated with an international transfer of personal data and the effectiveness of relevant safeguards.
Changing the structure, format or representation of data while preserving or intentionally modifying its meaning.
Checks performed to confirm that data is complete, accurate, consistent or correctly transformed for its intended purpose.
A structured collection of data managed so that information can be stored, queried, updated and retrieved.
A structured extraction of records, tables or fields from a database for review, analysis or migration.
A point-in-time copy or representation of a database used for recovery, analysis, preservation or comparison.
A culling or search condition limiting data to a defined time period based on specified date fields.
Limiting a collection or review set to documents falling within specified creation, sent or modified dates.
See Deduplication.
Techniques used to reduce or remove the association between data and identifiable individuals; the precise legal effect depends on the method and applicable law.
The removal of known system and application files (identified via NIST hash lists) from a collected data set to reduce volume before review.
A branching model used to represent rules or decisions and their possible outcomes; it may support legal workflows, investigations or machine learning.
The controlled retirement of a system, application or repository, including decisions about migration, preservation and disposal of its data.
A preservation control applied before retiring a system to ensure relevant data is not lost during migration or shutdown.
The process of expanding compressed data or archives into their constituent files or original representation.
The process of identifying and resolving conflicts, duplication or overlap between investigative, legal, technical or collection activities.
The process of converting encrypted data back into readable form using authorised credentials, keys or mechanisms.
The identification and removal or suppression of duplicate data, commonly using hash values or other comparison techniques, to reduce unnecessary processing and review.
The process of identifying and suppressing exact or near-duplicate documents within a collection so that only one instance is reviewed or produced.
A machine-learning approach using multi-layer neural networks to learn complex representations from data.
Synthetic or manipulated audio, video or imagery generated or altered using AI so that it appears authentic. Deepfakes create significant authentication and evidential challenges.
Techniques used to identify synthetic or manipulated media through forensic analysis, provenance, metadata or AI-based detection methods.
Forensic examination of potentially synthetic or manipulated media to assess authenticity and detect signs of alteration.
The ability to explain and support a process, decision or result as reasonable, proportionate, documented and appropriate under the circumstances.
The routine disposal of information under documented and consistently applied policies when no legal, regulatory, preservation or business obligation requires continued retention.
Data that has been removed from normal user access but may remain recoverable from storage, backups, application history or other sources.
Techniques used to recover data that has been deleted but not yet overwritten.
Recovery of database, file-system or application records that are no longer active but remain recoverable from underlying storage.
The removal of data from active use or storage. Technical deletion may not immediately erase all underlying information.
A control that suspends routine deletion for specified information while a preservation requirement remains active.
A request by an individual to erase qualifying personal information under applicable privacy law.
Collection limited to data added or changed since a previous collection, commonly used for rolling or periodic acquisitions.
Filtering files using known-file hash sets, commonly from NIST resources, to remove standard operating-system and application files unlikely to be user-created evidence.
A document or other item formally used or marked during a deposition or witness examination.
The organisation of deposition scheduling, exhibits, transcripts, video and related evidence.
Evidence created or obtained from another source item, such as a working copy, extracted report or converted image.
The process of linking activity or data to a particular device using identifiers, logs, metadata or forensic evidence.
A value used to identify a device, such as serial number, IMEI, MAC address or platform-specific identifier.
A mathematical approach designed to limit what can be inferred about individuals when analysing or releasing aggregate datasets.
Chain-of-custody records maintained electronically, documenting evidence possession, handling, transfer and verification.
Electronic information stored or transmitted in digital form that may have evidential value.
A controlled digital package or repository containing evidence and associated metadata, hashes and custody information.
A person trained to identify, secure and preserve digital evidence at the initial stage of an incident or investigation.
Formal recognition that a laboratory meets specified competence and quality requirements for digital forensic work.
The technical and procedural measures used to prevent alteration, loss or destruction of digital evidence.
Rapid initial examination used to prioritise devices, accounts or artefacts for deeper forensic analysis.
The application of scientific and investigative methods to identify, acquire, preserve, examine, analyse and report digital evidence.
A controlled environment, including procedures, tools and infrastructure, used for forensic examination of digital evidence.
An investigation in which digital systems, data or evidence play a material role in establishing facts or events.
A formal report describing the scope, methods, evidence, findings, limitations and conclusions of a digital investigation.
A cryptographic mechanism used to verify the origin and integrity of digital data, commonly based on public-key cryptography.
A digital representation of a physical object, system or process that may receive real-world data and support simulation or monitoring.
Embedding detectable information into digital content to support attribution, provenance, rights management or authenticity.
Marketing communications directed to specific individuals or organisations and subject to privacy and communications rules.
In England and Wales, the process by which a party states that documents exist or have existed and, where required, makes them available for inspection or provides them in accordance with court rules.
The process under the Civil Procedure Rules by which parties identify and make available documents that are or have been in their control and that meet the tests for disclosure.
A certificate or statement confirming compliance with disclosure obligations where required by the applicable procedural regime.
A formal list or statement of disclosed documents used under certain procedural regimes.
A defined approach governing the scope of disclosure. In the Business and Property Courts of England and Wales, Practice Direction 57AD provides several models for Extended Disclosure.
Under Practice Direction 57AD, a model generally limited to disclosure of known adverse documents, without search-based Extended Disclosure.
Under Practice Direction 57AD, a model generally involving limited disclosure of key documents necessary to understand the issues, with only a limited search.
Under Practice Direction 57AD, a request-led model involving disclosure of particular documents or narrow classes of documents relating to defined Issues for Disclosure.
Under Practice Direction 57AD, a search-based model requiring reasonable and proportionate searches for documents likely to support or adversely affect a party's position on defined Issues for Disclosure.
Under Practice Direction 57AD, the broadest model, used exceptionally, involving wide search-based disclosure where necessary for a fair resolution of the issues.
The document used under Practice Direction 57AD in the Business and Property Courts to identify and discuss disclosure issues, proposed models, searches and related matters.
A formal statement describing the extent of a party's search and disclosure and certifying compliance where required by the procedural rules.
The extent to which information is subject to discovery or disclosure obligations under applicable law and procedure.
The pre-trial process, particularly in US litigation, by which parties obtain information and evidence relevant to claims and defences.
Discovery requests directed at the opposing party’s methods of locating, preserving, searching, reviewing and producing information.
The final lifecycle action applied to records or information, such as deletion, destruction, transfer or permanent archival.
The policy, law, schedule or approval permitting records to be destroyed, deleted, archived or transferred.
A review conducted before records are destroyed, deleted or transferred to confirm that disposition is authorised and no holds apply.
A privacy control or notice associated with California privacy rights allowing consumers to opt out of specified sale or sharing of personal information.
A recorded communication or information item. In legal disclosure, the term can extend beyond paper and files to electronic communications, databases and other recorded information.
A documented plan defining custodians, sources, collection methods, timing, responsibilities, validation and chain-of-custody controls.
A group of related items treated together, commonly an email and its attachments or a parent file and embedded children.
A parent document and its attachments or embedded objects treated as a related unit for review, production or deduplication purposes.
A unique identifier assigned to a document within a processing, review or production system.
The complete set of documents within a defined scope for search, review, sampling or production.
Creating a visual representation of a native electronic file, commonly as TIFF, PDF or another image format.
A system or location used to store, index and manage documents and associated metadata.
The examination of documents to make legal or investigative decisions such as responsiveness, relevance, privilege, confidentiality, issue coding or redaction.
Software used to host, search, analyse, code, redact and manage documents during legal review.
A measure of how closely documents resemble one another based on content, structure, metadata or embeddings.
The process of determining where one document ends and another begins, particularly when scanning or processing paper-derived images.
The defined population of documents under consideration in a matter, search, review or analysis.
Reduction of the total document population through filtering, deduplication, analytics or targeted search before substantive review.
Metadata fields associated with a document as a review or production object, such as custodian, file name, dates, hash, path or coding values.
A logical naming space or administrative boundary in computing, such as an internet domain or network directory domain. Domain information may help identify accounts and infrastructure.
Use of workflow and case-management tools to track privacy requests, deadlines, search tasks, review and responses.
A legal basis allowing certain information or processing to be excluded from a subject-access response under applicable law.
Redaction of information from a subject-access response to protect third-party rights, privilege or other exempt information.
A search undertaken to identify personal data relevant to a data-subject access request.
The process used to receive, verify, search, review, redact, approve and respond to a data-subject access request.
Content that changes based on time, user interaction, underlying data or application state, creating special preservation challenges.
The early analysis of data, facts, scope, risk, volume and likely cost to support case strategy and discovery planning.
Electronic submission, review and processing of legal invoices, commonly using standardised formats and billing rules.
A term commonly used in the UK and other jurisdictions for the electronic handling of disclosure, including identification, preservation, collection, processing, review and production of electronically stored information.
A structured questionnaire used to exchange information about electronic documents, systems, search methods and disclosure issues under certain UK procedural regimes.
The process of identifying, preserving, collecting, processing, reviewing, analysing and producing electronically stored information for litigation, investigations, regulatory matters or other legal purposes.
The process of identifying, preserving, collecting, processing, reviewing, analysing and producing electronically stored information (ESI) for use in legal proceedings, investigations or regulatory matters.
A technical or operational role responsible for configuration, security, permissions and support of eDiscovery systems.
A structured matter or workspace in an eDiscovery system used to manage searches, holds, collections, review sets, exports and permissions.
A professional or platform role responsible for managing eDiscovery cases, workflows, teams and governance.
Software supporting one or more eDiscovery stages such as collection, processing, review, analytics, legal hold and production.
An agreement, court order or documented specification governing how ESI will be preserved, searched, collected, reviewed or produced.
The organisational capability to respond efficiently and defensibly to legal or regulatory information demands using prepared policies, data maps, roles, processes and technology.
A controlled collection of items gathered for review and analysis within an eDiscovery platform.
The ordered sequence of activities, decisions and controls used to manage ESI through a legal matter.
The Electronic Discovery Reference Model, a widely used framework describing stages of the eDiscovery lifecycle from information governance through presentation.
A widely adopted conceptual framework describing the major stages of the electronic discovery process from information governance through presentation.
The stage focused on understanding content, context, people, patterns and issues within the information under review.
The stage concerned with acquiring ESI from identified sources in a legally defensible manner.
The stage concerned with locating potential sources of electronically stored information and determining their scope, custodians, systems and relevance.
The stage involving use of evidence in depositions, hearings, trials, regulatory proceedings or other formal presentations.
The stage concerned with protecting potentially relevant ESI from alteration, loss or destruction.
The stage concerned with preparing collected ESI for search, analysis and review, including extraction, filtering and normalisation.
The stage in which responsive information is prepared and delivered in agreed or ordered formats.
The stage in which documents are evaluated for relevance, responsiveness, privilege, confidentiality and other legal or factual issues.
A term used for forensic methods applied to electronic information and digital evidence, often bridging digital forensics, investigations and eDiscovery.
See eDisclosure.
The full name of EDRM, a framework describing the principal stages and activities involved in electronic discovery.
A document in electronic form, including email, messages, word-processing files, databases, spreadsheets and other electronically recorded information.
A questionnaire formerly used under Practice Direction 31B to assist parties in addressing electronic disclosure issues and planning searches.
Information of evidential value that exists in electronic form or is derived from electronic systems, devices or communications.
Production of documents or data in electronic form rather than paper.
A record created, received, stored or maintained in electronic form.
Electronic data attached to or logically associated with other electronic data and used by a person to sign, approve or authenticate a record or transaction.
Information created, stored or communicated electronically and potentially subject to discovery or disclosure, including documents, email, chats, databases, cloud content and multimedia.
A validation method that samples documents predicted as non-relevant to estimate how many relevant documents were missed.
Standardising email addresses or identities so that different forms of the same participant can be analysed consistently.
A system that stores email and associated metadata for retention, search, compliance or recovery.
A file or repository containing multiple email items, such as PST, OST, MBOX or EML collections.
The portion of an email address after the @ symbol, often used to identify organisations, systems or communication patterns.
Structured metadata in an email message describing routing, sender, recipient, identifiers, timestamps and other transmission information.
A unique identifier included in email headers that can assist threading, deduplication and authenticity analysis.
A sequence of related email messages connected through replies, forwards and message-header relationships.
Rebuilding conversational relationships among email messages using headers, content and chronology.
A review technique that suppresses redundant earlier emails when later inclusive messages contain the same content, subject to workflow rules.
An analytics process that groups related email messages into conversational threads and may identify inclusive or most complete messages.
A file stored inside another file or container, such as a spreadsheet embedded in a presentation or an attachment inside an archive.
A file or object inserted within another file (e.g., a spreadsheet inside a Word document).
A numerical vector representation of text, images or other data that captures features or semantic relationships for machine learning and similarity search.
Numerical representations of content designed so that semantically or structurally similar items are located near one another in vector space.
A newer or non-traditional source of discoverable information, such as collaboration platforms, ephemeral messaging, IoT devices or AI systems.
A file format used to store an individual email message, commonly including headers and message content.
The transformation of readable information into protected form using cryptographic methods so that authorised credentials or keys are required to access it.
A user or system device connected to a network, such as a laptop, desktop, server, smartphone or virtual machine.
Collection of data from user or system endpoints, either locally or remotely, using forensic or targeted methods.
Security technology that monitors endpoint activity to detect, investigate and respond to threats. EDR logs can be valuable forensic evidence.
Analysis of people, organisations, locations and other entities and their relationships across a dataset.
Automated identification of people, organisations, locations, dates, amounts or other named entities from text.
The process of determining when different references, names, identifiers or records relate to the same real-world person, organisation or object.
A chat communication configured to disappear automatically after a set time or condition.
Data designed or configured to disappear, expire or become unavailable after a short period or triggering event.
Messaging designed to disappear or become inaccessible after a defined time or event, creating preservation and compliance challenges.
Deletion or removal of personal data or other information, including the privacy right to erasure where applicable.
A measure of incorrect outcomes produced by a process, sample, reviewer or model. The appropriate calculation depends on the task and evaluation design.
An agreement between parties setting out how electronically stored information will be handled, such as formats, custodians, search methods, metadata and production specifications.
A person designated to coordinate technical discovery issues among counsel, clients, vendors or the court.
See eDiscovery Protocol.
Any repository, application, device or system containing potentially relevant electronically stored information.
See Electronic Signatures.
A procedural and technical barrier restricting information flow between specified individuals or teams to manage conflicts or confidentiality.
Regulation (EU) 2016/679 governing protection of personal data in the European Union and European Economic Area.
Linking events from multiple sources based on time, identifiers, users, systems or behaviour.
Retention calculated from a defined event rather than from creation date alone.
Information, testimony, material or records used to establish or challenge facts in a legal, regulatory or investigative process.
A physical or digital control mechanism used to protect, identify and document evidence during custody and transfer.
A person or function responsible for controlling and documenting possession of evidence. This is distinct from an eDiscovery custodian whose information is relevant to a matter.
A cryptographic hash value associated with evidence to support integrity verification.
The assurance that evidence has not been improperly altered and that any changes, handling or transformations are controlled and documented.
A unique number assigned to an individual piece of evidence for tracking and chain-of-custody purposes.
A unique label containing identifiers and case information used to track an evidence item.
A secure physical or digital facility used to store evidence under controlled access and chain-of-custody procedures.
A structured table linking allegations, issues or facts to supporting or contradictory evidence.
Measures taken to prevent relevant evidence from being lost, altered, overwritten or destroyed.
A letter notifying another person or organisation of a duty or request to preserve potentially relevant evidence.
A court or authority order requiring specified evidence or information to be preserved.
A controlled location used to store and manage evidential material, associated metadata and chain-of-custody information.
A physical or digital control used to show whether evidence packaging or storage has been opened or altered.
See Spoliation.
The controlled movement of evidence between people, systems or locations with appropriate documentation and integrity safeguards.
A file or item that is identical to another according to the chosen comparison method, often cryptographic hash value.
A file that could not be processed, indexed, extracted or handled normally and therefore requires separate investigation or treatment.
Management of files that fail processing (password-protected, corrupted, unsupported formats, etc.).
A metadata standard commonly embedded in digital images, potentially containing camera, timestamp, location and technical information.
A filter used to remove data from a processing or review population based on defined criteria.
A list of files, terms, domains or other items intentionally removed from a process or search under defined criteria.
Logs, network records, files or other artefacts indicating that data was transferred out of an environment without authorisation.
Approaches intended to make AI outputs or behaviour understandable to people, including information about factors, evidence or processes influencing an output.
The transfer of data or documents out of a system into files, packages or another platform, often with associated metadata.
The collection of files, metadata, reports and directory structures generated when eDiscovery results are exported.
A report documenting items, counts, errors and settings associated with an eDiscovery export.
Under Practice Direction 57AD in the Business and Property Courts of England and Wales, disclosure beyond Initial Disclosure when ordered under an applicable disclosure model.
One of the disclosure models available under Practice Direction 57AD for disclosure beyond Initial Disclosure.
Making files or data accessible to users outside the originating organisation, commonly through collaboration platforms or cloud links.
The process of retrieving text, metadata, embedded objects or other information from files, applications, databases or devices.
A metric combining precision and recall as their harmonic mean, commonly used to evaluate classification performance.
Search that allows users to refine results by categories or metadata facets such as custodian, date, file type or sender.
The structured capture and analysis of facts, sources, issues, witnesses and chronology within a legal matter.
The degree to which AI-generated statements are supported by accurate facts or reliable sources.
An item incorrectly classified as not belonging to a target category when it actually does, such as a relevant document classified as non-responsive.
An item incorrectly classified as belonging to a target category when it does not, such as a non-responsive document classified as responsive.
See Family-Level Deduplication.
Maintaining or displaying related parent and child items together, such as an email with its attachments.
Deduplication performed so that entire document families (parent plus attachments) are treated as a unit rather than individual files.
A review approach in which decisions consider related document-family members together rather than independently.
An input variable or measurable characteristic used by a machine-learning model.
The process of selecting, transforming or creating input variables that improve model performance or suitability.
A data source queried or accessed through a federated search or integration without necessarily copying all content into a central repository.
A machine-learning approach in which models are trained across distributed data sources without centralising all raw data.
Search performed across multiple repositories or systems while presenting combined or coordinated results.
Providing a generative AI model with a small number of examples in the prompt to guide the desired behaviour or output format.
Search restricted to specified metadata or database fields, such as sender, date, custodian or file type.
A file-system timestamp indicating when a file was accessed, subject to operating-system and file-system behaviour and therefore requiring careful forensic interpretation.
A forensic technique for recovering files or file fragments from raw data based on content signatures or structures rather than relying only on file-system metadata.
A timestamp recording file creation within a particular file system or environment; it may not necessarily represent the original creation of the content.
A suffix in a file name indicating or suggesting file type, such as .docx, .pdf or .pst. Extensions can be misleading if changed or incorrect.
A cryptographic hash value calculated from file content and commonly used for integrity verification or duplicate identification.
A timestamp indicating when file content or certain attributes were last modified, subject to system-specific behaviour.
The logical location of a file within a file system or directory structure.
A structured classification scheme linking records categories to retention, ownership and disposition rules.
A characteristic sequence of bytes used to identify a file type independently of its filename extension.
Comparison of file contents with expected format signatures to identify true file type or detect extension mismatches.
Unused space between the logical end of file content and the end of the allocated storage unit, potentially containing residual data.
The structure and mechanisms an operating system uses to organise, name, store and retrieve files and directories.
Information maintained by a file system about files or directories, such as names, paths, timestamps, size and permissions.
Acquisition of a broader representation of the mobile file system, potentially including application data and system artefacts.
A rule or criterion used to include or exclude data from a search, collection, processing set or review population.
A condition in which automated selection repeatedly exposes a user or system to a narrowed subset of information, potentially affecting analysis or decision-making.
Narrowing a data set by criteria such as date, custodian, file type, path or keyword before or during processing.
Additional training of a pre-trained model on selected data to adapt it to a task, domain or behaviour.
Use of identifying characteristics, hashes or signatures to recognise data, files, devices or content.
The controlled acquisition of digital data using methods intended to preserve evidential integrity and enable repeatable examination.
Systematic examination and interpretation of digital evidence to answer investigative or legal questions.
A trace or data object created by system or user activity that can assist forensic analysis, such as logs, registry entries, browser records or application files.
Collection performed using forensic principles and tools to preserve relevant data, metadata, integrity and documentation.
A structured file format used to store acquired forensic data and associated metadata, such as E01 or AFF formats.
An accurate digital reproduction of information from an electronic device or medium whose integrity is verified, typically using accepted cryptographic hashing.
A validated reproduction of digital evidence intended to accurately preserve the information in the source.
A trained professional who acquires, examines, analyses and reports on digital evidence using forensic methods.
Comparison of cryptographic hash values to confirm that a forensic copy matches the acquired source within the hashed scope.
A forensic copy of a device, medium or defined storage source, often created bit-for-bit and verified by hash.
Contemporaneous records documenting forensic actions, observations, tools, settings and findings.
Preservation using forensic methods designed to maintain integrity, metadata and evidential traceability.
An organisation's preparedness to identify, preserve and use digital evidence efficiently when incidents or disputes arise.
A formal document describing the scope, methods, findings, limitations and conclusions of a forensic examination.
A commonly used concept describing methods that minimise alteration of source evidence, preserve integrity and allow actions to be explained and, where possible, reproduced.
A chronology derived from digital artefacts such as file events, logs, messages and system activity.
Testing and documenting whether a forensic tool performs its intended functions reliably under defined conditions.
A dataset with known characteristics used to test forensic tools, methods or analytical workflows.
A copy of evidence used for examination so the preserved master evidence remains unchanged.
A process or method that preserves the integrity, authenticity and reliability of digital evidence so that it can be relied upon in legal proceedings.
The Federal Rules of Civil Procedure governing civil litigation in US federal courts, including provisions relevant to discovery and electronically stored information.
The rules governing civil litigation in United States federal courts, including provisions that address the discovery of electronically stored information.
The Federal Rule of Civil Procedure governing pretrial conferences, scheduling and case management, which may include discovery and ESI issues.
The Federal Rule of Civil Procedure addressing, among other matters, discovery scope, proportionality, required disclosures and discovery planning.
The Federal Rule of Civil Procedure governing interrogatories to parties.
The Federal Rule of Civil Procedure governing requests for production and inspection of documents, ESI and tangible things.
The Federal Rule of Civil Procedure addressing failures to make disclosures or cooperate in discovery and related remedies or sanctions.
The Federal Rule of Civil Procedure governing subpoenas, including subpoenas seeking documents and ESI from non-parties.
Unstructured text entered without fixed categories or codes, such as notes or narrative fields.
A searchable index built from the textual content of documents and, often, selected metadata fields.
Similarity hashing techniques used to identify files or data that are related or substantially similar rather than exactly identical.
Search techniques that retrieve approximate matches, accounting for spelling variations or OCR errors.
The European Union regulation governing the processing of personal data, which significantly affects cross-border discovery, data minimisation and lawful bases for processing.
Providing generative AI with trusted source material so outputs can be tied more closely to specified evidence or knowledge.
Use of generative AI to assist with document review tasks such as classification, summarisation, issue spotting or explanation, with validation and human oversight.
Use of generative AI to create concise representations of longer documents, conversations or evidence sets, subject to validation.
Artificial-intelligence systems capable of producing new content (text, summaries, classifications, etc.) based on patterns learned from training data; increasingly used in document review and investigation support.
AI capable of generating new content such as text, images, audio, code or structured outputs in response to prompts or other inputs.
Governance specifically addressing generative AI use cases, prompts, confidential data, output validation, security, model selection and human oversight.
Information indicating or inferring the geographic location of a device, person or event.
A deduplication method that removes duplicate copies across the entire data population, typically retaining one representative copy while recording all associated custodians.
An organisation-wide or multi-jurisdictional hold process spanning multiple business units or regions.
The framework of authority, policies, responsibilities, controls and oversight by which an organisation directs and manages an activity or resource.
A defined area of business responsibility used to organise governance policies, ownership, data products or glossary concepts.
A hold scoped narrowly to defined custodians, data sources, dates, subjects or content instead of indiscriminately preserving all information.
A trusted reference set of known or adjudicated outcomes used to train, test or evaluate a model, reviewer or process.
The degree to which an AI output is supported by the source information or context supplied to the system.
Supplying an AI system with trusted source information or context to help constrain or support its outputs.
A rule, filter, technical control or policy constraint intended to limit unsafe, inappropriate or unauthorised AI behaviour.
A fixed-length digital fingerprint generated by a cryptographic algorithm (e.g., SHA-256) from a file or data set; used to verify integrity and identify exact duplicates.
A circumstance in which two different inputs produce the same hash value. Strong modern cryptographic hashes are selected to make collisions extremely unlikely for practical evidence-handling purposes.
Comparing cryptographic hashes to identify exact duplicate files or to verify that an image matches its source.
A collection of hash values used for comparison, such as known system files, known contraband, prior productions or duplicate detection.
A fixed-length value produced by applying a hash function to data. Hashes are widely used to verify integrity and identify exact duplicates.
The process of applying a mathematical hash function to data to generate a fixed-length value used for integrity checks, identification or deduplication.
An out-of-court statement offered for the truth of what it asserts, subject to jurisdiction-specific rules, exceptions and exclusions.
High Efficiency Image Container, a format commonly used by modern mobile devices for photographs and associated metadata.
Information not readily visible through ordinary viewing, such as metadata, hidden rows, deleted content, tracked changes or embedded objects.
A person subject to a legal or investigation hold because they possess or control potentially relevant information.
The written communication issued to custodians instructing them to preserve relevant information.
A final assessment performed before releasing a legal hold to confirm that no other matter or obligation requires continued preservation.
The defined subject matter, time period, people, systems and data categories covered by a preservation hold.
Encryption techniques allowing certain computations to be performed on encrypted data without first decrypting it.
Deduplication performed across an entire case or data set rather than within a single custodian’s data.
Data stored within a third-party or cloud-hosted platform for processing, review or other managed services.
Electronically stored information maintained in a hosted eDiscovery or review environment.
Document review conducted in a centrally hosted or cloud-based review environment where users search, analyse, code and manage documents.
Meaningful human supervision of automated or AI-supported processes, including review, escalation, approval and intervention when necessary.
Evaluation performed by people rather than automated systems, often used for legal judgement, validation or quality assurance.
A design approach in which human reviewers remain actively involved in training, validating or correcting AI/machine-learning outputs.
A system design in which human judgement, approval, review or feedback is integrated into automated or AI-supported decision processes.
Search combining lexical methods such as keywords with semantic or vector retrieval.
A reference in digital content that points to another location, file, webpage or resource. Modern collaboration data may use links instead of traditional attachments.
A file referenced through a link in a message or collaboration platform rather than stored as a static attachment. Its content may change independently of the message.
The process of locating potentially relevant people, systems, repositories, devices and data sources and determining what information may require preservation or collection.
The stage of locating potential sources of relevant ESI and the custodians or systems that hold them.
A digital file containing a visual image, such as JPEG, PNG, TIFF or HEIC.
Production of documents as static images (typically TIFF or PDF) rather than native files.
A document production delivered primarily as page images, commonly TIFF or PDF, usually with text and metadata.
Redaction applied to a rendered image or page representation rather than directly to the native source file.
A PDF whose pages are stored primarily as images rather than searchable text and may therefore require OCR.
The process of creating a forensic image of a storage device; also called acquisition.
International Mobile Equipment Identity, an identifier associated with mobile devices and useful in device identification.
Storage configured so that data cannot be modified or deleted for a defined period or under defined controls.
International Mobile Subscriber Identity, an identifier associated with a mobile subscriber identity module and cellular account.
A platform control that preserves content in its original service or repository despite user deletion or modification, subject to system behaviour.
Electronically stored information that is not reasonably accessible because of undue burden, cost or technical difficulty under the relevant legal standard.
Analysis of incident-related logs, files, messages or compromised datasets to identify affected systems, information and individuals.
The coordinated handling of a security incident from preparation and detection through containment, eradication, recovery and lessons learned.
In email-threading analytics, a message containing unique content that is not fully duplicated in later messages in the same thread and may therefore be prioritised for review.
A data structure created to enable efficient search and retrieval of content, metadata or terms.
Search performed against an index rather than by scanning source files in real time.
A document or data item successfully processed into a searchable index.
The process of analysing data and creating searchable representations of its content and metadata.
Technical artefacts or patterns suggesting that a system or network may have been compromised.
The process of using a trained model to generate predictions, classifications or other outputs from new input data.
A deployed service or interface through which users or applications submit inputs to an AI model and receive outputs.
Information or a collection of information that has value to an organisation and is managed as an asset.
A technical or procedural control preventing specified users or groups from communicating or accessing each other's information.
A structured set of categories used to classify information by sensitivity, business function, record type or handling requirements.
The UK independent regulator responsible for data protection and information-rights law, among other functions.
The authorised destruction or deletion of information when retention and hold requirements have been satisfied.
The coordinated framework for managing information according to business value, legal duties, privacy, security, retention and disposal requirements.
A framework describing stakeholders, responsibilities and processes involved in information governance across legal, records, IT, privacy and business functions.
The sequence of stages through which information passes from creation or receipt to use, storage, retention and final disposition.
Management of information from creation or acquisition through active use, retention, archival and eventual disposal.
A person or function accountable for how a defined information asset is used, protected, retained and governed.
The discipline of finding relevant information from a larger collection, including keyword search, ranking and machine-learning-assisted retrieval.
The risk of harm arising from poor information quality, security, privacy, retention, access or governance.
Protection of information and systems against unauthorised access, disclosure, alteration, destruction or disruption.
Responsible operational management of information assets according to governance, quality, privacy, security and lifecycle requirements.
The process of loading or copying data into an eDiscovery, analytics, storage or review system for further processing.
Under Practice Direction 57AD, limited disclosure generally provided without requiring a search beyond specified boundaries, subject to the rule's terms and exceptions.
In disclosure practice, the opportunity to examine a disclosed document or receive it in an appropriate form, subject to applicable rules and restrictions.
The property that data or evidence is complete and has not been altered improperly or without detection.
The data-protection principle requiring appropriate security of personal data, including protection against unauthorised processing and accidental loss, destruction or damage.
An organisation-led inquiry into suspected misconduct, regulatory issues, fraud, policy breaches, cyber incidents or other concerns.
Transfer of personal data or other regulated information across national or legal-jurisdiction boundaries.
A UK contractual safeguard used for certain restricted transfers of personal data outside the UK.
Network-connected physical devices and sensors that generate or exchange data and may become relevant sources of digital evidence.
A written question served on another party in US civil discovery that must be answered in accordance with applicable procedural rules.
A preservation instruction issued because of an internal, regulatory or other investigation, even where litigation is not yet anticipated.
A controlled system or case area used to collect, organise, analyse and document investigative information.
A network address assigned to a device or connection, used in communications and often relevant to attribution and log analysis.
An international standard specifying competence requirements for testing and calibration laboratories, relevant to some forensic laboratories.
The international standard series addressing electronic discovery, including concepts, governance, processes and guidance related to ESI.
Assigning labels to documents based on factual or legal issues relevant to the matter.
Under Practice Direction 57AD, an issue identified for the purpose of determining the scope of Extended Disclosure.
A structured list of factual, legal or disclosure issues used to guide search, review or case management.
The agreed or court-approved list of Issues for Disclosure used to frame requests, searches and disclosure models under Practice Direction 57AD.
A table or database mapping legal or factual issues to evidence, witnesses, arguments or tasks.
A review label used to identify documents associated with a particular allegation, topic, event or legal issue.
An attempt to induce an AI system to ignore or circumvent its safety policies, constraints or intended instructions.
Two or more controllers that jointly determine the purposes and means of processing personal data.
A file or log recording transactions or changes, potentially useful in reconstructing system or application activity.
JavaScript Object Notation, a structured text format commonly used by APIs and modern applications to exchange data.
A Windows artefact recording recent or frequent files and tasks associated with applications, potentially useful in user-activity analysis.
Techniques or tools that suggest related terms, synonyms or variations to improve search recall.
A document or item matching a specified keyword, phrase or search expression.
A report showing the number and distribution of results matching defined search terms.
Search that identifies items containing specified words, phrases, patterns or combinations of terms.
A structured collection of information used to support search, decision-making, user guidance or AI retrieval.
A structured representation of entities and relationships that can support investigation, search, reasoning and data intelligence.
The systematic capture, organisation, sharing and reuse of organisational knowledge to improve consistency, quality and efficiency.
The process of finding relevant facts, documents or data from a knowledge source for user queries or AI workflows.
A validation test using input with a known expected result to confirm a tool or process behaves correctly.
A file whose identity or category is established in advance, often through a known hash set.
A machine-learning model trained on large text or multimodal datasets to process and generate language and related outputs.
A mathematical technique for representing relationships between terms and documents in a lower-dimensional semantic space.
A legally recognised justification for processing personal data under applicable data-protection law.
A core UK/EU data-protection principle requiring personal data to be processed lawfully, fairly and transparently.
Legal Electronic Data Exchange Standard formats used for electronic legal billing and related data exchange.
Information stored in older systems, media or formats that may require specialised restoration or migration.
An older application, platform or technology still containing business or evidential data but potentially difficult to search or export.
Use of technology to automate legal tasks, workflows, documents, approvals or data processing.
The application of analytics and AI to legal datasets to identify facts, relationships, patterns, timelines and other decision-relevant insights.
A centralised repository used to hold large volumes of legal, investigation or eDiscovery data for analytics and reuse.
Migration of legal, matter or eDiscovery data between systems with attention to chain of custody, metadata, security and usability.
A professional applying statistics, machine learning and data analysis to legal or litigation datasets.
A professional combining legal knowledge with process, data or technical skills to design and improve legal systems, workflows or AI-enabled services.
A formal process used to notify people or functions that potentially relevant information must be preserved because of anticipated or ongoing litigation, investigation or regulatory action.
A directive suspending normal document destruction and requiring the preservation of potentially relevant information when litigation, investigation or regulatory action is reasonably anticipated.
A custodian's confirmation that a legal hold notice has been received, understood and will be followed.
The process of addressing non-response, non-compliance or uncertainty involving custodians or preservation actions through higher levels of legal or management oversight.
A discussion with a custodian or business function to clarify relevant information sources, preservation actions and compliance.
A communication directing recipients to preserve identified categories of potentially relevant information and explaining their obligations.
A control ensuring a legal hold takes precedence over ordinary retention or deletion rules.
Formal notification that a preservation obligation under a specific legal hold has ended, subject to any other applicable requirements.
A periodic communication reinforcing an ongoing preservation obligation and prompting custodians to confirm continued compliance.
Technology used to manage legal holds, custodians, notices, acknowledgements, reminders, questionnaires, reporting and audit trails.
Ongoing monitoring of hold recipients, acknowledgements, reminders, data sources, preservation actions and release status.
The process of capturing initial information about a new legal matter, request, complaint or investigation and routing it for assessment.
The systematic capture, organisation, sharing and reuse of legal know-how, precedents, research and institutional knowledge.
The application of business, process, technology, data and management disciplines to improve the delivery of legal services.
Software used to manage legal work, spend, matters, vendors, contracts, knowledge or workflow.
A legal protection that may shield certain confidential communications or work product from disclosure, depending on jurisdiction and circumstances.
Outsourcing legal or legal-support activities to external providers, including document review and some eDiscovery functions.
A term used in UK and other jurisdictions for privilege protecting specified lawyer-client communications or litigation-related material, subject to applicable law.
Application of project-management methods to legal matters, including scope, schedules, resources, budgets, risk and communication.
The structured tracking and handling of requests for legal support, advice, documents or approvals.
AI systems used to assist with searching, summarising or analysing legal authorities, subject to verification and professional responsibility.
The people, process, technology and operating model through which legal work is provided to clients or organisations.
Processes and technology used to budget, track, analyse and control legal expenditure.
Technology used to support, automate or improve legal work, legal services, compliance, investigations, litigation and legal operations.
Technology used to support the delivery of legal services, including eDiscovery platforms, contract tools, research systems and practice-management software.
A lawful basis for processing personal data where the controller or a third party has a legitimate interest and the processing is necessary and not overridden by the individual's rights and interests, subject to applicable law.
A documented assessment balancing the purpose, necessity and impact of processing relied upon under the legitimate-interests lawful basis.
Traditional document-by-document review without prioritisation by analytics or TAR.
Use of data analysis to understand litigation trends, courts, judges, parties, outcomes or case characteristics.
Information created, collected or managed for litigation, including pleadings, evidence, discovery material, transcripts and work product.
A database used to store, search, organise and manage documents, evidence and metadata for litigation.
Another common term for a legal hold, particularly in US practice.
See Legal Hold Notice.
The state of being prepared to identify, preserve, collect and manage potentially relevant information efficiently when litigation or investigation arises.
Professional and technical support for litigation, including document management, eDiscovery, databases, evidence preparation and trial technology.
A professional who supports legal teams with document databases, eDiscovery, case technology, evidence organisation and trial preparation.
A database designed for coding, searching and managing litigation documents and associated metadata.
Technology used to support litigation processes such as eDiscovery, case management, evidence presentation, transcription and trial support.
Testing a large language model or LLM-enabled workflow for task performance, factuality, safety, consistency and operational suitability.
A hallucinated or unsupported output generated by a large language model.
A structured file used to transfer metadata, text locations, image references or other information into or between eDiscovery review systems.
Standardised structures (e.g., Opticon, Concordance DAT/CSV) used to transfer document metadata and relationships.
A technical specification defining required metadata fields, delimiters, file paths, image references and encoding for a production load file.
Collection performed directly from a device or environment to which the collector has local access.
A repository physically or logically located on a user device or local network rather than a cloud service.
Combining and comparing logs from multiple systems to identify related events, sequences or anomalies.
A file or record containing chronological system, application, user or security events.
The period and policy under which logs are retained before deletion or archival.
Collection of selected files, folders or application data through the logical file system or application layer rather than acquisition of every underlying storage sector.
Acquisition of data exposed through the device operating system or authorised interfaces rather than raw physical storage.
Compression that allows the original data to be reconstructed exactly.
Compression that reduces file size by discarding some information, potentially affecting forensic or evidential detail.
Legal automation built using configurable tools with limited custom programming.
A network-interface identifier that can assist device attribution, though it may be randomised or spoofed.
A subset of AI in which algorithms improve performance on a task through experience (training data) rather than explicit programming; foundational to TAR and predictive coding.
A branch of AI in which systems learn patterns from data to make predictions, classifications or other outputs without being explicitly programmed for every case.
Automated translation of text or speech between languages using computational models.
A data format structured so software can process its contents without requiring human interpretation of presentation alone.
A repository associated with an email account containing messages, folders, calendar items and other communication data.
Preserving email and related mailbox content against deletion or modification for legal, regulatory or investigative purposes.
Software intentionally designed to disrupt, damage, spy on, gain unauthorised access to or otherwise misuse systems and data.
Document review performed by a specialised team (often outsourced) under defined protocols, quality controls and project management.
Document review performed primarily through human examination without machine-learning classification determining review decisions.
The preserved reference copy of acquired digital evidence, maintained under controlled access and not ordinarily used for routine analysis.
A legal, regulatory, compliance or investigative case or project managed as a distinct unit of work.
Information associated with a particular legal, compliance or investigative matter.
A legal hold associated with a specific legal matter and scoped to relevant custodians and data sources.
Software used to manage legal matters, deadlines, participants, budgets, documents and reporting.
A digital workspace containing matter documents, communications, tasks and structured information.
Access controls designed around individual legal matters so users only see information for matters they are authorised to access.
Cryptographic hash algorithms commonly used for integrity verification; SHA-256 is currently preferred for forensic work.
The process of removing data from storage media so that recovery is infeasible at the required security level.
A required or agreed discussion between parties or counsel about discovery issues such as scope, preservation, custodians, search methods, formats and ESI protocols.
The acquisition and analysis of volatile memory to identify running processes, network connections, credentials, malware artefacts and other transient information.
Records showing changes made to chat or collaboration messages over time, where the platform retains such history.
Metadata associated with a communication, such as sender, recipients, timestamps, channel, thread identifiers, reactions and message IDs.
A reaction or emoji response attached to a chat or collaboration message and potentially relevant to meaning or context.
A policy or system setting governing how long messaging content is retained before deletion or archival.
A linked sequence of messages, replies or comments within a messaging or collaboration system.
Data describing other data, such as file properties, timestamps, authorship, routing information, paths, message identifiers or system attributes.
A defined attribute used to store a particular type of metadata, such as author, sent date, file path or hash value.
The process of aligning metadata fields between source systems, processing tools, review platforms or production specifications.
Testing to demonstrate that a method is suitable and reliable for its intended forensic or analytical purpose.
Microsoft's eDiscovery capabilities for identifying, preserving, searching, collecting, reviewing and exporting content from Microsoft 365 environments.
Microsoft's family of data governance, compliance, risk and eDiscovery capabilities used to manage information across Microsoft environments.
Checks confirming that migrated data is complete, accurate and consistent with the source.
Forensic acquisition and analysis of smartphones, tablets and related mobile data, including messages, apps, media, location information and system artefacts.
Acquisition of data from a mobile device using logical, file-system, physical or cloud-based methods.
The specialised collection and analysis of data from mobile devices, including messages, apps, location data and deleted content.
Documentation summarising an AI model's intended use, characteristics, evaluation, limitations and relevant risks.
An open protocol for connecting AI systems with external tools and data sources through standardised interfaces.
Documentation describing an AI model's purpose, design, training, data, performance, limitations, risks and governance.
A decline or change in model performance over time because the data, environment, behaviour or relationship being modelled has changed.
Monitoring designed to identify changes in model performance or data relationships over time.
Assessment of a model's performance using defined metrics, test data, error analysis and task-specific criteria.
Ongoing observation of model performance, errors, drift and operational behaviour after deployment.
The risk of adverse outcomes arising from errors, limitations, misuse or incorrect assumptions in a model.
Governance processes for identifying, validating, monitoring and controlling risks arising from models.
The degree to which meaningful information about a model's purpose, data, design, limitations and performance is available to relevant users or overseers.
Testing and documenting whether a model performs sufficiently well for its intended purpose and operational context.
Document review involving material in more than one language, potentially using translation, language identification and specialist reviewers.
AI designed to process or generate more than one type of information, such as text, images, audio or video.
A cloud architecture in which multiple customers share underlying infrastructure while logical controls separate their data and services.
A file in the format used by the application that created or ordinarily opens it, preserving features and metadata that may be lost in an image conversion.
Redaction performed while preserving or modifying the native-format document, often requiring specialised tools and validation.
A file in the original application format in which it was created (e.g., .docx, .msg), as opposed to an imaged or converted version.
Production of documents in their native application format, usually accompanied by specified metadata and, where needed, placeholder images or text.
An analytics technique used to identify documents that are highly similar but not exactly identical.
A production format that preserves most usability of the original while not being the exact native file (e.g., PDF with extracted text).
A representation that preserves much of a source file's content or structure without being the exact original native format.
Data stored in a form that is not immediately active but can be restored or accessed with moderate effort.
A machine-learning model composed of interconnected computational units arranged in layers and trained to learn patterns from data.
Search using neural-network representations, often embeddings, to retrieve semantically related content.
Reference data sets of known software and system files used for de-NISTing.
Legal workflow or document automation created through visual configuration rather than traditional programming.
Data or information that is irrelevant, erroneous or obscures useful signals in analysis.
Potentially relevant information associated with shared systems or repositories rather than a single individual.
A potentially relevant repository not primarily associated with one individual custodian, such as a shared drive, database, collaboration site or departmental mailbox.
Discovery directed to a person or organisation that is not a party to the litigation, commonly through subpoena or equivalent procedure.
Transforming data into a more consistent form for processing, searching, comparison or review while preserving necessary source information.
A legal concept used for ESI whose retrieval would impose undue burden or cost, potentially affecting discovery obligations under applicable procedural rules.
A formal or informal communication placing recipients on notice that information must be preserved.
A communication requesting or requiring preservation of specified evidence or data.
Review and classification of breached data to determine notification obligations and affected individuals.
A storage architecture that manages data as objects with associated identifiers and metadata, commonly used in cloud environments.
Access controls applied to individual data objects, documents or records rather than only to an entire system.
Coding of factual, non-interpretive attributes such as date, author and document type.
Technology that converts images of printed or handwritten text into machine-readable text for searching, analysis or accessibility.
Data not readily accessible through live systems, such as removable media, offline archives or disconnected storage.
Microsoft's cloud file-storage and collaboration service, frequently relevant as a source of user files in Microsoft 365 eDiscovery.
A formal representation of concepts, categories and relationships within a domain of knowledge.
Information collected from publicly available sources and analysed for investigative, security, legal or intelligence purposes.
A privacy right allowing individuals to object to or opt out of specified data uses, such as certain sale, sharing or targeted advertising practices, depending on jurisdiction.
See OCR.
The principle of collecting the most volatile evidence (e.g., RAM) before less volatile sources (e.g., disk, backups).
The source evidence or original item from which copies, images or derivative working copies may be made, subject to the applicable evidential context.
An Offline Storage Table file used by Microsoft Outlook to cache mailbox data locally.
Processes and systems used to instruct, monitor and evaluate external legal providers.
Collection of substantially more data than is reasonably necessary for the defined legal or investigative purpose.
A machine-learning problem in which a model fits training data too closely and performs poorly on new or unseen data.
Discovery or disclosure centred on physical paper documents rather than electronically stored information.
The primary item to which one or more child items are related, such as an email containing attachments.
The link between a container or email and its attachments or embedded objects.
Analysing data or a file format to identify its structure and extract meaningful fields, text or objects.
A data item that cannot be fully indexed because of encryption, corruption, unsupported format or other limitations and may require special handling.
Independent review of forensic methods, analysis or conclusions by another qualified professional.
A record designated for indefinite or archival retention because of legal, historical or institutional value.
Information relating to an identified or identifiable individual under applicable data-protection law.
A breach of security leading to accidental or unlawful destruction, loss, alteration, unauthorised disclosure of or access to personal data.
A record of categories, locations, purposes, owners, recipients and retention associated with personal data processing.
Data that can identify an individual; definitions vary by jurisdiction and are generally narrower than GDPR “personal data”.
A broad term for information about an individual; the precise legal meaning varies by jurisdiction.
A management-system framework for governing privacy and personal information, often associated with ISO/IEC 27701.
Information that identifies or can be linked to a specific person. The precise definition varies by jurisdiction, standard and context.
A social-engineering technique using deceptive messages or websites to induce users to disclose information, execute malware or take other harmful actions.
Acquisition targeting raw physical storage or a near-complete physical representation, where technically possible and lawful.
A document format designed to preserve page appearance across systems. PDFs may contain searchable text, images, metadata, attachments and interactive content.
The Civil Procedure Rules practice direction addressing disclosure of electronic documents in certain English and Welsh proceedings outside the principal PD57AD regime.
The practice direction governing disclosure in the Business and Property Courts of England and Wales, including Initial Disclosure, Issues for Disclosure and Extended Disclosure models.
In information retrieval, the proportion of items identified as relevant that are in fact relevant.
A machine-learning approach used in document review to predict document categories, such as responsiveness, based on human-labelled examples.
Ordering documents by predicted relevance score so that the most likely relevant items are reviewed first.
A model-generated estimate of how likely a document is to be relevant or responsive based on training data and algorithmic analysis.
A Windows artefact that may contain information about application execution and file access, useful in forensic timeline analysis.
The process of protecting potentially relevant information from alteration, deletion, overwriting or loss when a legal, regulatory or investigative duty requires it.
A collection made primarily to preserve data in a stable form rather than immediately process it for review.
A communication directing relevant people or functions to preserve identified information. It may be part of a formal legal hold process.
An event or circumstance that causes a duty or need to preserve potentially relevant information.
Preserving information within the source system rather than creating a separate copy, commonly through platform hold or retention features.
The proportion of relevant documents within a collection; an important factor in TAR design and cost estimation.
An approach in which privacy protections are incorporated into systems, processes and products from the earliest design stages.
The identification and review of personal data across organisational systems for privacy compliance or rights requests.
The application of engineering methods to design, implement and assess systems that meet privacy objectives and requirements.
Technology designed to reduce privacy risk, such as differential privacy, secure multiparty computation, federated learning or privacy-preserving encryption.
A structured assessment of privacy implications and risks associated with a project, system or processing activity; terminology and legal requirements vary by jurisdiction.
Information provided to individuals explaining how and why their personal data is processed, including relevant rights and other required details.
A request by an individual to exercise rights under applicable privacy law, such as access, deletion, correction or objection.
The risk that data processing causes adverse consequences to individuals through inappropriate collection, use, access, disclosure or inference.
A controlled technical environment or set of technologies intended to support data use while reducing privacy exposure.
A legal protection permitting certain communications or material to be withheld from disclosure, such as attorney-client privilege or protected work product, subject to applicable law.
A legal right to withhold certain communications or materials from discovery (attorney-client, work product, etc.).
A determination that a document is privileged or not privileged, usually made during legal review.
The recovery or return of privileged information that was inadvertently disclosed, under applicable law, agreement or court order.
Coding documents to indicate whether privilege applies and the basis or category of privilege.
A list describing documents or communications withheld from production on privilege grounds with sufficient information to support the claim without revealing protected substance.
Redaction used to withhold privileged portions of an otherwise producible document.
The review process used to identify, confirm and manage potentially privileged information before production.
A workflow or technical control separating potentially privileged material from users who should not access it.
A designated group responsible for reviewing potentially privileged material, often separated from the main review team.
A report generated by an eDiscovery platform describing what occurred during a search, collection, ingestion, export or other process, including counts and errors.
The eDiscovery stage in which collected data is prepared for search and review through extraction, filtering, indexing, deduplication, conversion and exception handling.
The stage after collection in which data is filtered, deduplicated, text-extracted, imaged and prepared for review and analytics.
Software that extracts, expands, normalises, filters and indexes data for eDiscovery or investigation.
A file or data item that cannot be processed normally because of corruption, encryption, unsupported format, password protection or other technical issue.
Files that cannot be fully processed by standard tools and require special handling.
A defined processing operation applied to a dataset or batch in an eDiscovery system.
A set of configured processing options controlling extraction, deduplication, filtering, OCR, time zones and other processing behaviour.
Under data-protection law, a person or organisation that processes personal data on behalf of a controller.
The delivery of responsive documents or ESI to another party, regulator or authority in an agreed or ordered form.
The format in which documents are delivered, such as native, TIFF, PDF, text and metadata load files.
A list or database of documents included in a production, commonly recording Bates ranges, filenames, custodians or other metadata.
A structured inventory of files and records included in a production or data delivery.
An agreement or specification governing how responsive material will be prepared and delivered.
The defined group of documents and associated files prepared for delivery in a particular production.
Agreed or ordered requirements governing production formats, metadata, numbering, text, images, natives, redactions and delivery structure.
The number or size of documents and associated files included in a production.
Automated processing of personal data to evaluate or predict aspects of an individual, such as behaviour, preferences, performance, reliability or location.
The planning, coordination, budgeting, quality control and stakeholder management of eDiscovery workflows.
Input instructions, context or examples supplied to a generative AI system to influence the output it produces.
The structured design, testing and refinement of prompts and workflows to obtain more reliable, useful and controlled outputs from generative AI.
An attack or manipulation in which malicious or untrusted input attempts to override, alter or exploit an AI system's instructions or tool use.
Unintended disclosure of hidden instructions, confidential context or system information through an AI interaction.
A governed collection of reusable prompts, templates and instructions for approved AI use cases.
Tracking changes to prompts so that AI workflows can be reproduced, audited and compared over time.
The principle that the scope, cost and burden of discovery or disclosure should be reasonable in relation to the needs, value and circumstances of the matter.
The party that issues a discovery request.
A label or marking applied to information to indicate confidentiality, sensitivity or handling restrictions.
A court order restricting how discovered or disclosed information may be used, shared, handled or protected.
Information about the origin, history, custody and transformations of data or evidence.
Metadata recording the origin, custody, transformation or history of data, documents or AI outputs.
A search requiring terms to appear within a specified distance of each other.
Processing personal data so that it cannot be attributed to a specific person without additional information kept separately and protected by appropriate controls.
Microsoft Outlook personal storage and offline storage file formats commonly encountered in collections.
A Personal Storage Table file used by Microsoft Outlook to store email, calendar and other mailbox data.
The data-protection principle that personal data should be collected for specified, explicit and legitimate purposes and not used incompatibly with those purposes.
Planned processes intended to provide confidence that a workflow, system or output meets defined quality requirements.
Operational checks used to detect and correct errors or inconsistencies in data, review coding, productions or other work products.
A documented system of policies, procedures, responsibilities and controls used to achieve and demonstrate consistent quality.
An attribute that may not directly identify a person alone but can contribute to identification when combined with other information.
A structured request used to retrieve or analyse information from a database, search index, application or AI system.
An AI or search capability that generates or retrieves answers to user questions from data or knowledge sources.
An arrangement allowing limited review of another party's material before privilege screening, typically with protections against waiver where legally permitted.
An AI architecture that retrieves relevant source information and supplies it to a generative model to help produce better-grounded answers.
The process of linking de-identified or pseudonymised data back to an identifiable individual.
A search considered proportionate and appropriate to the issues, sources and circumstances under the applicable disclosure or discovery regime.
Actions considered reasonable under applicable legal standards to meet duties such as preservation, disclosure or compliance, assessed in context.
In information retrieval, the proportion of all truly relevant items that the process successfully identifies.
The inverse relationship often observed between maximising recall and maximising precision in search or TAR.
The process of comparing source and destination counts, identifiers, volumes or records to confirm completeness and explain differences.
Information created, received and maintained as evidence of business, legal or organisational activity.
The authoritative version of a record designated for official retention.
A formal record describing personal-data processing activities, often required under data-protection law for relevant controllers and processors.
A group of related records managed together because they arise from the same function, process or retention requirement.
Assignment of records to categories based on function, subject, legal requirement or retention rule.
The act of designating information as an official record subject to records-management controls.
The authorised destruction, deletion, transfer or archival of records when retention requirements are satisfied.
A temporary suspension of scheduled disposition, commonly because of legal, regulatory or investigative requirements.
The systematic control of records throughout their lifecycle, including creation, classification, retention, access, storage and disposition.
A schedule defining retention periods and final disposition actions for classes of records.
Metadata and residual files associated with deleted items placed in a system recycle bin or trash location.
The removal or obscuring of information from a document before disclosure, production or publication, commonly to protect privilege, privacy, confidentiality or protected information.
A record describing redactions applied to documents, often including document identifier, page, reason and legal basis.
A logical section of the Windows Registry stored in one or more files and commonly examined in digital forensics.
A preservation hold issued in response to actual or anticipated regulatory inquiry, examination or enforcement activity.
An investigation conducted by or in response to a regulator concerning potential breaches of law, rules or standards.
A database organising data into related tables that can be queried using structured relationships.
A widely used eDiscovery platform supporting review, analytics, workflow management and production. It is a product name rather than a generic eDiscovery term.
See Legal Hold Release.
The relationship between information and the factual or legal issues under consideration. In discovery, relevance is applied according to the governing procedural standard.
Documents or data that are relevant to the claims or defences in a matter or that respond to a discovery request.
A document bearing on a matter, issue or request under the applicable legal or investigative standard.
Corrective action taken to address a legal, technical, security, privacy or process deficiency identified during investigation or review.
Collection performed from a device or system over a network without the collector being physically present at the source.
Forensic collection performed over a network from a device or system while preserving evidential integrity and documentation.
A generated representation of a native file, such as a TIFF or PDF image used for review or production.
The ability to obtain consistent results when the same method is performed under the same conditions.
A location or system used to store and manage data or documents.
A central store of collected or processed ESI used for review and analysis.
Documenting the relationship between business activities, custodians and the repositories in which their information is stored.
The ability for independent practitioners or systems to obtain consistent results under appropriately similar conditions using documented methods.
A US civil discovery request asking another party to admit or deny specified facts or the authenticity of documents.
A formal discovery request seeking documents, ESI or tangible items within the scope permitted by procedural rules.
A retrieval step that reorders initial search results using an additional relevance model or scoring method.
The party that receives and must respond to a discovery request.
A document that falls within the scope of a discovery request, disclosure issue, investigation criterion or other defined requirement.
Electronically stored information that falls within the scope of a discovery request, disclosure issue or other defined requirement.
The determination of whether a document or item satisfies the criteria of a request, issue or production obligation.
A privacy right or processing state in which personal data is retained but its use is limited while specified conditions apply.
Continued storage of information for a defined period or until specified conditions are met.
A situation in which two or more legal, regulatory, business or hold requirements prescribe different retention outcomes for the same information.
A justified departure from normal retention or disposal rules, such as legal hold, regulatory requirement or business need.
A metadata label used by information-governance systems to apply retention or disposition rules to content.
An organisation’s rules governing how long different categories of information are kept and when they may be disposed of.
A structured schedule specifying how long categories of records or information should be retained and what should happen at the end of the period.
The event from which a retention period begins, such as contract termination, employee departure or case closure.
An architecture in which relevant external information is retrieved and supplied to a generative model to support more grounded and context-specific outputs.
The eDiscovery stage in which documents are examined and coded for responsiveness, relevance, privilege, confidentiality, issues or other matter-specific criteria.
The examination of documents for relevance, privilege, confidentiality and other coding decisions, performed by humans, AI systems or a combination of both.
Analytics used to support document review, such as threading, clustering, prioritisation, similarity analysis and reviewer metrics.
A discrete set of documents assigned to a reviewer for coding within a managed review.
Defined conditions used to determine when a review can reasonably end, including coverage, QC, validation, issue closure or production readiness.
A reviewer's determination about a document, such as responsive, non-responsive, privileged, confidential or issue-related.
Written instructions explaining legal issues, coding criteria, examples, escalation rules and QC expectations for reviewers.
Software used to host, search, analyse, code and produce ESI during the review stage of eDiscovery.
The set of documents made available for review after collection, processing and culling.
Written instructions defining how reviewers should assess and code documents, including scope, issues, privilege, escalation and QC.
The number of documents or pages reviewed per unit of time, often used for planning but influenced heavily by complexity and workflow.
A defined population of documents assembled for examination, coding and analysis within a review platform.
The process of copying collected data into a review set and preparing it for search and analysis.
A query executed within a defined review set rather than against the original enterprise data sources.
The sequence of review activities, including batching, coding, escalation, privilege, QC, redaction and completion.
Digital content combining or using media such as audio, video, animation, interactive elements or complex web content.
A privacy right allowing individuals to obtain confirmation and access to personal data processed about them, subject to applicable exemptions and procedural rules.
A privacy right allowing individuals in qualifying circumstances to receive certain personal data in a structured, commonly used, machine-readable format and transmit it to another controller.
A privacy right allowing individuals in qualifying circumstances to request deletion of personal data, subject to legal exceptions.
A privacy right allowing individuals in specified circumstances to object to certain processing of their personal data.
A privacy right allowing individuals to seek correction of inaccurate personal data and completion of incomplete data where appropriate.
A privacy right allowing individuals in specified circumstances to require limitation of personal-data processing.
The ability of a system or model to maintain acceptable performance when inputs, conditions or environments vary.
Techniques for managing versions or restoring previous states of data or coding decisions.
Collection performed in stages over time as additional custodians, systems or periods become relevant.
Production delivered in multiple tranches rather than one final delivery.
The US federal civil-procedure conference in which parties discuss discovery planning, including issues concerning ESI and preservation.
The Federal Rule of Civil Procedure governing requests for production of documents, electronically stored information and tangible things, among other matters.
The Federal Rule of Civil Procedure addressing failure to preserve electronically stored information and potential remedial measures or sanctions.
The selection of a subset of items from a population for estimation, validation, quality control or investigation.
Using statistical samples to estimate the accuracy of a TAR model or review process.
A formal description of the structure, fields, relationships or allowed values within a dataset or database.
The process of retrieving information that matches defined criteria from a data source, index or document collection.
An item returned because it matches a search query or search condition.
A preliminary view of search results used to assess query effectiveness before export or collection.
A subset of search results presented for inspection or validation rather than the complete result population.
Metrics describing search results, such as hit counts, locations, file types, custodians or estimated volume.
The rules and operators used to construct valid queries in a search system.
A word, phrase, expression or query element used to identify potentially relevant information.
A report showing search terms, hit counts, unique hits, overlap or other statistics used to evaluate proposed searches.
Testing whether a search strategy retrieves expected responsive material and avoids unreasonable over- or under-inclusion.
A PDF containing machine-readable text, either natively or through OCR, enabling text search.
Deletion intended to make data unrecoverable or impracticable to recover using appropriate technical methods.
Controlled transmission of files using encryption, authentication and other safeguards to protect confidentiality and integrity.
Cryptographic techniques allowing multiple parties to jointly compute results while limiting disclosure of their private inputs.
A label indicating the sensitivity and handling requirements of information.
Microsoft's AI-assisted security capability; where used with compliance or eDiscovery workflows, outputs should be governed and validated like other AI-assisted work.
Technology that collects, correlates and analyses security logs and events across systems to support monitoring, detection and investigation.
A document selected as an example for training, searching, similarity analysis or workflow calibration.
A group of human-reviewed documents used to begin training or calibrating a predictive coding or machine-learning model.
Search based on meaning and contextual similarity rather than exact lexical matching alone.
The degree to which two pieces of content express related meaning, regardless of exact wording.
Information requiring heightened protection because of legal, contractual, security or business sensitivity.
A general expression for personal information requiring heightened protection. The precise legal categories vary by jurisdiction; under UK/EU GDPR, 'special category data' has a specific meaning.
A defined category under some privacy laws, including the CCPA/CPRA, covering specified highly sensitive types of personal information.
A label used to classify and protect information according to sensitivity, commonly controlling encryption, access or handling requirements.
Automated analysis intended to classify emotional tone or opinion in text. Its reliability and legal relevance require careful validation.
A computer system or service providing data, applications or functionality to other devices or users.
Microsoft's collaboration and content-management platform, commonly a significant data source in Microsoft 365 eDiscovery.
A logical Microsoft SharePoint workspace containing pages, lists, libraries and files that may be relevant to eDiscovery.
A Windows artefact recording folder-view information that can provide evidence of folder access, including removable or network locations.
A Windows shortcut artefact containing target paths, timestamps and other metadata that may show file or program access.
Forensic examination of subscriber identity modules and associated data such as identifiers, contacts or network information.
A collaboration and messaging platform whose channels, direct messages, threads, reactions and shared files may be relevant to eDiscovery.
Collection and preservation of messages, files and metadata from Slack or similar collaboration platforms.
A shared communication space within Slack containing messages, threads, reactions and files.
A private Slack conversation between two or more users, potentially subject to preservation and collection.
The preservation, collection, search, review and production of potentially relevant Slack data.
Unused space within a file or disk cluster that may contain residual data from previous files; of interest in forensic examinations.
A set of replies associated with a parent Slack message, preserving conversational context.
Potentially evidential content from social networks or online platforms, including posts, messages, comments, media, profiles and associated metadata.
Specialised handling and review of software source code when it is relevant to a dispute.
Data obtained from the original or authoritative source before processing, conversion or transformation.
Preserving data at or near its original source rather than relying only on later exported copies.
The application, device, platform or repository from which data originates.
Maintaining traceability from original source data through processing, review and final production.
Under UK/EU GDPR, personal data revealing specified sensitive characteristics or concerning specified sensitive matters, subject to additional legal protections.
Personal data falling within specified sensitive categories under UK/EU GDPR, such as health, biometric or political-opinion data, subject to additional protections.
Review performed by subject-matter specialists, such as privilege lawyers, foreign-language reviewers, forensic experts or regulatory specialists.
The loss, destruction, alteration or failure to preserve evidence that should have been preserved for litigation or another legal process.
Court-imposed penalties (adverse inference, cost shifting, dismissal, etc.) for the loss or destruction of relevant evidence.
A lightweight relational database format widely used by mobile and desktop applications and frequently encountered in forensic analysis.
European Commission-approved contractual clauses used as a safeguard for certain international transfers of personal data.
A disclosure concept under the Civil Procedure Rules historically associated with disclosure of documents supporting or adversely affecting a party's case; its application depends on the procedural regime governing the case.
A documented procedure describing how a recurring task should be performed consistently and controlled.
Selection and analysis of a sample using statistical methods to estimate characteristics of a larger population.
Techniques for hiding information within other files or media so that the existence of the hidden content is not obvious.
The data-protection principle requiring personal data not to be kept in identifiable form for longer than necessary for the purposes for which it is processed.
Data organised according to a defined schema, such as rows and columns in a relational database.
AI-generated output constrained to a defined schema, format or set of fields to improve consistency and downstream processing.
A language used to define, query and manipulate data in relational databases.
A document review process using defined coding fields, protocols, workflows and quality controls.
Review of material collected for a subject-access request to identify responsive personal data, exemptions, privilege and third-party information.
Coding that requires legal or factual judgment (relevance, privilege, issue tags) rather than purely objective attributes.
A compulsory legal demand requiring a person to attend, testify or produce documents, ESI or other evidence, subject to applicable law.
A processor engaged by another processor to process personal data on behalf of the controller.
Machine learning in which models learn from examples paired with known labels or outcomes.
An independent public authority responsible for monitoring and enforcing data-protection law, such as the UK Information Commissioner.
Additional contractual, technical or organisational safeguards used to support international personal-data transfers where required.
A search configuration treating defined terms as equivalent or related for retrieval purposes.
Artificially generated data designed to resemble characteristics of real data, often used for testing or training while reducing exposure of real information.
Media generated or materially altered by AI or other computational techniques rather than captured directly from the represented event.
Artificially generated data that resembles characteristics of personal data but is not directly copied from real individuals; re-identification risk still requires assessment.
A file primarily used by an operating system or application rather than created as substantive user content.
Operating-system and application files that are typically removed during de-NISTing because they lack user content.
Metadata generated by systems or applications about files, events, accounts, devices or transactions.
The authoritative system designated as the primary source for a defined category of data or records.
A high-priority instruction supplied to an AI model or agent to define behaviour, role, constraints or policy.
A label applied to a document, record or data item to classify it or record a review decision.
An iterative TAR workflow in which humans code successive sample sets to train a model that is then applied to the remaining collection.
A non-iterative or continuously updating TAR approach in which the model improves and re-prioritises documents as human coding progresses.
Statistical and practical methods used to confirm that a technology-assisted review process has achieved acceptable results.
Collection limited to defined data sources, folders, accounts, dates or categories rather than acquisition of an entire device or repository.
A controlled hierarchical classification structure used to organise information, records, knowledge or data.
A Microsoft Teams workspace for group communication containing posts, replies, files and other collaborative content.
A Microsoft Teams private or group chat containing messages, reactions, shared files and related metadata.
The preservation, search, collection, review and production of Microsoft Teams messages, channels, meeting content and associated files.
Audio/video content captured from a Microsoft Teams meeting, potentially accompanied by transcript, attendance and chat data.
Machine-generated or edited text representing spoken content in a Teams meeting, subject to accuracy and preservation considerations.
Review focused on correctness of technical methods, calculations, tools or findings.
A process in which computer technology classifies or prioritises documents based on input from human reviewers to support document review.
An informal industry term commonly used for predictive-coding workflows based on seed or training sets and iterative model training before or during review.
An informal industry term commonly associated with continuous active learning and ongoing reprioritisation as reviewers make decisions.
An agreed or court-ordered description of how TAR will be trained, applied, validated and documented.
An approach focused on principles and outcomes rather than dependence on a particular software product or vendor.
Metadata representing dates, times or sequences associated with files, messages, events or transactions.
A dataset kept separate from training and used to evaluate model performance on unseen examples.
Computational analysis of text to identify terms, entities, topics, sentiment, similarity or other patterns.
The process of obtaining machine-readable text from electronic documents for indexing, search or analysis.
Standardising text representation, such as encoding, whitespace, punctuation or character forms, to improve processing or comparison.
Documents whose extracted text is highly similar even if binary hashes differ.
Excluding non-inclusive email messages from review when an inclusive message already contains the full conversation.
The grouping of related communications into conversational sequences, most commonly email threading.
See Email Threading.
A person, group or system responsible for malicious or unauthorised cyber activity.
Tagged Image File Format, an image format widely used in traditional eDiscovery image productions.
The process of aligning system clocks to a common time source to improve consistency of event records.
Converting timestamps to a consistent time zone so that events from different systems or users can be compared reliably.
The difference between local time and UTC recorded or applied to timestamps.
The organisation and examination of events in chronological order to identify sequences, gaps, relationships and key periods.
Rebuilding the sequence of relevant events from logs, metadata, communications and forensic artefacts.
Intentional or accidental alteration of timestamp information, requiring careful forensic validation.
A unit of text processed by many language models, such as a word, subword or symbol. Token counts affect context limits and processing cost.
The maximum amount of tokenised input and output an AI model can process within a request or context window.
The process of splitting text into units, such as words, subwords or symbols, for search, analysis or language-model processing.
An AI capability that allows a model or agent to invoke external functions, applications or services to perform actions or retrieve information.
Broader testing demonstrating that a software or hardware tool performs its intended functions reliably under defined conditions.
Checks confirming that a specific tool instance is operating as expected before or during use.
The process by which a machine-learning model learns patterns or parameters from data; in professional contexts, the term also refers to developing human competence.
Data used to train or adapt a machine-learning model.
A labelled or unlabelled dataset used to fit a model's parameters or behaviour.
Linking transcript text to corresponding audio or video timecodes for navigation and presentation.
An assessment of risks associated with transferring personal data internationally and whether safeguards provide adequate protection.
A neural-network architecture using attention mechanisms and widely used in modern large language models and other AI systems.
The principle of providing understandable information about processes, decisions, data use or system behaviour. In AI and privacy, transparency supports accountability and informed oversight.
The requirement or expectation that individuals and stakeholders receive clear, accessible information about data processing, systems or decisions.
The organisation and display of documents, images, video and other evidence during hearings or trials using litigation technology.
An item correctly classified as not belonging to the target category.
An item correctly classified as belonging to the target category.
A dataset considered sufficiently reliable, documented and controlled for a defined analytical, evidential or AI purpose.
A protected processing environment designed to isolate code and data from other parts of a system.
The United Kingdom version of the General Data Protection Regulation forming a central part of UK data-protection law.
Data remnants located in storage areas not currently allocated to active files.
Storage space not currently assigned to active files but which may contain remnants of deleted data or prior content.
A document that should have been disclosed under the applicable obligation but was not included in the disclosure provided.
A value assigned to distinguish one item, record, user, device or document from others.
Information that does not follow a fixed tabular schema, such as documents, email, chat, audio and images.
Machine learning that identifies structure or patterns in data without predefined labelled outcomes.
System records indicating connection or use of removable USB devices, including identifiers and timestamps.
A record of user actions within a system, potentially showing logins, searches, downloads, edits or administrative events.
The process of linking digital activity to a particular user, account, device or person using technical and contextual evidence.
Uniform Task-Based Management System codes used to categorise legal work and expenses for billing and analysis.
Coordinated Universal Time, commonly used as a reference time standard when normalising timestamps across systems.
The process of checking whether a method, tool, model, dataset or result performs as intended and meets defined requirements.
A sample selected to test the accuracy, completeness or reliability of a search, review process or model.
A set of documents with known coding used solely to evaluate model performance and not for training.
A numerical representation of data. In modern AI and semantic search, vectors often represent meaning or features in a multidimensional space.
A database optimised to store and search vector representations, commonly used for semantic search and retrieval-augmented generation.
A vector representation of data designed so that semantic or feature similarity can be measured mathematically.
Search that compares numerical embeddings to identify semantically similar items.
A system used to store and retrieve vector embeddings for semantic search, RAG or similarity analysis.
Independent of a specific technology vendor and focused on transferable principles, methods and competencies.
A particular state or revision of a document, file, dataset, model or application.
The systematic management of changes to documents, code, prompts, models or other artefacts so that revisions can be tracked and restored.
A record of earlier states or revisions of a document or item, often stored by collaboration and content-management platforms.
Deduplication performed within a single custodian’s or source’s data set rather than across the entire case.
Video content that may be used to establish facts in a legal, investigative or regulatory matter.
Forensic examination and analysis of video evidence for authenticity, enhancement, timing, content or other investigative questions.
A controlled online repository used to share sensitive documents for transactions, investigations, litigation or due diligence.
A record essential to continued operations, legal rights or recovery after disruption.
Recorded or generated audio containing speech, such as calls, voicemail, meeting recordings or voice messages.
A recorded spoken message in voicemail, chat or collaboration systems and a potential source of discoverable evidence.
Data that exists in temporary storage (especially RAM) and is lost when power is removed; prioritised in order-of-volatility forensic collection.
Visible or hidden information embedded in a document or media item to identify source, ownership, status or distribution.
A preserved collection of web content captured at a point or over a period of time.
A preserved copy or acquisition of web content, potentially including HTML, images, scripts, metadata and screenshots.
Capturing web content and relevant metadata in a stable form so it can be reviewed or used as evidence later.
Email accessed through a web service or browser-based interface, such as Gmail or Outlook on the web.
A hierarchical database used by Microsoft Windows to store system and application configuration, user settings and activity-related artefacts.
A structured collection of information, documents, statements and evidence associated with a witness.
Material prepared by or for lawyers in anticipation of litigation that may receive protection from disclosure under applicable law.
A US legal doctrine protecting qualifying materials prepared in anticipation of litigation from discovery, subject to applicable rules and exceptions.
A protection that shields materials prepared in anticipation of litigation by or for an attorney from discovery by opposing parties.
Use of rules, scripts, APIs or AI to perform or coordinate repetitive tasks across a legal, eDiscovery or governance process.
A defined stream of related tasks within a project, such as collection, review, privilege or production.
A hardware or software control designed to prevent data being written to a source storage device during forensic acquisition or examination.
Controls that prevent data from being altered on a storage device or evidence source.
A database log recording changes before they are committed to the main database, potentially containing recent or deleted information.
Measures that prevent any write operations to storage media during forensic acquisition.
Extensible Markup Language, a structured text format used to represent and exchange hierarchical data.
A security model based on continuously verifying users, devices and access rather than assuming trust based on network location.
Classification performed by an AI model without task-specific labelled examples in the prompt or fine-tuning set.
A compressed archive format capable of containing multiple files and folders. ZIP files are commonly expanded during eDiscovery processing.