Electronic document management system

The cloud-based electronic document management system addresses document management challenges by automating processing and organization, enhancing eDiscovery efficiency and reducing duplication, suitable for legal document workflows.

WO2025179335A1PCT designated stage Publication Date: 2025-09-04LAW BAA PTY LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/AU2025/050166
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-26
Filing Date
2025-02-26
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Managing large volumes of electronic documents is challenging, particularly in legal environments, due to difficulties in organizing, verifying veracity, and reducing duplication, which leads to increased costs and complexity in eDiscovery processes.

Method used

A cloud-based electronic document management system utilizing a coordinated set of microservices and queues to automate document processing, including ingestion, metadata extraction, and formatting, enabling non-technical users to manage and review documents efficiently.

Benefits of technology

Facilitates effective management and organization of electronic documents, reducing duplication and irrelevant data, and simplifying eDiscovery processes, allowing non-technical users to handle legal document workflows without expert assistance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure AU2025050166_04092025_PF_FP_ABST
    Figure AU2025050166_04092025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed herein are an electronic document management system and associated method. The system includes: a co-ordinator; a set of request queues; a set of response queues; and a set of microservices, each of which is configured to perform one of a set of predefined processing tasks on one of either a set of predefined item categories or a set of predefined item subtypes. The request queues and the response queues manage communications between the co-ordinator and each of the microservices. The co-ordinator is configured to: create a project, based on user input; ingest data items by: identifying an item category of the data item, and assigning a set of the processing tasks to be performed on the data item, based on the identified item category; and add the ingested data items to the project.
Need to check novelty before this filing date? Find Prior Art

Description

ELECTRONIC DOCUMENT MANAGEMENT SYSTEMRelated Application

[0001] This application is related to Australian Provisional Patent Application No. 2024900466 titled “ELECTRONIC DOCUMENT MANAGEMENT SYSTEM” and filed 26 February 2024 in the name of Law Baa Pty Ltd, the entire content of which is incorporated by reference as if fully set forth herein.Technical Field

[0002] The present disclosure relates to an electronic document management system.Background

[0003] Electronic documents provide great flexibility in allowing users to edit and format content with great ease. Depending on the software used to create and manage a particular document type, electronic documents may also incorporate version control, wherein different versions of a document are stored over time and a user can subsequently access those different versions.

[0004] Using networked computers, multiple users can access a single electronic document contemporaneously. Further, different permissions and security settings can be associated with an electronic document or folder to control the manner by which users are able to interact with the document or folder, dependent on a level of permission allocated to the respective users. For example, different users may have different levels of permission in relation to opening, reading, editing, deleting, and / or moving a document or folder.

[0005] The great flexibility that electronic documents offer and the ease with which documents can be generated and propagated has resulted in extremely large numbers of electronic documents of many different types being stored and exchanged. In the context of this application, electronic documents are any data type that may be stored on a computer and may include, for example, but are not limited to: emails, text files, word-processing documents, Computer Aided Design (CAD) files, database files, spreadsheet files, images, video files, audio files, presentation files, webpages, and the like. In addition to actual content, an electronic document may include or be associated with metadata that provides further information about the document. Metadata may include, for example, creation timestamp, modification timestamp(s), document type, size, location, creator name, modifier name, and access permissions.

[0006] It has become increasingly difficult to manage large volumes of electronic documents. Traditional approaches stored individual documents in a single file repository, such as an online folder. However, the distribution of documents via email and other document serving facilities, such as Dropbox and the like, presents great difficulty in ensuring that multiple copies of the same document are not being stored unnecessarily and potentially in different places, which can cause confusion and incur storage costs.

[0007] In legal environments, it is often necessary to manage large data sets of electronic documents for different legal matters. Further, it is important that the veracity of electronic documents be verifiable. Electronic discovery, also known as eDiscovery, refers to discovery of electronic documents in legal proceedings. Some jurisdictions have strict regulations and procedures for eDiscovery, which makes the collection, labelling, and organisation of relevant electronic documents difficult.

[0008] eDiscovery may relate to any electronic data present on a computer system that is to be used as evidence or supporting material in legal matters. Such electronic data may be particularly voluminous, covering many thousands of electronic files. Contributing to the difficulty of correctly compiling legal documents for eDiscovery is the presence of unrelated files, such as system files resident on the computer and files that contain content unrelated to a particular matter in question or that contain matter relating to internal administration that should not form part of the eDiscovery.

[0009] When performing an eDiscovery activity, it is often the case that materials are located on multiple computers or in directories belonging to different users. Sourcing materials from different locations frequently means that there may be significant duplication of data across the different sources. For example, different users working on a single legal matter may each store a copy of the same document in their own respective user directories, resulting in multiple copies of the same document.

[0010] eDiscovery in relation to a particular legal matter must reduce all collected source data to contain only data relevant to that particular legal matter. This is done by removing all irrelevant files, removing duplicates, and only retaining files that contain relevant content. In this context “content” need not refer strictly to data contained within a file, as the existence of a file itself, or the name of the file, or properties relating to the creation or modification of the file may constitute relevant content for consideration by the courts.

[0011] Once the relevant data has been collected, it is necessary to organise and label the data so that the data is in a form presentable for a legal team to review and so that the data is in a form that is suitable for use in legal proceedings.

[0012] The procedures associated with eDiscovery are typically so complex that specially trained staff are required, sometimes even requiring the services of external eDiscovery service providers, resulting in additional costs and administrative overheads. eDiscovery service providers often utilise their own proprietary complex software and extensive experience to manage the processes associated with eDiscovery. However, the complexity of the eDiscovery process is a barricade to legal professionals being able to manage and make changes to documents without the intervention of expert assistance. Further, existing approaches require various separate tools to facilitate an end-to-end workflow. For example, creating and editing hyperlinked indices, converting documents to a displayable format, and applying finishing operations are all handled by separate, external tools.

[0013] Thus, a need exists to provide a method and system for managing electronic documents.Summary

[0014] The present disclosure relates to an electronic document management method and system.

[0015] A first aspect of the present disclosure provides an electronic document management system comprising: a co-ordinator; a set of request queues; a set of response queues; a set of microservices, each of which is configured to perform one of a set of predefined processing tasks on one of either a set of predefined item categories or a set of predefined item subtypes; wherein said request queues and said response queues manage communications between said co-ordinator and each of said microservices; and wherein said co-ordinator is configured to: create a project, based on user input; ingest data items by: identifying an item category of said data item; and assigning a set of said processing tasks to be performed on said data item, based on said identified item category; and add said ingested data items to said project.

[0016] A second aspect of the present disclosure provides a method of managing a set of electronic documents utilising the electronic document management system described herein.

[0017] According to another aspect, the present disclosure provides an apparatus for implementing any one of the aforementioned methods.

[0018] According to another aspect, the present disclosure provides a computer program product including a computer readable medium having recorded thereon a computer program that when executed on a processor of a computer implements any one of the methods described above.

[0019] Other aspects of the present disclosure are also provided.Brief Description of the Drawings

[0020] One or more embodiments of the present disclosure will now be described by way of specific example(s) with reference to the accompanying drawings, in which:

[0021] Fig. 1 is a flow diagram illustrating a basic workflow performed for ingesting data into an electronic document management system in relation to an eDiscovery exercise for a legal matter;

[0022] Fig. 2 is a flow diagram illustrating flows in relation to request queues, response queues, and a set of microservices;

[0023] Fig. 3 is a schematic representation of a workflow for creating a project, distinguishing between user input and system performed functionality;

[0024] Fig. 4 is a schematic representation of a workflow for ingesting data, distinguishing between user input and system performed functionality;

[0025] Fig. 5 is a schematic representation of a workflow for ingesting filtering items during review of a project;

[0026] Fig. 6 is a schematic representation of a workflow for viewing items during review of a project;

[0027] Fig. 7 is a schematic representation of a workflow for tagging items during review of a project;

[0028] Fig. 8 is a schematic representation of a workflow for generating a hyperlinked index;

[0029] Fig. 9 is a schematic representation of a workflow for sharing data;

[0030] Fig. 10 is a schematic representation of a workflow for a presentation mode;

[0031] Figs 11a and 11 b are a schematic representation of a workflow for recruiting temp reviewers;

[0032] Fig. 12 is a schematic block diagram representation of a system on which the electronic document management system of the present disclosure may be practised; and Fig. 13 is a schematic block diagram representation of ingestion.

[0033] Method steps or features in the accompanying drawings that have the same reference numerals are to be considered to have the same function(s) or operation(s), unless the contrary intention is expressed or implied.Detailed Description

[0034] Cloud-based computer services are delivered over the Internet (“the cloud”), providing great flexibility and accessibility, as the computing power and storage can be hosted remotely. A user is able to access cloud-based computer services using any computing device that is connected to the Internet.

[0035] The present disclosure provides a cloud-based electronic document management system and associated method. The cloud-based platform of the present disclosure provides a user interface that enables users, including casual and non-technical users, to manage the processing, review, and production of electronic data without requiring specialist training or expert assistance.

[0036] The electronic document management system is configured to handle a predefined set of item categories. For each item category, there is a predefined set of metadata that can be extracted for that item category and a predefined set of processing tasks that can be performed on that item category. For example, item categories may include images, videos, word-processing documents, emails, spreadsheets, and the like.

[0037] The item categories may be defined for a particular implementation.Alternatively, for compatibility purposes, the set of item categories may correspond to the Multipurpose Internet Mail Extensions (MIME) types defined by the Internet Assigned Numbers Authority (IANA). MIME types are standardised identifiers for file formats. Each MIME type is a two-part identifier for identifying file formats, each MIME type comprising a type and a subtype. The type represents the general category of the file format, such as video or text. The subtype identifies the exact variant of the specified type that the MIME type represents. For the MIME type “text”, the subtype might be, for example, plain (plain text), html (HTML source code), or calendar (for iCalendar / .ics) files. Different formats of emails may be classified as different MIME types. For example,• “.msg” file mime type is: “application / vnd. ms-outlook”• “.eml” file mime type is: “message / rfc822”The different types and subtypes within the set of file categories used by the application can be utilised to identify different programming code required to handle the individual files. For example, one microservice is programmed to read an MS Outlook email, and another microservice is programmed to read an RFC822 email.

[0038] For the purposes of displaying the information to the user, on the other hand, the item categories are simplified and the documents are displayed as a high-level category, e.g., “Email”, “Spreadsheet”, etc. Depending on the implementation, the high-level file category may correspond to the “type” sub-field of a MIME type or may be a more general field covering multiple MIME types.

[0039] The electronic document management system functions as a result of combining a set of practices: a) automated pre-processing of ingested data b) user interface simplifications c) simplified workflows for common tasks d) an integrated automatic-hyperlinking word processor e) automated finishing operations

[0040] The electronic document management system is implemented using a customised computer architecture that is programmed to perform automatically services that would otherwise require significant human intervention by an eDiscovery expert. The computer architecture includes a modular pre-processing system that includes: a data co-ordinator module; a set of microservices modules; and a set of processing queues.

[0041] The data co-ordinator module co-ordinates ingestion of data from a set of predefined computer sources. For each item of data that is ingested (i.e., each email, document, image, etc.), a set of microservices execute a predefined list of processing tasks, based on the item category. Microservices are services that each relate to a self-contained piece of functionality with well-defined interfaces that enable different microservices to engage and interact with each other in a microservices architecture. Each microservice is associated with a corresponding request queue and a response queue. The request and response queues manage communication between the data co-ordinator module and the set of microservices modules.

[0042] A message broker is a program that manages all of the message queues. In one implementation, the message broker is implemented using RabbitMQ. It will be appreciated that other message brokers may equally be utilised, including, for example,but not limited to, Amazon Web Services (AWS), HornetQ, IBM MQ, Open Message Queue, and TarantooL Alternatively, a custom-built message broker may be programmed and utilised.

[0043] The message broker manages the sending and consuming of messages in and out of queues. Each message queue functions in a manner similar to a post office box, wherein one program (e.g., a pre-processing co-ordinator) can put messages onto that queue, and another program (e.g., a microservice) can take messages off the queue. Each part of the system that needs to co-ordinate with other parts of the system knows how to send and consume messages from the queues via the message broker.

[0044] There are many microservices, with each microservice being specific to a processing task and an item category or a processing task and item subtype. That is, for an item category for which the same code and libraries can handle all subtypes of that item category, then a single microservice may be programmed to handle a nominal processing task for all subtypes of that item category. For an item category having distinct subtypes that require particular handling, then different microservices are required for each processing task to be performed on each item subtype. This is the case, for example, in relation to the item category of “Email” having item subtypes of .msg and .eml, for which separate microservices are required in order to process the different item subtypes. For example, an item category relating to a compressed “ZIP” file has an associated microservice “decompose” to extract files from the compressed ZIP file:“microservice: task=decompose, item type=application / zip"

[0045] An item type that is an RFC822 formatted email has an associated microservice “extract_metadata” to extract metadata from RFC822 format emails:“microservice: task=extract_metadata, item type=message / rfc822’

[0046] Each microservice is accompanied by a corresponding request queue and response queue. Using the ZIP file example, the associated decompose microservice has a request queue and a response queue:

[0047] There may be multiple instances of each microservice running at any time. The microservices programmed for an implementation of the electronic document management system depend on the item types that are to be processed by that particularsystem and the types of processing tasks that are able to be performed on each of the item types.

[0048] In one implementation, each microservice is a discrete program called to execute on one or more processors in order to perform the associated processing task. In an alternative implementation, each microservice is a functional module within a larger program. Such functional modules may be implemented, for example, as discrete subroutines, functions, or the like, depending on the particular coding. In a further alternative, one or more microservices are implemented as functional modules within a program, with other microservices implemented as discrete programs.

[0049] Fig. 12 is a schematic block diagram representation of a system 1200 that features an electronic document management system 1250 in accordance with the present disclosure. The electronic document management system 1250 is configured to manage a predefined set of item types by utilising a set of microservices to perform a predefined set of processing tasks in relation to each of the item types.

[0050] The electronic document management system 1250 is coupled to a communications network 1299. The communications network 1299 may comprise one or more wired communications links, wireless communications links, or any combination thereof. In particular, the communications network 1299 may include a local area network (LAN), a wide area network (WAN), a telecommunications network, or any combination thereof. A telecommunications network may include, but is not limited to, a telephony network, such as a Public Switch Telephony Network (PSTN) or a cellular mobile telephony network, the Internet, or any combination thereof.

[0051] The electronic document management system 1250 includes each of an ingestion co-ordinator 1252, a request queues module 1254, a response queues module 1256, a set of microservices 1258, and a storage device 1260. The co-ordinator 1252 may be implemented as a single co-ordinator or, alternatively, may be implemented using multiple distinct components that perform different sub-functions of the co-ordinator. For example, in one implementation the co-ordinator includes: a first module to receive uploaded documents, a second module to co-ordinate the microservices, and a third module to update the database with the processed data.

[0052] Similarly, the storage device 1260 may be implemented using a single storage device or a number of distributed storage devices. For example, in one implementation the storage device 1260 is implemented using a single hard drive. In an alternative implementation, the storage device 1260 is implemented using several distinct datastorage components in order to support one or more features, such as database sharding, failover, load balancing, and the like. For example, the storage device 1260 may be implemented using an array of hard drives, co-located or distributed across multiple locations.

[0053] The electronic document management system 1250 also includes a web server 1262 and an application server 1264. The web server 1262 stores, processes, and delivers web elements to a browser executing on browsers executing on computing devices accessed by users. When a user accesses a page of a website hosted by the web server 1262, the browser used by the user sends a request to the web server 1262. The web server 1262 processes the request and then sends back data to be displayed in the browser. The data to be displayed is typically a combination of HTML code with supporting JavaScript code and content file (e.g., media files, text files, and the like) that form a user interface to be displayed to the user and with which the user will interact.

[0054] The application server 1264 receives operation requests from the user interface and executes the requested operations against the database. For example, when the user requests to tag an item, or generate an index, or add a user to a project, the application server receives the request and then executes code to update the database 1260 to effect the change, and then send push notifications back to the browser.Depending on the particular implementation and application, the application server is configured to perform a range of functions. Such functions may include, for example, but are not limited to: user access management; account & project management; service discovery; and database operations (create, read, update, delete).

[0055] Communication among the various components 1252, 1254, 1256, 1258, 1260, 1262, and 1264 occurs via one of more connections, illustrated as a bus 1260. The request queues module 1254 and response queues module 1256 manage communications between the co-ordinator 1252 and the set of microservices 1258. The storage device 1260 is utilised by the electronic document management system 1250 to store data items ingested into the system 1250.

[0056] One or more users 1205a... 1205n are authorised to access the electronic document management system 1250 by accessing computing devices, such as computing devices 1210a... n, each of which is coupled to the communications network 1299. In the example of Fig. 12, the users 1205a... 1205n form part of a legal team compiling documents in relation to a legal matter. By utilising the computing devices 1210a... n, the users 1205a... n are able to create and manage projects. The users 1205a... n are also able to share documents within the legal team and externally of thelegal team by creating and sharing hyperlinked indices. As discussed above in relation to the web server 1262, the users 1205a... 1205n access a user interface displayed on a display device of the respective computing devices 1210a... n to interact with the electronic document management system 1250, with content displayed to the users 1205a... 1205n being provided by the web server 1262. In the embodiment of Fig. 12, the system utilises a web server 1262 to serve content to web browsers executing on the user computing devices 1210a... n. Alternatively, the system 1200 is implemented utilising one or more native applications that are downloaded to the user computing devices 1210a... n for executing on those computing devices 1210a... n, with the native applications pushing and pulling data, as required, from the server 1250. In a further alternative, a combination of web-based browser content and app-based content is utilised.

[0057] In the example of Fig. 12, external user 1220 utilises a computing device 1225 to access documents via hyperlinks embedded in a hyperlinked index document sent to the external user by the electronic document management system 1250 in response to an authorised user 1205a... n providing the external user 1220 with access to that hyperlinked index document.

[0058] The system 1200 also includes a reviewer 1230 who accesses a computing device 1235 to access a data set of documents on the electronic document management system 1250 in order to review those documents in accordance with a proposal presented and approved by one of the authorised users 1205a... n.

[0059] In order to manage large volumes of data items, it is preferable to be able to remove all “junk” data items that are not relevant to the particular legal matter or project, the electronic document management system is optionally configured to identify junk data items automatically and generate a report indicating why a data item has been marked as junk.

[0060] The electronic document management system is able to identify junk data items using a number of different techniques, depending on the implementation. A first approach is to identify duplicate item families. For example, an email and any attachments are marked as belonging to a single item family. The system computes a cryptographic hash for each item family. Any suitable cryptographic hash can be used, such as MD5 which produces a 128-bit hash value. Once the system has generated a hash for a new item family, the system compares the hash against hashes that are already present in the system. The presence of a matching hash indicates that the data item is a duplicate and the system marks the data item accordingly.

[0061] In one implementation, the ingestion co-ordinator generates the hash and performs the comparison against other hashes stored in the storage device of the electronic document management system. In an alternative implementation, the electronic document management system includes a hashing microservice that is programmed to generate the hash and perform the hash comparison. In such an implementation, the ingestion co-ordinator sends a request to the hash microservice, via a request message queue, and waits for a response from the hash microservice, via the response message queue.

[0062] The system may also identify common non-user generated items, such as system files and the like. In general applications, system files are not likely to be relevant to the legal proceedings or project. Accordingly, the system filters system files as data items are ingested into the system. Filtering may be performed in many ways and may include, for example, analysing file properties (e.g., file type), analysing file size, comparing a cryptographic hash of the data item to a list of known non-user generated file hashes, or a combination thereof. A list of known non-user generated file hashes may be, for example, the National Software Reference Library list of hashes.

[0063] In order to identify junk images, one implementation analyses property data associated with a data item, such as image size and data context. Source data, which often contains logos and / or icons, may also be analysed to identify junk images.

[0064] The electronic document management system is configured to process data in accordance with a predefined workflow, with different portions of the workflows utilising specially programmed microservices to interact with user input.

[0065] Fig. 1 is a flow diagram illustrating a basic workflow 100 performed for ingesting data into an electronic document management system in relation to an eDiscovery exercise for a legal matter. The workflow 100 begins at a Start step 105 and passes to step 110, in which a user uploads data items for ingestion by the system. The electronic document management system provides a user interface by which the user is able to select data items to be uploaded. Depending on the implementation, the user interface may enable the user to drag and drop selected data items to be uploaded. The user interface may also enable the user to browse contents of a computing device to select data items to be uploaded. The user interface may also provide searching and / or filtering utilities to assist the user to find and identify data items to be uploaded. As describe above, eDiscovery is a difficult task that typically involves receiving data items from multiple sources. The data uploaded by the user in step 110 may be from a single source or from many sources.

[0066] During an initial upload of data items, the system prompts the user to create a project against which data items will be stored. Alternatively, the system automatically generates unique project identifiers against which data items are stored. It is common for data items to be uploaded over multiple sessions. In subsequent sessions, the user enters the project identifier in order to store data items against the correct project.

[0067] Data uploaded by the user in step 110 is presented to the co-ordinator 115. In some implementations, the co-ordinator extracts metadata from the uploaded data items. The co-ordinator 115 checks to determine whether the data items uploaded by the user in step 110 are new items or whether any of those data items have already been uploaded to the electronic document management system. The co-ordinator 115 utilises the extracted metadata to identify duplicate data items and prevent further unnecessary processing of any such duplicate items. In one implementation, the co-ordinator 115 only displays the new data items to be processed. In another implementation, the co-ordinator 115 displays all data items uploaded by the user in step 110, but marks duplicate items that have already been uploaded. Duplicate items may be identified, for example, by displaying the names of those data items in a different colour, against a different background, with a check mark next to them, or any combination thereof.

[0068] The co-ordinator 115 processes each data item, identifies the processing tasks to be performed on that data item, and then sends microservice requests for each processing task of each data item to a microservices request queuing module 120 of a message broker 160. The message broker 160 manages all message queues within the electronic document management system through the request queuing module 120 and a response queuing module 130. The microservices request queuing module 120 manages request queues for all of the different microservices.

[0069] For example, a user uploads a data item in the form of a Microsoft (MS) Word document in step 110. In step 115, the ingestion co-ordinator identifies the type and subtype of the data item and determines that “decompose”, “extract_metadata”, and “render” tasks are to be performed by the relevant microservices for the uploaded MS Word document. Accordingly, the ingestion co-ordinator 115 sends processing request messages to each of the MS Word decompose queue, the MS Word extract_metadata queue, and the MS Word render queue within the message broker 160. It will be appreciated that some tasks to be performed on a received data item can be performed in parallel, whilst other tasks may need to be performed in series. For example, an encrypted file first needs to be decrypted, before tasks can be performed on the content of the encrypted file.

[0070] The microservices request queuing module 120 manages the requests and forwards the requests to a set of microservices 125. The set of microservices 125 includes one or more microservices to perform processing tasks for the different item types presented in the incoming requests.

[0071] A microservice receives a request from the queuing module and reads the received request. The microservice processes the request and then sends the results of the processing to an accompanying response queue. During the processing, the microservice may identify further sub-items to be processed. Such sub-items may include, for example, attachments to emails or embedded documents or images. Any sub-items identified by the microservice are sent back to the co-ordinator to be processed in the same way as the uploaded data items.

[0072] The response queuing module 130 of the message broker 160 receives results from the set of microservices 125 and then returns the results to the co-ordinator 115. The co-ordinator 115 reads the microservice responses for each task of each item. In many cases, different tasks associated with a data item are performed in parallel by different microservices. Using the example of the MS Word document from above, the coordinator 115 sends request messages to the “decompose” message queue, the “extract_metadata” message queue, and the “render” message queue, all at substantially the same time. The co-ordinator 115 then monitors the response queues 130 for any responses. Whenever a response is received, there will be some additional data resulting from the processing (for example, the “decompose” response message contains information about the children of the data item, and the “extract_metadata” response message contain any metadata that was extracted). The co-ordinator 115 adds the new data returned from the response queues 130 to the data item. When all tasks associated with a data item have been completed, the co-ordinator 115 marks that data item as “processing completed”. When all uploaded data items have been marked as “processing completed”, the ingested data is stored in the database 1260 against the relevant project identifier.

[0073] The co-ordinator 115 determines at step 135 that all tasks have been completed for all data items and control passes to step 140, in which the ingested data is added to a selected case. In step 145, a user’s browser is updated to reflect the newly ingested data. That is, the system sends a push notification to the user interface displayed on the computing device accessed by the user, wherein the push notification updates the user interface to indicate that ingestion of the new data has been completed. In some implementations, the user interface displays a dialog box indicating that ingestion hascompleted. In other implementations, the user interface displays information pertaining to the newly ingested items in a region of the user interface, which may be a newly opened window, for example.

[0074] The electronic document management system utilises the microservices 125 to perform processing tasks on ingested data items. The set of processing tasks performed by the microservices depends on the implementation of the electronic document management system. In particular, the set of processing tasks depends on the item types for which the electronic document management system is configured and the set of processing tasks defined for each of those item types. As an example, the set of processing tasks may include, but is not limited to:• automatic decomposition of items• automatic extraction of metadata• automatic generation of Summary metadata• automatic extraction of text• automatic conversion of items to a common display format• automatic identification of junk items

[0075] In some embodiments, the electronic document management system utilises the microservices 125 to perform automated pre-processing of data items ingested by the user on step 110. The system performs automatic decomposition of data items, such that any nested items are processed as individual data items. Thus, email attachments or the contents of compressed archival files, such as ZIP files, are processed individually. The system then performs automated extraction of metadata from each data item, based on the item type.

[0076] The relevant “extract_metadata" microservice categorises the extracted metadata by context so that it is easier for a user to understand from where the metadata originated. For example, metadata can be categorised into “File System” for file system properties, and “MS Word” for metadata embedded in the MS Word properties of an MS Word document: a. File System > Last Modified b. MS Word > Last Saved

[0077] The categories of metadata depend on the origin of the metadata, and this includes the item type. For example, one “.msg” email file ingested by the electronic document management system may contain the following pieces of metadata (not a complete list):Application > MD5 Hash (hash value generated by the system)Application > Mime (the mime type generated by the system)Application > Item Type (high level item type generated by the system)File System > Filename (the original filename of the item)File System > Path (full file path of the file)File System > Last Modified (the file system’s last modified date)Email > Date (the date the email was sent)Email > Subject (the subject of the email)Email > From (person who sent the email)Email > To (people the email was sent to)Email > CC (people the email was CC’d to)Summary > Title (editable summarised value for Title)Summary > Date (editable summarised value for Date)Summary > To (editable summarised value for To)Summary > From (editable summarised value for From)

[0078] Some embodiments generate “Summary” metadata, which is a simplified set of metadata extracted from the ingested data items. eDiscovery processes often involve extracting many pieces of metadata from each data item, resulting in a lot of information for a user to comprehend in order to understand that data item and its context within the legal proceedings. Defining a set of Summary metadata distils the full set of metadata down to a common consolidated set of basic values. The Summary metadata values may change, depending on the implementation

[0079] In one example, the Summary metadata values include: Date, Title, To, and From. The Date value may be chosen from several available metadata dates, based on context. For example, in some circumstances the system selects the creation date associated with a MS-Word document, rather than the last accessed date or last modified date. In other circumstances, the system selects the last modified date associated with the MS-Word document. The system generates the Title value based on either a file name already allocated to a data item or a portion of title metadata. The format of the Title value depends on the implementation. In some embodiments, the Title value is a standardised format across all categories of data items. For example, all Title values areunique alphanumeric character strings of a predefined length. In other embodiments, the Title values depend on the category of data items. For example, for an email, the Title value corresponds to the email subject or a truncated version of the email subject, based on a predefined Title value length. MS Office documents often include a “Title” metadata field and the electronic document management system can use the embedded metadata field for ingested MS Office documents.

[0080] The system generates the To value based on any available metadata that indicate recipients associated with the data item. Similarly, the system generates the From value based on any available metadata that indicate an author or authors associated with the data item. Depending on the particular implementation, long-form data, such as email addresses, may be simplified. For an MS Word document, the system utilises the document author, if available, to populate the “From” value. The “To” value is left blank or populated with a default value. For an email item, the “From” value is the sender and the “To” value is one or more of the recipients.

[0081] The electronic document management system provides a user with a user interface by which to upload data items and navigate through data items associated with one or more projects. In some embodiments, the Summary metadata is displayed prominently in the user interface. The user can choose to drill-down further into the full set of metadata, if required.

[0082] The electronic document management system performs automatic extraction of text from an ingested data item. Extracting the textual content from each data item enables filtering and searching to be performed based on keywords. In some circumstances, such as images and flattened PDF documents, the electronic document management system performs optical character recognition (OCR) in order to extract the textual content. The extracted textual content is stored in a manner similar to other metadata associated with the data item. In one implementation, the electronic document management system stores the extracted textual content in two formats: (i) plain text; and (ii) fast text searching. Fast searching may be performed, for example, as a database function, depending on the implementation. Such database functions typically analyse and index text up-front, so that the relevant database does not need to read and process the full text content when a search is conducted during run time. Suitable database functions that provide fast text searching include, but are not limited to, Apache Solr, Elasticsearch, PostrgeSQL tsvector, MySQL fulltext index, and MongoDB text index.

[0083] The electronic document system may also be configured to automatically convert data items into a common display format. Ingested data items may be of many differentformats. In order to facilitate handling and viewing, the system converts all data items into a common display format, such as Portable Document Format (PDF) from Adobe Inc, that is readily viewable on different computing system platforms and operating systems.

[0084] Fig. 2 is a flow diagram illustrating an arrangement 200 of the request queues 120, response queues 130, and the set of microservices 125 of Fig. 1 . As described above, the co-ordinator 115 processes an incoming data item and determines a set of tasks to be performed on that data item, based on the item category and any type or subtype associated with the data item. Each combination of item category and associated task has a corresponding request message queue 220, response message queue 230, and microservice 225 to perform the task on that particular item category.

[0085] For example, for a task “extract_metadata” in relation to a data item of category Email and type / subtype “message / rfc822”, there are the following:• a request message queue named: “task=extract_metadata, item type=message / rfc822, queue type=request”,• a response message queue named: “task=extract_metadata, item type=message / rfc822, queue type=response”,• a microservice named:“task=extract_metadata, item type=message / rfc822”,

[0086] There may be multiple instances of each microservice running, corresponding to processing different data items of the same type / subtype combination. In such scenarios, the instances of the microservice share the same request and response message queues. Further, there are instances in which a task is not appropriate for a particular item category and / or type / subtype combination. For example, there is no need to render a compressed file, such as a ZIP file, to PDF format. In such a case, the ingestion coordinator skips the rendering task when processing a ZIP data item and does not send a rendering request.

[0087] For each data item, there is a set of tasks 205 to be performed on that data item. Tasks that are independent of each other can be processed in parallel, if the scheduling and computer resources are available. In such scenarios, those independent tasks can be sent to corresponding message queues concurrently for execution by the respective microservices in parallel with each other. For example, a first microservice can decompose a data item at the same time as a second microservice extracts metadatafrom the data item and a third microservice renders the data item. If computational resources are limited, then the microservices can be executed sequentially.

[0088] At step 210, the system determines the item type for the data item. In this example, the data item is an RFC822 formatted email, so a request “item type = message / rfc822” 215 is sent to a Request Queue 220 in relation to a set of microservices 225. The relevant microservice for extracting metadata from an RFC822 formatted email performs the request task by extracting metadata from the email data item. The microservice 225 sends a response to a Response Queue 230, which then returns the response to the co-ordinator 105 of Fig. 1 .

[0089] When a microservice has processed a request, the microservice posts a message onto the response queue (130 and 230). The message on the response queue contains new data resulting from the processing (e.g., an “extract_metadata” microservice provides the metadata that was extracted from the data item). The ingestion co-ordinator 115 continually monitors all of the response queues. When the ingestion co-ordinator 115 sees that a message has been posted to the response queue, the ingestion co-ordinator 115 consumes this message, reads the new data, and appends the new data to the data item. When all of the tasks from the set 205 have been executed, the ingestion coordinator 115 marks the data item as “processing completed”. When all of the ingested items (including child items discovered in the “decompose” stage) have been marked as “processing completed”, the ingestion / pre-processing coordinator adds the new items to the case data by storing those items into the database 1260.

[0090] As indicated above, once data items have been ingested into the system and processed according to item type, a dashboard presented by the electronic document management system to a user accessing a computer device is updated to reflect the newly ingested data items. Traditional eDiscovery processes are complicated and require training and considerable experience in order to interact with the data. The dashboard of the present disclosure provides an intuitive interface suitable for access by legal practitioners without requiring extensive training or experience in eDiscovery.

[0091] As described above, the electronic document management system of the present application is configured to identify duplicate items and junk items. Managing which items should be part of a working data set in relation to a particular project and which items should be excluded from the working data set is a fundamental aspect of eDiscovery processing. Existing approaches rely on a user to accurately set parameters when filtering or viewing a data set. Using the wrong parameters can inadvertently result inunintended junk or duplicate items being included in results displayed to the user. Displaying unintended junk or duplicate items can cause confusion and embarrassment.

[0092] The dashboard provides an efficient user interface for display within a browser executing on a computing device accessed by a user or within a software application executing on a computing device accessed by the user. The computing device may be a smart phone, personal computer, phablet, tablet computing device, or the like.

[0093] In one implementation, data items that have been marked as junk are not displayed in a main display region of the dashboard. Consequently, any data items that have been identified as being junk will not be retrieved by the system and shown as a result of any operations selected by the user in the main section of the user interface. This assists in preventing confusion between relevant data items and junk data items, which may otherwise result in accidental deletion of relevant data items or unnecessary processing of junk data items.

[0094] The user interface provides a separate section through which a user is able to interrogate data items that have been marked as junk. The user is able to swap between the main section of the user interface and the junk section of the user interface. This enables the user to view data items that have been marked as junk and view the reasons and context relating to why the data item was classified as being junk. The user can override the automatic classification of any of the data items marked as being junk, un-marking those items at which point the system moves the un-marked data items to the main section of the user interface. Conversely, the user is able to view items in the main section of the user interface and, if the user determines that a data item in the main section is actually a junk item, mark that data item as junk, whereupon the system moves the newly marked junk data item to the junk section of the user interface.

[0095] The electronic document management system classifies data items into high level item categories. Thus, emails from different email systems that would be classified as “Microsoft Outlook Note” or “application / vnd. ms-outlook” in other eDiscovery systems are classified as “Email”. This high level classification makes it easier for users to interact with the system and data items. In one implementation, the following item categories are defined: Archive, Audio, Calendar, Code, Document, Drawing, Email, Email Archive, Html, Image, Multimedia, PDF, Presentation, Spreadsheet, System, Text, and Unknown. It will be appreciated that some item categories may be omitted in some implementations.Further, some implementations can have additional item categories and may optionally provide a user with the ability to define custom item categories in order to process data items relating to a particular industry or environment.

[0096] The user interface allows users to view and manage data items by using simple Boolean operators. Boolean operators are a powerful tool that can be utilised to explore a data set. However, the underlying logic can be unintuitive for a non-technical user and thus can result in unexpected results when filtering items. This is particularly the case when a non-technical user combines multiple operators and is unaware of the precedence of logical operators. In order to prevent inaccurate results, the user interface supports a set of predefined logical conditions, which are presented in unambiguous natural language.

[0097] In some embodiments, the set of predefined logical conditions include: a. Item contains any of these keywords b. Item contains all of these keywords c. Item does not contain any of these keywords

[0098] The electronic document management system provides a portion of the user interface that enables a user to edit the “Summary” metadata section. In some circumstances, it is useful for a user to be able to edit the Summary metadata automatically compiled during ingestion of data items. As described above, the Summary metadata includes a set of predefined values, such as Date, Title, To, and From. The user interface enables the user to change the automatically generated Summary metadata using custom content. The user is also able to add customised values. For example, if a case dataset contains a number of documents in a particular format, such as lots of purchase orders, then the user may want to capture a certain piece of information from those documents, such as a purchase order number. The user is able to add a customised field to the Summary and the purchase order number can be input manually to the respective field.

[0099] The electronic document management system has been implemented to improve workflows within an eDiscovery process. Consequently, the workflows require minimal user input, resulting in a more consistent and accurate process that provides a self-service platform that can be accessed and used by non-technical users.

[0100] The simplified workflows implemented in the electronic document management system include: a. Creating a case b. Ingesting data c. Filtering data d. Review i. Filtering items ii. Viewing itemsiii. Tagging items e. Generating a Hyperlinked Index (including specifying finishing options) f. Sharing Data

[0101] Fig. 3 is a schematic representation of a workflow 300 for creating a project, distinguishing between user input and system performed functionality. The workflow 300 begins at a Start step 305 and moves to step 310, in which a user accesses a computing device to interact with a user interface of the electronic document management system. The user clicks a “Project Creation” button on the user interface and enters a name to be used for the newly created project.

[0102] Depending on the implementation, the system optionally governs the structure of project names, such as the minimum and / or maximum number of characters that are required, the type of characters that can and / or cannot be used and any other formatting requirements. For example, one implementation requires that all project names be 8 characters long and may consist of alphanumeric characters only. This can be useful in large organisations that have many projects, as it ensures that there is some consistency in the naming conventions. Another implementation restricts use of some special characters, such as @, / , and %.Control passes from step 310 to step 315, in which the system creates a project on a backend server. The system allocates storage space in a storage medium for the project and associates the storage space with the user-entered project name. The front-end user interface sends a new project creation request to the application server. The application server populates the database with a set of entries for the new project. The back end server then sends a push notification to update the user interface that the project has been created and the browser executing on the user computing device updates to display the new project.

[0103] Control passes from step 315 to step 320, in which the system browser auto-navigates to a new project. That is, the system displays a user interface to the user, wherein the user interface shows a view of the new project. The view of the new project may include controls to enable the user to change attributes of a profile associated with the project, including the project name, access levels, and access rights for different users. The view may also include controls that enable the user to interact with the project, such as by selecting data items to be uploaded to the project. Control passes to an End step 325 and the workflow for creating a project is complete.

[0104] Data items come in many different formats and yet in legal proceedings must be presented in a common format for review. Processing different formats can involve manydifferent methods, each of which can affect the output. Further, it is difficult to access and select an appropriate rendering type to process a particular document. The electronic document management system addresses these issues by defining a set of processing tasks to be performed on each of a set of predefined item types that can be handled by the system. The system utilises different microservices to perform the processing of the tasks. Depending on the implementation, the microservices may be located on a single computing device or distributed across multiple computing devices or a combination thereof. In one example, the microservices are stored on a cloud computing platform, such as Amazon Web Services (AWS), Microsoft Azure, Google Cloud Platform, vmware, AlertLogic, SoftLayer, or the like.

[0105] Fig. 4 is a schematic representation of a workflow 400 for ingesting data, distinguishing between user input and system performed functionality. The workflow 400 begins at a Start step 405 and moves to step 410, in which the user clicks a data ingestion button on the user interface to begin uploading data items into the electronic document management system. In step 415, the user selects data items and submits those selected items for data ingestion. Depending on the implementation, the user interface provides various options by which the user is able to select data items for ingestion. For example, options may include a “Drag and drop” region of the user interface into which the user can upload files, a file dialog window, or the like.

[0106] Control passes from step 415 to step 420, in which the user interface is updated to indicate that ingestion of the user selected data items is in progress. All affected users, meaning all users with access to the project and who are logged in, will see an update. Control passes to step 425, in which the system uploads the selected data items to a backend server. In step 430, the system applies automated ingestion processing to the uploaded data items, as described with reference to Figs 1 and 2. In step 435, the browser is updated with the newly ingested items. Control passes to step 440 and the ingestion of data workflow 400 completes. As it may be necessary to ingest new data multiple times during a lifecycle of a project, the workflow 400 may be repeated at any time that new data items need to be ingested.

[0107] Fig. 5 is a schematic representation of a workflow 600 for ingesting filtering items during review of a project. The workflow 500 begins at a Start step 505 and moves to step 510, in which a user selects filters to apply to a data set associated with a project. The filters may include, for example, date ranges, keywords, sender, receiver, and the like.

[0108] In step 515, the system translates the user input into a database query and in step 520 the system executes the query on the database and retrieves matching items. One implementation utilises a custom “Domain Specific Language” (DSL) to aid in communicating the filtering criteria between the user interface and the application server. The front end user interface is programmed to encode the Ul filtering controls (buttons, checkboxes, textboxes, etc) into a query string in the DSL syntax. The front end user interface then sends the query string to the application server. The application server translates the DSL query string into database queries that can be executed against the database 1260 to retrieve the matching documents.

[0109] In such an implementation, the DSL is programmed to make encoding of filtering criteria unambiguous for the programming code to understand, but also easy for a human to understand. This enables the system to display the query to the user in a format that can be easily read.

[0110] Step 525 updates the browser to display the matching items and control passes to an End step 530 at which the workflow 500 terminates.

[0111] Fig. 6 is a schematic representation of a workflow 600 for viewing items during review of a project. The workflow 600 begins at a Start step 605 and moves to step 610, in which the user selects a document to view by clicking on a document from a document list, such as may be returned from the end of workflow 500. In step 615, the browser updates to display preview, metadata, and family relationships for the selected document. Control passes to step 620 and the workflow 600 terminates.

[0112] Fig. 7 is a schematic representation of a workflow 700 for tagging items during review of a project. The workflow 700 begins at a Start step 705 and moves to step 710, in which the user selects one or more documents for tagging. Step 715 determines whether a tag already exists. If the tag already exists, Yes, control passes to step 720 in which the user tags the document(s). Tagging the document(s) may be effected in different ways, depending on the implementation, including, for example, utilising a shortcut key or a button on the user interface. Control passes from step 720 to step 735, in which the backend server applies the user selected tag to the selected document(s).

[0113] Returning to step 715, if the tag does not exist, No, control passes to step 725, in which the user generates a new tag, such as by clicking on a “new tag” button, and assigns a name to the new tag. Examples of tags that might either be predefined or user-defined include “Privileged”, “Relevant”, “Fraud”, or the like. User-defined tags may relate, for example, to internal protocols. Control passes from step 725 to step 730, inwhich the back end server creates a new tag based on the user input. Control then passes to step 735, in which the backend server applies the user selected tag to the selected document(s).

[0114] Control passes from step 735 to step 740, in which the system updates the browser to reflect the tagging applied by the user to the selected document(s). Control then passes to End step 745 and the workflow 700 terminates.

[0115] Fig. 8 is a schematic representation of a workflow 800 for generating a hyperlinked index. The application server reads a template into memory and then loads the relevant data for a set of selected items into memory. The application server iterates through the selected items and inserts entries into the template document for each item. The application server saves the updated template into the database and the updated version of the template is then available for viewing by a user. The workflow 800 begins at a Start step 805 and moves to step 810, in which the user utilises a user interface displayed in a browser executing on a user computing device to select documents. In step 815, the user clicks an index generator button and specifies a name and template for the index that is to be generated.

[0116] Control passes from step 815 to step 820, in which the system generates an index, based on a request generated by step 815. The request contains identifiers for each of the items selected in step 810, along with the template selected by the user in step 815. In step 820, the backend server of the system processes the request to generate an index. The backend server creates a new entry in a database for the new index. The server then retrieves data associated with the selected items from the database. The server opens a copy of a template index and executes logical code to populate data from the selected items into the new index. The logical code also populates hyperlinks into the new index from each of the selected items. The server stores the populated index in the database and then sends a push notification to a browser window to notify the user of the Universal Resource Locator (URL) of the newly generated index. The browser executing on the user computing device then displays the new index for viewing and editing by using a built-in word processor.

[0117] Control passes to step 825, in which the system navigates the browser to the index. Control then passes from step 825 to step 830, in which the user edits the index as require using a built-in word processing interface of the user interface. In step 835, if required, the user specifies finishing options to apply to hyperlinked documents. In step 840, if required, the user shares the index with other parties by clicking a share index button. Control then passes to End step 845 and the workflow 800 terminates.

[0118] Fig. 9 is a schematic representation of a workflow 900 for sharing data. The workflow 900 begins at a Start step 905 and moves to step 910, in which the user selects a documents list or hyperlinked index to be shared with other users. Control passes to step 915, in which the user clicks a sharing button and selects a set of recipients with whom to share data. The users may be registered users with the electronic document management system, in which case internal usernames may be utilised to select recipients. Alternatively, users may be identified by email addresses, or other unique identifiers that are suitable for document sharing (e.g., Dropbox account, or the like). The user optionally selects whether the link is to be read-only or whether the recipients have other authority to modify, share, or otherwise deal with the link. The system optionally sets links to be shared as read-only.

[0119] From step 915, control passes to step 920, in which the system generates a sharing link, based on a request sent from the browser executing on the user computing device at the end of step 915 to the application server. The request includes a resource identifier of the resource that is to be shared. For example, if the user in step 910 selects a hyperlinked index, then the resource identifier in the request sent from step 915 is the identifier of that selected hyperlinked index. The request also includes a list of users with whom the resource is to be shared, and, optionally, a set of sharing options. The sharing options may define, for example, the level of access to the shared resource, such as readonly, ability to copy, ability to share, ability to modify, and the like.

[0120] In step 920, the application server receives the request and inserts a new entry into the database that includes an automatically generated URL address for the link, the resource identifier, the user list, and the set of sharing options. When the link is accessed, the system utilises the database entry to check that the person accessing the link corresponds to one of the users in the set of users and checks the access rights, as defined by the set of sharing options, before actually retrieving the resource, when the user is verified as an authorised user. In step 925, recipients are granted access to the link and in step 930 the system sends an email to the recipients, wherein the email contains an access link.

[0121] In step 935, the recipients have received the emails sent from the system and click the access link, whereupon the system delivers the link contents in a browser. In step 940, the recipients consume data from within the browser or download the data, as required. Control passes to End step 945 and the workflow 900 terminates.

[0122] In some scenarios, it is desirable to present documents in a briefing situation, such as an in a courtroom setting. The system enables users to log in to a digital “virtualroom” (a web page), where document display is synchronised among all participants in the room. A user loads a list of selected presentation documents into the room and one or more participants log in to the room. One or more controlling users then control the display of the presentation documents by selecting a document to be displayed at any given time. When a controlling user selects a document to be displayed, that selected document is displayed to all participants in the virtual room as a result of the system updating the browsers viewed by the participants to show the selected document.

[0123] Fig. 10 is a schematic representation of a workflow 1000 for a presentation mode. The workflow 1000 includes the involvement of a controller (being a user in control of presenting a presentation), one or more viewers (being users who are to receive and view a presentation), and the system. The workflow 1000 begins at a Start step 1005 and moves to step 1010, in which a controller selects documents for presentation and then in step 1015 the controller clicks a new presentation virtual room button, specifies a name, and assigns room access to parties.

[0124] In step 1020, a server of the electronic document management system sets up a virtual room, access link, and authorisation for parties. In step 1025, the system sends an access link to the selected parties. In step 1030, the viewers log in to the virtual presentation room created by the system in step 1020.

[0125] Once the presentation is to commence, the controller in step 1035 selects a document to be displayed in the presentation. Once the controller has selected a document to be displayed, the system in step 1040 updates the browsers for all users in the virtual room to display the selected document. Steps 1035 and 1040 repeat as needed during the presentation. Once the presentation has concluded, control passes from step 1040 to an End step 1045.

[0126] Legal teams often need to scale up their review teams in response to incoming requests. In order to do so, temporary reviewers are sometimes recruited. However, doing so often takes time to train the reviewer. The system of the present disclosure optionally includes a database of registered ‘Temp’ reviewers that can be recruited by users of the system on an as-needed basis to complete review tasks. The process follows the following steps: a. Reviewers (e.g., paralegals) register with the system b. Users (e.g., lawyers) submit review task requests to the system c. The system matches up reviewers with review tasks d. The reviewer and the user can negotiate the terms of the review task, including payment termse. On acceptance, the system grants temporary access to the reviewer to the project f. The reviewer executes the specified task g. Once the reviewer has completed the task to the satisfaction of the user, the user can authorise payment to the reviewer. The system then revokes the temporary reviewer access

[0127] Figs 11 a and 11 b are a schematic representation of a workflow 1100 for recruiting temp reviewers. The workflow 1000 includes the involvement of a user, one or more reviewers, and the system. At a first Start step 1105, the user accesses the system and at step 1110 the user sends a request for reviewers to assist with document review, wherein the request includes proposal terms. Control passes from step 1110 to step 1125. In parallel, at a second Start step 1115, a reviewer accesses the system and in step 1120 the reviewer signs up with the system as a registered reviewer, which may include providing personal information and payment details. Control passes from step 1120 to step 1125.

[0128] In step 1125, the system matches reviewers with requests and in step 1130 the system sends proposals to the reviewers. In step 1135, the reviewer negotiates the proposal with the user, which involves interaction between step 1135 and step 1140, in which the user negotiates the proposal with the reviewer. In step 1145, the reviewer accepts the proposal. The workflow 1100 continues on Fig. 11b, in which at step 1150 the user accepts the proposal.

[0129] From step 1150, control passes to the system at step 1155, in which the system grants the reviewer temporary access to a project. In step 1160, the reviewer performs a review task in relation to that project, in accordance with the terms of the proposal.Control passes to step 1165, in which the user monitors the review process. During the review process, there may be various interactions between the user and the reviewer, shown by the arrows between steps 1160 and 1165. Such interactions may include instructions and guidance from the user to the reviewer, and questions and interim results from the reviewer.

[0130] Control passes from step 1165 to step 1170, in which the user verifies that the review task is successfully completed and authorises payment to the reviewer. Control then passes to step 1175, in which the system revokes the reviewer’s temporary access. Control passes from step 1175 to step 1180, in which the system processes payment from the user to the reviewer, in accordance with the terms of the proposal. Payment may be effected using an internal payment gateway or an external, third-party payment gateway,such as SecurePay, eWay, Square, WorldPay, and the like. Control then passes to End steps 1185 and 1190, which terminate the workflow 1100 for the user and the reviewer.

[0131] eDiscovery often involves sharing document lists with other parties. This is commonly done using an index document containing hyperlinks to a number of supporting items that are supplied together in such a way that clicking a hyperlink in the index document will open the supporting items. However, the traditional process of producing these linked lists of documents is technical and usually involves collaboration between a legal team, who are required to edit the content of the index document, and an eDiscovery expert, who facilitates the technical aspects of hyperlinking. This collaboration is time consuming and cumbersome, particularly when index documents evolve over time, which happens quite frequently. Further, the hyperlinks in these index documents can be easily invalidated if the documents are mis-handled. Such invalidation may occur, for example, if the index is moved without moving the supporting documents or if a user clicks “Save As” in MS Word.

[0132] In order to remove the risk of inadvertent user error and to remove the need for separate teams to collaborate to generate a hyperlinked index, the electronic document management system provides an integrated word processor that enables a user, such as a non-technically trained lawyer or legal team, to prepare an index without requiring the assistance of an eDiscovery expert.

[0133] The integrated word processor automatically handles all hyperlinking and the index document and supporting hyperlinked items are managed such that the links cannot be accidentally invalidated. The word processor enables text editing in a manner that is similar to common word processors, such as Microsoft Word or Google Docs.

[0134] In order to create an index document, the user can follow the flow of Fig. 8 or alternatively can create an index from scratch, beginning with a blank document. The word processor enables collaboration from multiple users, such that multiple users can access and edit the index document contemporaneously.

[0135] The word processor is “aware” of the data items in the project data set to which the index document relates, in that the word processor has access to the project data set. Accordingly, the word processor is able to compare text within the index to the identifiers of the items in the project in real-time or near real-time, as the user types. The word processor automatically creates a hyperlink on identifying a match.

[0136] When a user edits an index document by entering a reference to an item that is present in the project data set, such as by typing the reference manually or by selectingthe item from a list, the system automatically creates a hyperlink to the item. Clicking the hyperlink within the index document automatically takes a user to the linked document. The user is able to edit one or more properties of the hyperlink, such as to specify which particular version of an item is to be linked (e.g., an original native version of the item or a converted display format version of the document).

[0137] It is possible to apply automated finishing operations to the hyperlinks.Automated finishing operations may relate to particular formatting or labelling, such as page numbering, document stamping, sizing of documents, and rotation of documents. Automated finishing operations are described in greater detail below.

[0138] The hyperlinked index documents can be shared with other parties, as described in relation to the workflow 900 of Fig. 9. Sharing a hyperlinked index document readily provides links to a set of relevant documents to a set of recipients.

[0139] In many situations, documents need to comply with certain protocols. This is particularly common in tender submissions and eDiscovery for legal proceedings. Protocols often require certain document finishing operations to be applied to documents. Document finishing operations may include, for example: a. Stamping document identifiers onto the produced documents b. Stamping running numeric pagination in the header or footer for all documents across a document set c. Sizing documents to a standard format (e.g., resizing to A4, A3, or Legal) d. Rotating documents to a common orientation

[0140] Document finishing operations are traditionally executed by eDiscovery experts or other third party service providers, as applying such document finishing operations are either beyond the capabilities of non-technical users or require too much time of expensive legal professionals.

[0141] The electronic document management system optionally applies one or more document finishing operations automatically, in accordance with options specified by a user. For example, when a user is utilising the system to created an index document, the system provides the user with options for document finishing operations to be applied to items in the index document.

[0142] The system automatically applies the selected document finishing operations by allocating an electronic storage area, such as a portion of a hard drive, to a list of items to which document finishing operations are to be applied. For each item in the list, the system stores a copy of the item in the storage area. The system then processes thecopied items according to the specified document finishing operations For example, the system applies running numeric pagination to the items or resizes all items to A4.

[0143] In order to ensure that the original information is retained, the finishing operations are performed on a copy of the item, rather than the original item. That is, the system generates a copy of the item and then applies the selected document finishing operations to the copy. If edits are made to the Index (for example, if auto-pagination is enabled and an item is added into the middle of an index, then the pagination of following items may need to be recalculated), then some of these finished copies may become invalidated (because their pagination is now wrong). In this case, the invalidated copy is discarded and a new finished copy is regenerated (a new copy of the original is made and the correct pagination is applied to it). The index’s hyperlinks retrieve these finished copies, rather than the un-finished originals.

[0144] When a user subsequently accesses an item from the index document, the system retrieves the finished item for display. If any edits are made that may affect the finishing operations, such as a newly introduced document inserted half way through the index document, then the system automatically re-processes any affected items. This ensures that all documents are correctly finished. Thus, a user can specify a set of required document finishing operations for a set of items in a document list and the system not only applies the required finishing operations by ensures that the finishing operations are correctly applied in light of any subsequent modifications.Industrial Applicability

[0145] The arrangements described are applicable to the computing, legal and information technology industries.

[0146] Although the invention has been described with reference to specific examples, it will be appreciated by those skilled in the art that the invention may be embodied in many other forms. The foregoing describes only some embodiments of the present invention, and modifications and / or changes can be made thereto without departing from the scope and spirit of the invention, the embodiments being illustrative and not restrictive.

[0147] In the context of this specification, the word “comprising” and its associated grammatical constructions mean “including principally but not necessarily solely” or “having” or “including”, and not “consisting only of’. Variations of the word "comprising", such as “comprise” and “comprises” have correspondingly varied meanings.

[0148] As used throughout this specification, unless otherwise specified, the use of ordinal adjectives "first", "second", "third", “fourth”, etc., to describe common or relatedobjects, indicates that reference is being made to different instances of those common or related objects, and is not intended to imply that the objects so described must be provided or positioned in a given order or sequence, either temporally, spatially, in ranking, or in any other manner.

[0149] Reference throughout this specification to “one embodiment,” “an embodiment,” “some embodiments,” or “embodiments” means that a particular feature, structure or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment, but may. Furthermore, the particular features, structures or characteristics may be combined in any suitable manner, as would be apparent to one of ordinary skill in the art from this disclosure, in one or more embodiments.

[0150] While some embodiments described herein include some but not other features included in other embodiments, combinations of features of different embodiments are meant to be within the scope of the invention, and form different embodiments, as would be understood by those skilled in the art. For example, in the following claims, any of the claimed embodiments can be used in any combination.

[0151] Furthermore, some of the embodiments are described herein as a method or combination of elements of a method that can be implemented by a processor of a computer system or by other means of carrying out the function. Thus, a processor with the necessary instructions for carrying out such a method or element of a method forms a means for carrying out the method or element of a method. Furthermore, an element described herein of an apparatus embodiment is an example of a means for carrying out the function performed by the element for the purpose of carrying out the invention.

[0152] In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the invention may be practised without these specific details. In other instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure an understanding of this description.

[0153] Note that when a method is described that includes several elements, e.g., several steps, no ordering of such elements, e.g., of such steps is implied, unless specifically stated.

[0154] The term “coupled” should not be interpreted as being limitative to direct connections only. The terms “coupled” and “connected,” along with their derivatives, maybe used. It should be understood that these terms are not intended as synonyms for each other, but may be. Thus, the scope of the expression “a device A coupled to a device B” should not be limited to devices or systems wherein an input or output of device A is directly connected to an output or input of device B. It means that there exists a path between device A and device B which may be a path including other devices or means in between. Furthermore, “coupled to” does not imply direction. Hence, the expression “a device A is coupled to a device B” may be synonymous with the expression “a device B is coupled to a device A”. “Coupled” may mean that two or more elements are either in direct physical or electrical contact, or that two or more elements are not in direct contact with each other but yet still co-operate or interact with each other.

Claims

We claim:1 . An electronic document management system comprising: a co-ordinator; a set of request queues; a set of response queues; a set of microservices, each of which is configured to perform one of a set of predefined processing tasks on one of either a set of predefined item categories or a set of predefined item subtypes; wherein said request queues and said response queues manage communications between said co-ordinator and each of said microservices; and wherein said co-ordinator is configured to: create a project, based on user input; ingest data items by: identifying an item category of said data item; and assigning a set of said processing tasks to be performed on said data item, based on said identified item category; and add said ingested data items to said project.

2. The system according to claim 1 , wherein said system further includes an application server configured to: receiving operation requests from a user interface displayed by said system to a user computing device; and execute said received operation requests against said project.

3. The system according to claim 1 , wherein said application server is further configured to perform at least one of: user access management; account & project management; service discovery; and database operations.

4. The system according to any one of claims 1 to 3, wherein said5. The system according to any one of claims 1 to 4, wherein said co-ordinator is further configured to identify an item subtype of said data item, wherein: assigning said set of said processing tasks to be performed on said data item is based on said identified item category and said identified item subtype.

6. The system according to any one of claims 1 to 5, further comprising: a web server for serving content to a user computing device.

7. The system according to any one of claims 1 to 6, wherein each of said predefined item categories is associated with: a predefined set of metadata that can be extracted from that category; and a predefined set of processing tasks that can be performed on that category.

8. The system according to any one of claims 1 to 7, wherein said item categories are selected from the group consisting of: images, videos, documents, emails, spreadsheets.

9. The system according to any one of claims 1 to 8, wherein said item categories correspond to Multipurpose Internet Mail Extensions (MIME) types, each MIME type including a type and a subtype.

10. The system according to any one of claims 1 to 9, further comprising: a message broker adapted to manage sending and consuming of messages by said request queues and said response queues.11 . The system according to any one of claims 1 to 10, wherein each microservice has a corresponding request queue and a corresponding response queue.

12. The system according to any one of claims 1 to 11 , further comprising: a web server configured to store, process, and deliver web elements via a user interface presented to a browser executing on a user computing device via a communications network; and an application server configured to receive operation requests received from said user interface and execute said operation requests on said project.

13. The system according to any one of claims 1 to 12, wherein said co-ordinator identifies duplicate data items.

14. The system according to claim 13, wherein said co-ordinator utilises a cryptographic hash for each item family ingested to said project.

15. The system according to any one of claims 1 to 14, wherein said co-ordinator analyses property data associated with each data item to identify junk data items.

16. The system according to any one of claims 1 to 15, wherein said system is configured to process ingested data items in accordance with a predefined workflow.

17. The system according to any one of claims 1 to 16, wherein said co-ordinator extracts metadata associated with ingested data items to identify an item category of said data item.

18. The system according to any one of claims 1 to 17, wherein said set of processing tasks are selected from the group consisting of:• automatic decomposition of items• automatic extraction of metadata• automatic generation of Summary metadata• automatic extraction of text• automatic conversion of items to a common display format• automatic identification of junk items19. The system according to any one of claims 1 to 18, wherein said co-ordinator automatically decomposes data items during ingestion to process nested items as individual data items.

20. A method of managing a set of electronic documents utilising the electronic document management system of any one of claims 1 to 19.

Citation Information

Patent Citations

  • Electronic document content classification and document type determination

    US10699065B2

  • Collaborative matter management and analysis

    US20210065320A1

  • Document management system using multiple threaded processes and having asynchronous repository responses and no busy cursor

    US5544051A