Multi-modal case analysis intelligent generation system and method oriented to lawyer of litigation
The intelligent generation system for multimodal case analysis automates the processing of multimodal case materials, solving the problem of low efficiency in existing technologies and enabling efficient and standardized generation of litigation documents, thereby improving lawyers' case-handling efficiency and document quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI SHENGXI FASO ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
- Filing Date
- 2026-01-23
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies cannot efficiently process multimodal case materials, resulting in low efficiency for lawyers in information integration and document drafting, and the generated documents lack accuracy and logical rigor, making it impossible to achieve an integrated automated process.
Design a multimodal case analysis intelligent generation system, including modules such as multimodal data acquisition, preprocessing, structured representation, case element extraction, claim basis analysis, evidence association and document generation. Combined with a predefined intelligent workflow, it automatically processes multimodal materials and generates litigation documents that meet the requirements of judicial practice.
It significantly improved the efficiency of case information input, realized the unified processing and integrated management of multimodal materials, and generated documents that conform to judicial norms, reduced operational complexity, and improved case handling efficiency and document quality.
Smart Images

Figure CN122020530A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer-based judicial practice technology, specifically to a multimodal case analysis intelligent generation system and method for litigation lawyers. Background Technology
[0002] With the increasing complexity of socio-economic activities, the number of civil and commercial litigation cases continues to rise. The lengthening of case fact chains, the abundance of evidence, and the disorganized nature of parties' oral accounts place immense pressure on frontline litigation lawyers in terms of information integration and legal document drafting. Under current technological conditions, lawyers primarily rely on the following methods to handle cases, but all have limitations: Currently, lawyers need to obtain oral information through phone calls and face-to-face meetings, and then manually compile fragmented materials from multiple sources such as WeChat, emails, contracts, and financial documents to form a timeline, case summary, and evidence list. This method is highly dependent on the lawyer's personal experience and energy, is inefficient and time-consuming, and is prone to omissions of key information or misunderstandings due to human factors.
[0003] Furthermore, some legal information service providers currently offer database search functions for laws, regulations, and judgments, as well as standardized document templates. Lawyers can manually rewrite and splice the search results. However, such tools are essentially still in the rudimentary stage of "search-template filling." They cannot automatically extract core elements of a case, conduct in-depth analysis of the basis of a claim, or automatically link evidence and facts based on specific, unstructured case materials. They also cannot achieve the coherent and integrated generation of multiple types of litigation documents.
[0004] Furthermore, with the development of artificial intelligence technology, general-purpose large language models (LLM) have begun to be used to assist in document drafting. Lawyers can input case details in natural language to obtain preliminary drafts. However, general-purpose models cannot directly process multimodal raw materials such as audio, scanned documents, and images, resulting in a high input threshold. They lack a precise grasp of the case elements, evidence chain logic, and factual relationships unique to the legal field, leading to insufficient factual accuracy and logical rigor in the generated content. They also lack in-depth customization tailored to the characteristics of litigation procedures, court document format norms, and legal terminology habits, often resulting in discrepancies between the generated results and practical requirements. The lack of a unified structured case knowledge representation as an intermediate layer leads to the fragmentation of the "case analysis - evidence organization - document generation" process, making it impossible to form a traceable and editable integrated workflow.
[0005] In summary, existing technologies still force lawyers to spend a significant amount of time on repetitive basic organization and document polishing when handling complex litigation cases, resulting in low overall case-handling efficiency and difficulty in standardizing and scaling service quality. Therefore, there is an urgent need for a systematic technical solution that can automatically extract elements from multimodal materials, construct a unified structured case model, and intelligently generate a series of documents and analysis reports in accordance with Chinese litigation practice procedures to overcome the above-mentioned shortcomings. To this end, a multimodal case analysis intelligent generation system and method for litigation lawyers is proposed. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention provides a multimodal case analysis intelligent generation system and method for litigation lawyers, thereby resolving the problems mentioned in the background.
[0007] To achieve the above objectives, the present invention provides the following technical solution: Firstly, a multimodal case analysis and intelligent generation system for litigation lawyers includes: The multimodal case data acquisition module is used to receive raw materials in various forms related to the target case; The multimodal preprocessing and unified representation module is connected to the multimodal case data acquisition module and is used to perform speech recognition, optical character recognition and document parsing on the received raw materials, and to uniformly store the processed text and associated metadata into a preset case raw corpus. The case element extraction module is connected to the original case corpus and is used to automatically extract case elements from the original case corpus based on a pre-trained language model and domain rules, and store the extraction results in a structured form in the case structure database; the case elements include at least party information, case type, litigation claims, key facts and timeline, and points of contention; The module for analyzing the basis of claims and applicable law is connected to the case structure database. It is used to automatically identify the legal relationship corresponding to the claim and match legal clauses based on the elements in the case structure database and a pre-set legal knowledge base, and generate a mapping relationship between the claim, legal relationship, legal clauses and key facts. The evidence association and evidence list generation module is connected to the case structured database. It is used to identify the types of evidence materials in the case and extract key points, establish an association model between facts, evidence and litigation claims, and generate a standardized evidence list based on the association model, mark evidence gaps and provide supplementary evidence collection prompts. The document generation and multi-document type arrangement module is connected to the case structured database, the claim basis and legal application analysis module, and the evidence association and evidence list generation module, respectively. It is used to automatically generate corresponding litigation document drafts based on the case structured database, the mapping relationship, and the association model, according to the selected document type. The interactive editing and traceable display module is connected to the document generation and multi-document type arrangement module. It provides a visual interface to display and support the modification of the generated document drafts, and synchronously updates the modified content to the case structured database.
[0008] Preferably, the raw materials received by the multimodal case data acquisition module include at least one of the following: audio files, electronic documents, image materials, and structured data.
[0009] Preferably, the process of generating litigation documents by the document generation and multi-document type arrangement module adopts a combination of template constraints and intelligent generation, including: loading a structured format template according to the selected document type; filling the framework of the template with the content of the case structured database; and calling a domain-tuned language model to polish and logically optimize the filled natural language part.
[0010] Preferably, it further includes: The quality assessment and iterative optimization module is connected to the document generation and multi-document type arrangement module. It is used to automatically verify the generated document drafts according to preset rules and mark or trigger corrections for parts that do not meet the requirements.
[0011] Preferably, the system predefines a mapping relationship between document types and analysis modules; when a user selects a target document type in the interface of the interactive editing and traceable display module, the system automatically triggers and calls one or more pre-analysis modules required to generate the target document type to perform analysis tasks according to the mapping relationship.
[0012] Secondly, a multimodal case analysis intelligent generation method for litigation lawyers, based on the intelligent generation method of the system described in the first aspect, includes the following steps: S1: Obtain multimodal input materials of the target case, and convert them into unified text and metadata form through multimodal preprocessing and unified representation processing, and store them in the original case corpus; S2: Extract case elements from the original case corpus to form structured case data and store it in the case structure database; S3: Based on the aforementioned case structure database, conduct analysis of the basis of the claim and the applicable law, and generate a mapping relationship between the litigation claim, legal relationship, legal clause and key facts; S4: Process the evidence materials in the case, establish a correlation model between facts, evidence and litigation claims, and generate a standardized evidence list; S5: In response to the user's selection of the target document type, automatically generate the corresponding draft litigation document based on the case structure database, the mapping relationship, and the association model; and S6: Provides an interactive editing interface, responds to user modifications to the document draft, and synchronously updates the modified content to the case structured database.
[0013] Preferably, in response to the user's selection of the target document type, a corresponding draft litigation document is automatically generated, specifically including: Based on the predefined mapping relationship between document types and analysis modules, determine one or more pre-analysis modules that need to be called to generate the selected target document type; The system automatically triggers and executes each of the determined pre-analysis modules in sequence, generating corresponding intermediate analysis results; the intermediate analysis results include at least one or more of the following: case summary, case timeline, claim analysis report, and evidence list; Based on the structured database of the cases and the intermediate analysis results, a draft document of the target document type is generated.
[0014] Preferably, the pre-analysis module includes multiple modules such as a case summary generation module, a case timeline review module, a claim basis pre-analysis module, a legal rationality review module, a cause of action determination module, a defense analysis module, and an evidence list review module.
[0015] Preferably, the specific process of generating a draft litigation document in step S5 includes: Load the structured format template corresponding to the selected target document type; The contents of the case structure database are populated into the corresponding frame of the template; The domain-adjusted language model is invoked to generate or optimize the natural language portion of the filled template, forming a document draft.
[0016] Preferably, after step S6, the method further includes: S7: Automatically validates the format and logical consistency of the final generated document and alerts you to potential defects.
[0017] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention possesses multimodal integrated processing capabilities, enabling simultaneous processing of heterogeneous original case materials from multiple sources, such as audio, images, scanned documents, and various electronic documents. Through built-in speech recognition, optical character recognition, and document parsing technologies, it automatically converts these materials into unified and processable text data, thereby significantly reducing the large amount of manual work that lawyers spend on material transcription, format conversion, and preliminary organization, and improving the efficiency and convenience of case information input.
[0018] 2. This invention constructs a unified structured case knowledge representation model integrating "parties, legal relationships, litigation claims, key facts, and evidence materials." By standardizing, associating, storing, and managing the extracted case elements, it provides a solid and consistent data foundation for subsequent case analysis, ensuring the interconnectivity and traceability of data across the entire chain from claim analysis and evidence list compilation to the generation of various documents, thus solving the problem of information fragmentation in traditional methods.
[0019] 3. The document generation capability of this invention is deeply aligned with the specific requirements of Chinese judicial practice. It adopts a technical approach that combines "structured template constraints" with "domain-specific intelligent generation." While ensuring that the generated complaint, answer, appeal, and other litigation documents meet the format specifications and column requirements stipulated by the court, it uses a language model finely tuned in the legal field to polish the content and organize the logic, making the language of the final documents more professional and the logic more rigorous, thus significantly improving the practical usability and professionalism of the documents.
[0020] 4. This invention achieves automated workflow scheduling based on user intent by predefining the intelligent mapping relationship between "document type - analysis module". When a lawyer only needs to select the target document type, the system can automatically trigger and run all the necessary pre-analysis modules (such as case summary generation, claim analysis, evidence sorting, etc.), without requiring the user to manually configure and execute each item, greatly reducing operational complexity and error risk, and making the entire process from case analysis to document output more coherent, efficient and automated.
[0021] 5. This invention supports the automation of multi-task, multi-stage workflows covering the entire litigation cycle. The system can guide and automatically complete a series of tasks from case material entry, element extraction, legal analysis, evidence organization to the generation of various litigation documents, providing lawyers with an end-to-end integrated solution. This not only significantly improves the overall efficiency of handling complex cases, but also effectively reduces repetitive work between different stages, and enhances the standardization and scalability of legal service workflows.
[0022] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of the overall system structure of the present invention; Figure 2 This is a schematic diagram of the multimodal preprocessing and unified representation process of the present invention; Figure 3 This is a schematic diagram of the workflow execution based on intelligent linkage according to the present invention; Figure 4 This is a schematic diagram illustrating the overall process and interactive optimization of the method of the present invention. Detailed Implementation
[0024] Please see Figures 1-4 This invention discloses a multimodal case analysis intelligent generation system and method for litigation lawyers. It aims to provide litigation lawyers with an integrated solution encompassing automated case material organization, intelligent legal analysis, and standardized document generation by integrating multimodal information processing, artificial intelligence analysis, and legal knowledge. The method and its corresponding system can automatically receive and process heterogeneous case materials from multiple sources, including audio, images, and electronic documents. Through intelligent extraction and modeling, it constructs a unified, structured case knowledge system. Based on a predefined intelligent workflow, after the lawyer simply selects the document type, the system automatically schedules and completes all correlation analysis tasks, ultimately generating various litigation documents that meet the requirements of judicial practice. The technical solution, implementation process, and beneficial effects of this invention will be described in detail below with reference to specific embodiments.
[0025] Example: Intelligent system and process for handling shareholder qualification confirmation disputes based on cloud architecture I. System Deployment and Initialization The system of this invention is deployed using a cloud-based distributed architecture based on microservices to ensure high availability, elastic scaling, and secure isolation. The specific deployment environment and initial configuration are as follows: 1. Hardware and Infrastructure Layer: The system is deployed in the data center of a cloud service provider. The underlying infrastructure includes: a virtual machine cluster for hosting stateless application services; computing nodes with GPU acceleration capabilities for running artificial intelligence models; object storage services for storing massive amounts of unstructured data (such as audio and images); and relational database clusters and graph database instances for storing structured data.
[0026] 2. Software service layer: Application services are orchestrated and deployed in a containerized manner, mainly including: API Gateway and Service Registry: Serving as the unified entry point for all external requests, it is responsible for routing, authentication, and load balancing.
[0027] Multimodal Acquisition and Preprocessing Service Group: This is a file upload and management service that implements the functionality of a multimodal case data collection module. The service provides a RESTful API interface and WebSocket connection, allowing lawyers to upload audio files (such as MP3 and WAV), electronic documents (such as PDF and DOCX), image files (such as JPG and PNG), and structured data files (such as CSV and XLSX) via a web interface or client. The service scans uploaded files for viruses, validates their format, and assigns them globally unique identifiers.
[0028] A speech recognition service that integrates a speech recognition engine specifically designed for Chinese legal scenarios. This engine employs an end-to-end deep learning model, and its training data includes a large amount of transcribed text from legal consultations and court hearing recordings to improve the accuracy of recognizing legal terminology and entities such as personal names and company names.
[0029] An optical character recognition and document parsing service that implements OCR and document parsing functions in a multimodal preprocessing and unified representation module. For images and scanned PDFs, it employs an OCR pipeline that integrates text detection and recognition models, with optimizations specifically for printed text, handwritten text, and tables. For native PDF and DOCX documents, a dedicated parser is used to extract text content, font styles, paragraph structure, table data, and metadata (such as author and creation date).
[0030] A unified representation service is responsible for receiving text results from the aforementioned processing services. It normalizes all text content, along with its source file ID, position information within the source file (such as page numbers and line numbers), processing timestamps, and other associated metadata, according to a predefined JSON Schema format, and then calls a storage interface to persistently write it to the original case corpus. The corpus is implemented using a document database that supports full-text search (such as Elasticsearch) to facilitate subsequent rapid retrieval and analysis.
[0031] Case Intelligent Analysis Service Group: Case Element Extraction Service: This service loads a large language model (e.g., a Transformer-based model) pre-trained and fine-tuned on a general corpus and millions of judgments and legal documents. Upon receiving an analysis request, the service reads the corpus of the specified case from the "raw case corpus." The extraction process combines model prediction with rule post-processing. 1) Named Entity Recognition: The model recognizes entities in the text, such as "Company A" (ORG), "B" (PERSON), "January 1, 2020" (DATE), and "500,000 yuan" (MONEY).
[0032] 2) Relation Extraction and Element Classification: The model performs semantic understanding of sentences and combines it with predefined rule templates (e.g., matching the sentence structure "request...order..." to identify the litigation claim) to classify entities and text fragments into preset element categories. For example, from the sentence "Plaintiff B requests confirmation that he holds 15% of the equity of Company A", the litigation claim is extracted as follows: {Subject: "B", Content: "confirmation of holding 15% of the equity of Company A", Share: "15%"}; from multiple sentences containing dates, a timeline is formed by arranging them in chronological order: [{Time: "2020-01-01", Event: "signing of nominee shareholding agreement"}, ...].
[0033] 3) All extracted elements are converted into structured data objects (e.g., Python dictionaries or Java objects) and written to a case structured database. This database uses a relational database and is designed with standardized table structures to store tables such as "Parties Table," "Petition Table," "Factual Node Table," and "Points of Contention Table," and establishes relationships through foreign keys.
[0034] Claim Analysis Service: This service maintains an internal legal knowledge base, which stores structured legal provisions such as the Civil Code, Company Law, and Civil Procedure Law, along with their revision history, relevant judicial interpretations, and key judgment points extracted from precedents. The service receives a list of "Case Types" (e.g., "Shareholder Qualification Confirmation Disputes") and "Litigation Claims" from the "Case Structure Database." The analysis process is as follows: 1) Legal Relationship Identification: Based on case type and request keywords, a pre-defined legal relationship graph is matched. For example, "shareholder qualification confirmation" is associated with the "shareholder qualification confirmation dispute" subclass under "corporate legal relationship".
[0035] 2) Legal Provision Matching and Mapping: In the legal knowledge base, semantic similarity calculations (such as vector-based retrieval) and keyword indexing are used to find the legal provisions most relevant to the current request and legal relationship. For example, for a request to "confirm shareholding," Article 32 of the Company Law regarding the validity of shareholder register entries is matched. The system generates a mapping table, recording each claim ID, the corresponding legal relationship ID, and the supporting legal provision ID (multiple are possible), and links it to the list of "key facts" IDs supporting the request in the "case structured database." This mapping table is also stored in the structured database.
[0036] Evidence Association Service: This service processes uploaded original files. First, it pre-determines the evidence type of the file (documentary evidence, physical evidence, audiovisual materials, electronic data, etc.). Second, it extracts key evidence points by analyzing the file content (using the element extraction service) and filename. The core step is building an association model: 1) For each key fact registered in the “Case Structure Database”, the service attempts to find textual evidence in the evidence content that supports or mentions that fact.
[0037] 2) Establish a relational triple: (Fact F, Evidence E, Proof Relation R). For example, (Fact: "B invested 500,000 yuan in a certain month of a certain year", Evidence: "Screenshot of bank transfer voucher", Relation: "Direct proof").
[0038] 3) All related triples are linked to the claims to form a three-dimensional network of "facts-evidence-claims". This network can be stored in a graph database for easy complex association queries and reasoning.
[0039] 4) Based on this network, an "Evidence List" form conforming to the court's format requirements is automatically generated, listing the evidence number, name, source, page number, and object of proof (i.e., a summary of the related facts). Simultaneously, the system runs an "Evidence Gap Analysis" sub-process: checking whether each key fact, especially those involving amounts, time, and identity verification, is supported by sufficiently strong evidence. If a fact is found to be supported only by the party's statement, lacking documentary evidence or third-party evidence, it is marked "Evidence Needs Supplementation" on that fact and the evidence list, and a prompt can be generated according to preset rules, such as "It is recommended to supplement the original bank transfer record for the fact of 'investing 500,000 yuan in May 2021'."
[0040] Document generation service group: Document formatting and generation service: This document generation and multi-document type formatting module is the core output unit of this invention. Its workflow strictly follows a model that combines template constraints with intelligent generation.
[0041] 1) Template Loading: The service has a built-in "document template library" where each type of document (such as a civil complaint, answer, or application for property preservation) has a strictly corresponding structured template. The templates are defined in XML or JSON format, specifying the document's chapter structure (such as "preamble," "claims," "facts and reasons," and "closing"), the required fields under each chapter (such as "plaintiff's name," "defendant's name," and "cause of action"), and the optional fields.
[0042] 2) Data Population: When a document needs to be generated, the service loads the corresponding template based on the document type. Then, as a "fill-in" process, it accurately extracts data from the "Case Structured Database," "Mapping Relationship Table," and "Evidence Association Network" to fill in the template fields. For example, the party information is filled in the "Header," and the list of claims is filled in the "Plaintiff" section.
[0043] 3) Intelligent Polishing: For natural language sections requiring coherent statements, such as the "Facts and Reasons" section, the service invokes a large, domain-tuned language model. The model's input is a carefully crafted prompt containing: a chronologically ordered list of key facts extracted from a database, a summary of relevant legal provisions, and key evidentiary points to emphasize. The model's task is to generate a logically clear and legally grammatically correct argument based on this information. The generation process is constrained to ensure that facts are not fabricated and that the text does not deviate from legal grounds. The generated text is then inserted into the appropriate locations in the template.
[0044] 4) Format Combination: Finally, the service combines all the filled and generated content into a standard document file (such as DOCX format) according to the format defined in the template (font, font size, paragraph spacing, header and footer, etc.).
[0045] Workflow orchestration engine: This is a control system that maintains a configuration file that defines the mapping relationship between "document types" and "required analysis modules". For example, the mapping relationship is defined as follows: "Civil complaint" needs to call the case element extraction module, the claim basis and applicable law analysis module, and the evidence association and evidence list generation module, and can selectively trigger tasks to generate reports such as case summary and case timeline; Application for evidence preservation: element extraction, evidence correlation.
[0046] When a user selects only to generate a civil complaint on the interface, the workflow engine is triggered. It automatically parses the mapping relationships and then sequentially and asynchronously calls the case element extraction service (implementing the case element extraction module function), the claim analysis service (implementing the claim basis and applicable law analysis module function), and the evidence association service (implementing the evidence association and evidence list generation module function), and may trigger a subtask to generate an intermediate report (such as a case summary). After all the preliminary tasks are completed, the engine automatically calls the document formatting and generation service to produce the final draft. This process is transparent to the user; the user only needs to see a progress bar and the final result, without having to manually execute each analysis step.
[0047] User interaction and quality assurance services: Interactive Editing Service: This service powers the front-end web interface. The interface clearly displays the generated document draft. Lawyers can directly edit the text on the interface. The service records all differences in modifications. More importantly, when a lawyer modifies content in the document that originates from the structured database (e.g., correcting the date of a fact), the service attempts to back-synchronize this modification back to the case's structured database. For example, through natural language understanding, it identifies the original data ID corresponding to the modified text fragment and then updates the record in the database. This ensures the consistency of the data source.
[0048] Quality assessment service: This service runs automatically after the document is generated or edited. It loads a series of preset rules for verification. Format compliance check: Verify that the document contains all the required fields as specified in the template.
[0049] Logical consistency check: For example, check whether the amount mentioned in the claim has a corresponding calculation basis or factual support in the facts and reasons section; check whether the evidence cited exists in the list of evidence.
[0050] Completeness check of elements: ensure that the information of the parties involved is complete and the cause of action is clear.
[0051] For issues detected (such as missing client addresses), the service generates specific annotations and prompts, which are then fed back to the front-end interface to guide lawyers in making corrections. The service also possesses basic model self-evaluation capabilities; for example, it can calculate the perplexity of generated fact and reason paragraphs or compare them with the distribution of training data, marking paragraphs with low confidence.
[0052] 3. Data storage layer initialization: Original case corpus: As mentioned earlier, an index is created using an Elasticsearch cluster. The index mapping is predefined to store fields such as text content, file metadata, and case ID.
[0053] Case structured database: Using a database such as MySQL or PostgreSQL, all data tables are pre-created according to the relational model of the system design.
[0054] Legal knowledge base: This uses a combination of graph databases (such as Neo4j) and relational databases for storage. The graph database stores the relationships between legal concepts, legal provisions, and judicial interpretations; the relational database stores the full text of the legal provisions and their metadata.
[0055] Document Template Library: Stores template definitions for various documents in the form of files or a dedicated configuration database.
[0056] Workflow configuration library: Stores configuration information such as the mapping relationship between "document type - analysis module".
[0057] System log library: centrally records all user operations, service calls, and error messages for auditing and iterative optimization.
[0058] II. Specific Implementation Steps and Processes (Taking a Shareholder Qualification Confirmation Dispute Case as an Example) Suppose lawyer A is handling a shareholder qualification dispute case. The following is a complete and detailed operational and technical implementation process of their use of the system of this invention: Step 0: Case Creation and Material Upload 1. Lawyer A logs into the system's web interface, creates a new case, and temporarily selects "Shareholder Qualification Confirmation Dispute" as the cause of action.
[0059] 2. Attorney A submitted the following materials in batches on the case workbench by dragging and dropping or clicking to upload: consultation_20231015.mp3 (A consultation recording with client B, 30 minutes in length).
[0060] articles_of_association.pdf (Scanned copy of Company A's articles of association, 10 pages).
[0061] share_transfer_agreement.pdf (scanned copy of the shareholding agreement between B and C, 5 pages).
[0062] meeting_minutes_photo1.jpg, photo2.jpg (Photos of the minutes of the two shareholders' meetings).
[0063] wechat_chat_screenshots.zip (contains multiple screenshots of WeChat chat history).
[0064] company_info.xlsx (Company A's registration information table exported from the business registration system).
[0065] 3. The multimodal case data acquisition module (file upload service) receives these files, performs security checks, generates a unique ID for each file (such as file_001, file_002, etc.), and temporarily stores the file binary stream in object storage. Subsequently, the system asynchronously triggers the preprocessing process.
[0066] Step 1: Multimodal Preprocessing and Unified Representation 1. Speech Recognition: The speech recognition service reads the file "consultation_20231015.mp3" from object storage, performs preprocessing such as noise reduction and sentence segmentation, and then inputs it into the acoustic and language models for decoding, outputting the Chinese transcribed text. The model has been specifically optimized for legal terms such as "nominal shareholder," "actual investor," and "defective capital contribution," resulting in improved accuracy.
[0067] 2. OCR and Document Parsing: The OCR service processes .jpg images and scanned pages from PDFs, performs text detection and recognition, and outputs the text content and text coordinates for each page.
[0068] The document parsing service processes the original articles_of_association.pdf and share_transfer_agreement.pdf, extracting plain text, identifying heading levels, and extracting the shareholder list and investment amount from the table.
[0069] For company_info.xlsx, directly parse its cell data.
[0070] 3. Unified Representation and Storage: The unified representation service receives all the processing results described above. It organizes all text content, along with its source file identifier, location information within the source file, processing timestamps, and other associated metadata, according to a structured data format (e.g., a format can be defined for each text unit, including fields such as case identifier, content, and metadata). For example, a data unit can be represented as follows: case identifier "case_20240001", document identifier "file_002", content "Article 3: The shareholders' meeting is composed of all shareholders…", and metadata records information such as the source file and page number. All standardized data units are batch-indexed into the original case corpus, and an inverted index is built for full-text retrieval. Thus, the scattered and heterogeneous materials are transformed into a unified, machine-readable text corpus.
[0071] Step Two: Case Element Extraction and Structured Storage 1. Lawyer A clicks "Start Intelligent Analysis" on the interface, or the system automatically triggers it after preprocessing is complete.
[0072] 2. The case element extraction service is invoked, with the parameter case_id: “case_20240001”. The service queries all data related to this case from the corpus.
[0073] Named entity recognition and classification: The model analyzes the corpus sentence by sentence.
[0074] From the transcript of the audio recording, “I am B, and I have entrusted C to hold 15% of the shares of Company A on my behalf,” we can identify “B” (PERSON), “C” (PERSON), “Company A” (ORG), and “15%” (PERCENT).
[0075] Based on the rules, "B" is classified as the "plaintiff", while "C" and "Company A" are classified as the "defendants".
[0076] Claim Extraction: The model identifies sentences like "My request is to confirm my shareholder status and have my name added to the shareholder register," and extracts structured claims through sequence labeling and intent classification. Type: Declaratory judgment action; Target: Shareholder status; Subject: B (the party in question); Object: Company A; Share: 15%; type: payment claim, action: change of registration, target: shareholder register.
[0077] Timeline and Fact Building: The model identifies date entities and event descriptions from all the corpus, sorts them by time, and builds a fact chain. Date: 2020-01-01; Event: B and C signed the "Equity Holding Agreement"; Date: 2021-05-10, Event: B transferred 500,000 yuan to C's account as an investment; Date: August 20, 2022; Event: B issued a "Letter of Request for Disclosure" to Company A and other shareholders; Date: 2022-09-05, Event: Company A responded in writing, refusing to process the roster change for Mr. B.
[0078] Summary of points of contention: Based on the analysis of the facts and the statements of both parties, the model summarizes the points of contention, such as "whether the nominee shareholding agreement is valid", "whether B's capital contribution has actually been paid in" and "whether more than half of the other shareholders of Company A have agreed".
[0079] Data persistence: All extracted structured information is converted into a relational data model and written into the case structured database. For example, three records are inserted into the parties table; four records are inserted into the facts table, each fact having a unique ID; two records are inserted into the claims table, and a link is established between the claims and relevant fact IDs through a claims-fact association table.
[0080] Step 3: Analysis of the Basis of the Claim and Applicable Law The claim analysis service is invoked to read case information from the structured database.
[0081] Legal Relationship Positioning: Based on the cause of action of "shareholder qualification confirmation dispute" and the requests for "confirmation of equity" and "change of registration", the service locates the path "company dispute → shareholder rights dispute → shareholder qualification confirmation dispute" in the internal knowledge graph.
[0082] Legal provision retrieval and matching: The service uses keywords such as "actual investor," "shareholder register," and "change registration" to conduct semantic searches in the legal knowledge base. Search results may include Article 32 of the Company Law and Article 24 of the Judicial Interpretation (III) of the Company Law. The system calculates the semantic similarity between each legal provision and the description of the current case, and ranks them according to relevance.
[0083] Generate mapping relationships: The service creates and stores a mapping relationship table. For example:
[0084] Step 4: Evidence Association and Inventory Generation Technical implementation: The evidence association service is invoked.
[0085] Evidence Identification and Numbering: The service lists all uploaded files, performs preliminary classification based on file type and content, and assigns evidence numbers: E001 (transcribed audio recordings), E002 (company articles of association), E003 (shareholding agreement), E004 / E005 (shareholders' meeting minutes photos), E006 series (WeChat screenshots), E007 (company information sheet).
[0086] Association Model Construction: The service analyzes the content of each piece of evidence (feature extraction models can be invoked to assist in understanding) and matches it with records in the "Fact Table".
[0087] The content of E003 (the nominee agreement) is directly matched with Fact1 (the signing of the agreement), establishing a connection (Fact1, E003, "direct proof").
[0088] E004 (Minutes of the First Shareholders' Meeting) mentions "deliberation on B's capital contribution", which is related to Fact2 (transfer of capital contribution), establishing a connection (Fact2, E004, "indirect evidence").
[0089] A screenshot in E006 shows that C replied "I agree to let you show your name", which is related to Fact3 (making a request) and the other party's possible attitude, establishing a connection (Fact3, E006 - specific screenshot, "proving that the other party had agreed").
[0090] Evidence Gap Analysis and Hints: The system analyzes the "fact-evidence" correlation model and finds that the key fact "Fact2 (B transferred 500,000 yuan on May 10, 2021)" lacks direct evidence such as bank transfer vouchers.
[0091] After constructing the association model, the evidence association service executes an automated evidence strength analysis sub-process: the system checks whether each key fact is associated with sufficiently strong evidence (such as direct proof or original vouchers). If it finds that the existing evidence for a certain fact (such as Fact2) is weak or lacks key evidence, the system will: 1) Mark the "Evidence to be supplemented" status in the view associated with the facts and the generated list of evidence; 2) Based on a pre-defined rule base, it automatically generates and outputs supplementary evidence prompts, such as "To prove... facts, it is recommended to supplement with the original bank transfer voucher." This achieves intelligent identification and auxiliary prompts for evidence gaps.
[0092] Step 5: Document generation based on intelligent linkage (core automation) After completing steps two (element extraction), three (claim analysis), and four (evidence association), the system has constructed a complete structured case model of 'parties-legal relationship-claim-facts-evidence' and generated all necessary analysis results and mapping relationships. At this point: User action: Lawyer A selects the target document type as 'civil complaint' in the system interface and triggers the generation command.
[0093] Workflow engine trigger: User operation events are captured by the workflow orchestration engine. The engine queries the mapping relationship between predefined document types and analysis modules in its configuration library and finds that the prerequisite task list corresponding to "civil complaint" is: case element extraction module, claim basis and legal application analysis module, evidence association and evidence list generation module, as well as optional report generation modules such as case summary generation module and case timeline sorting module.
[0094] Task orchestration and execution: The engine begins asynchronously scheduling these tasks. Since the analyses in steps two through four may already be complete (results stored in the database), the engine will directly read existing results from the cache or database. If some analyses are not yet complete, the engine will immediately invoke the corresponding microservices to execute them. For report generation tasks such as generating case summaries and timelines, the engine will invoke specialized services to create a concise case summary and a visual timeline based on existing data. All these backend interactions are completely transparent to Attorney A; he only sees a progress bar on the interface: "Preparing materials required for the complaint… Generating the complaint…".
[0095] Document content generation: Once all previous tasks are marked as completed, the engine calls the document formatting and generation service.
[0096] Template loading: The service loads the "Civil Complaint" template.
[0097] Data population: For the parties involved: Read the plaintiff and defendant information from the party table in the database and fill it into the template.
[0098] Claims section: Read two claims from the claims form and format them in legal document language, such as: 1. Confirm that Plaintiff B holds a 15% equity stake in Company A; 2. The court orders Defendant Company A to complete the shareholder register change procedures for Plaintiff B within ten days of the judgment taking effect.
[0099] Intelligent Fact and Reason Writing: Based on the results generated from the aforementioned steps—namely, extracting elements and facts from the structured case database, obtaining legal basis from the mapping table, and determining evidentiary relationships from the evidence association network—the service integrates all information to construct a structured prompt input to a large-scale language model fine-tuned in the legal domain. This prompt integrates a chronologically ordered list of key facts, claims, summaries of relevant legal provisions, and key evidentiary points. Based on this context, the model generates (or refines and optimizes) a logically clear and legally grammatically correct "fact and reason" argument, for example: Background: A dispute over the confirmation of shareholder qualifications.
[0100] Plaintiff: Mr. B. Defendants: Company A and Mr. C.
[0101] The core facts are listed in chronological order: [Fact1 description, Fact2 description, Fact3 description, Fact4 description].
[0102] Plaintiff's claims: [Description of Req1, description of Req2].
[0103] Legal basis: mainly involves Article 32 of the Company Law (validity of the shareholder register) and Article 24 of the Judicial Interpretation III of the Company Law (conditions for the actual investor to be named).
[0104] Key points of evidence: Fact 1 is directly proven by E003; Fact 2 is indirectly corroborated by the party's statement and E004, but the direct transfer voucher needs to be supplemented; Fact 3 is proven by E006 to show that the other party had given consent.
[0105] The key issues in dispute are: the validity of the nominee shareholding agreement, the status of capital contribution, and the consent of other shareholders.
[0106] Based on the above information, please draft the "Facts and Reasons" section of the civil complaint, ensuring it is logically clear, well-founded, and written in legal language.
[0107] The system invokes a language model finely tuned for the legal field. Based on this prompt, it generates (or refines and optimizes) a coherent argument, for example: Plaintiff B and Defendant C signed an "Equity Holding Agreement" on January 1, 2020, stipulating that C would hold 15% of the equity of Company A on behalf of B (Evidence E003). On May 10, 2021, B paid C RMB 500,000 as agreed (see attached evidence list, bank vouchers to be supplemented). On August 20, 2022… Company A refused to pay on September 5 without cause. The plaintiff argues… According to Article 32 of the "Company Law of the People's Republic of China"… and relevant judicial interpretations… therefore, this lawsuit is filed with your court…
[0108] Evidence citation insertion: The system will automatically associate the facts mentioned in the generated argument with the corresponding evidence number and indicate them in the text or at the end of the text in the form of (see evidence E003 for details).
[0109] Format assembly: After combining all parts, the final civil complaint draft file is generated according to the requirements of court document format.
[0110] Results presentation: The generated documents are presented to Attorney A for preview on the interface. At the same time, the automatically generated "Case Summary Report", "Case Timeline", "Claim Analysis Report" and "List of Evidence" are also available as attachments for viewing.
[0111] Step Six: Interactive Editing and Data Synchronization 1. Attorney A reviewed the draft complaint. He found that the statement of the contribution date in the Facts and Reasons section could be more precise, changing May 10, 2021 to 3:32 PM on May 10, 2021.
[0112] 2. The interactive editing service records this modification. It not only updates the text in the DOCX file but also attempts to understand the changes. Through parsing, the service identifies that the modified text fragment corresponds to the original description of the date field in the record with ID Fact2 in the database fact table.
[0113] Data write-back: The service sends an update request to the "Case Structured Database," updating the event field of the Fact2 record to "B transferred 500,000 yuan to C's account as investment at 15:32 on May 10, 2021." This step synchronously updates the modified content to the case structured database, ensuring the accuracy of the data source. Subsequently, if a defense statement or pleadings are generated based on this case, the description of the investment time will automatically use the updated, precise time.
[0114] Step Seven: Quality Assessment and Finalization After Attorney A confirmed the editing was complete and clicked "Final Check," the quality assessment service was triggered.
[0115] Rule verification: Service operation verification rules.
[0116] Check the format: Ensure that the beginning and end of the document are complete.
[0117] Inspection logic: Verify whether the "15% equity" in the "claims" is supported by corresponding evidence in the "facts and reasons" (yes, the nominee agreement and capital contribution are mentioned).
[0118] Inspection elements: It was found that the method of bearing the case acceptance fee was not clearly stated in the "claims" section (a common format requirement).
[0119] The system highlights the "Claims of Litigation" section of the document and provides the following prompt: Please supplement your request regarding the allocation of case acceptance fees in accordance with the "Measures for Payment of Litigation Costs," for example: The litigation costs in this case shall be borne by the defendant.
[0120] Attorney A supplemented the request according to the prompts. The system performed a quick verification again, and after passing the verification, generated the final version of the "Civil Complaint".
[0121] All operation logs, model-generated content, manual modification records, and verification results are recorded in the system log library for subsequent analysis, model retraining, and system iterative optimization.
Claims
1. A multimodal case analysis and intelligent generation system for litigation lawyers, characterized in that, include: The multimodal case data acquisition module is used to receive raw materials in various forms related to the target case; The multimodal preprocessing and unified representation module is connected to the multimodal case data acquisition module and is used to perform speech recognition, optical character recognition and document parsing on the received raw materials, and to uniformly store the processed text and associated metadata into a preset case raw corpus. The case element extraction module is connected to the original case corpus and is used to automatically extract case elements from the original case corpus based on a pre-trained language model and domain rules, and store the extraction results in a structured form in the case structure database; the case elements include at least party information, case type, litigation claims, key facts and timeline, and points of contention; The module for analyzing the basis of claims and applicable law is connected to the case structure database. It is used to automatically identify the legal relationship corresponding to the claim and match legal clauses based on the elements in the case structure database and a pre-set legal knowledge base, and generate a mapping relationship between the claim, legal relationship, legal clauses and key facts. The evidence association and evidence list generation module is connected to the case structured database. It is used to identify the types of evidence materials in the case and extract key points, establish an association model between facts, evidence and litigation claims, and generate a standardized evidence list based on the association model, mark evidence gaps and provide supplementary evidence collection prompts. The document generation and multi-document type arrangement module is connected to the case structured database, the claim basis and legal application analysis module, and the evidence association and evidence list generation module, respectively. It is used to automatically generate corresponding litigation document drafts based on the case structured database, the mapping relationship, and the association model, according to the selected document type. The interactive editing and traceable display module is connected to the document generation and multi-document type arrangement module. It provides a visual interface to display and support the modification of the generated document drafts, and synchronously updates the modified content to the case structured database.
2. The intelligent generation system for multimodal case analysis for litigation lawyers according to claim 1, characterized in that, The raw materials received by the multimodal case data acquisition module include at least one of the following: audio files, electronic documents, image materials, and structured data.
3. The intelligent generation system for multimodal case analysis for litigation lawyers according to claim 1, characterized in that, The process of generating litigation documents by the document generation and multi-document type arrangement module adopts a combination of template constraints and intelligent generation, including: loading a structured format template according to the selected document type; filling the framework of the template with the content of the case structured database; and calling a domain-fine-tuned language model to polish and optimize the logic of the filled natural language part.
4. The intelligent generation system for multimodal case analysis for litigation lawyers according to claim 1, characterized in that, Also includes: The quality assessment and iterative optimization module is connected to the document generation and multi-document type arrangement module. It is used to automatically verify the generated document drafts according to preset rules and mark or trigger corrections for parts that do not meet the requirements.
5. The intelligent generation system for multimodal case analysis for litigation lawyers according to claim 1, characterized in that, The system predefines a mapping relationship between document types and analysis modules. When a user selects a target document type in the interface of the interactive editing and traceable display module, the system automatically triggers and calls one or more pre-analysis modules required to generate the target document type to perform analysis tasks according to the mapping relationship.
6. A multimodal case analysis intelligent generation method for litigation lawyers, characterized in that, The intelligent generation method based on the system described in any one of claims 1-5 includes the following steps: S1: Obtain multimodal input materials of the target case, and convert them into unified text and metadata form through multimodal preprocessing and unified representation processing, and store them in the original case corpus; S2: Extract case elements from the original case corpus to form structured case data and store it in the case structure database; S3: Based on the aforementioned case structure database, conduct analysis of the basis of the claim and the applicable law, and generate a mapping relationship between the litigation claim, legal relationship, legal clause and key facts; S4: Process the evidence materials in the case, establish a correlation model between facts, evidence and litigation claims, and generate a standardized evidence list; S5: In response to the user's selection of the target document type, automatically generate the corresponding draft litigation document based on the case structure database, the mapping relationship, and the association model; and S6: Provides an interactive editing interface, responds to user modifications to the document draft, and synchronously updates the modified content to the case structured database.
7. The intelligent generation method for multimodal case analysis for litigation lawyers according to claim 6, characterized in that, Upon responding to the user's selection of the target document type, the system automatically generates the corresponding draft litigation document, which includes: Based on the predefined mapping relationship between document types and analysis modules, determine one or more pre-analysis modules that need to be called to generate the selected target document type; The system automatically triggers and executes each of the determined pre-analysis modules in sequence, generating corresponding intermediate analysis results; the intermediate analysis results include at least one or more of the following: case summary, case timeline, claim analysis report, and evidence list; Based on the structured database of the cases and the intermediate analysis results, a draft document of the target document type is generated.
8. The intelligent generation method for multimodal case analysis for litigation lawyers according to claim 6, characterized in that, The pre-analysis module includes several modules such as: case summary generation module, case timeline sorting module, claim basis pre-analysis module, legal rationality review module, cause of action determination module, defense analysis module, and evidence list sorting module.
9. The intelligent generation method for multimodal case analysis for litigation lawyers according to claim 6, characterized in that, The specific process of generating a draft litigation document in step S5 includes: Load the structured format template corresponding to the selected target document type; The contents of the case structure database are populated into the corresponding frame of the template; The domain-adjusted language model is invoked to generate or optimize the natural language portion of the filled template, forming a document draft.
10. The intelligent generation method for multimodal case analysis for litigation lawyers according to claim 6, characterized in that, Following step S6, the following is also included: S7: Automatically validates the format and logical consistency of the final generated document and alerts you to potential defects.