Method, device, medium and product for processing legal documents
Patent Information
- Application Number
- CN202610894457.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-18
- Publication Date
- 2026-09-04
AI Technical Summary
[0003](1)处理效率极低:非结构化文档涵盖Word、扫描件、图片等多种格式,扫描件需人工转录、不同格式需手动适配,且合同类型和案件类型多样导致整理规则不统一,人工逐份处理需耗费大量时间,难以应对大量法务文档的处理需求,严重拖慢法务工作推进节奏
[0018]本申请通过所述的法务文档的处理方法,实现了法务文档的自动化结构化处理:接收待处理文档并获取元数据、将多源异构数据转换为统一格式并预处理、利用大模型按目录分块形成层级化的法务知识块、对长文本按语义进一步分块后保存。由此带来了显著的有益效果:处理效率大幅提升,彻底摆脱对人工逐份整理的依赖;通过标准化清理、统一目录解析,保障了数据结构化的高度准确性与一致性,避免人工主观误差;构建的层级目录与语义关联体系支持精准检索及文档联动分析,充分释放信息化系统的智能应用价值;依托元数据中的版本、时间信息实现文档修订追溯,结合安全管控满足法务工作的合规化要求;同时结构化知识块支持同类信息跨案件复用,并可灵活调整参数适配新业务类型,具备极强的数据复用性与业务扩展性,持续支撑法务工作数字化升级。
Smart Images

Figure CN122692014A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular to a method, device, medium, or product for processing legal documents. Background Technology
[0002] In current legal work, original legal documents exist in unstructured form and contain various different types. During the information technology transformation of legal work, legal document information relies on manual processing before structured information is input into the legal information system. This method has many drawbacks:
[0003] (1) Extremely low processing efficiency: Unstructured documents cover a variety of formats such as Word, scanned documents, and images. Scanned documents need to be transcribed manually, and different formats need to be adapted manually. In addition, the variety of contract types and case types leads to inconsistent sorting rules. Manual processing of each document takes a lot of time and is difficult to meet the processing needs of a large number of legal documents, which seriously slows down the pace of legal work.
[0004] (2) Insufficient data accuracy: Subjective errors are easily generated during manual sorting, such as typos in text entry, incorrect classification of case or contract content, incorrect calculation of case amount, and incorrect calculation of contract tax amount; moreover, it is difficult for manual sorting to guarantee the consistency of structured standards. Different sorting personnel have different classification logic and information extraction dimensions, which leads to data chaos in the information system and affects the accuracy of subsequent business judgment.
[0005] (3) Limited retrieval and linkage capabilities of information systems: manual sorting can only achieve simple information entry and does not build a unified association structure for information systems, resulting in the system being unable to support core functions such as accurate semantic retrieval and rapid location of key information. Retrieval requires browsing through historical documents or filtering data one by one; at the same time, fragmented data cannot achieve cross-document and cross-module information linkage, and the value of data analysis and intelligent application of information systems cannot be brought into play.
[0006] (4) Lack of traceability and compliance control of information system version: The manual transfer and data entry process lacks a system-level standardized approval and operation record mechanism, which does not meet the core requirements of legal information technology for traceability and compliance control of case and contract materials. Summary of the Invention
[0007] To avoid the problems of low efficiency, low data accuracy, insufficient information retrieval and linkage capabilities, and lack of version traceability and compliance control in manual legal document processing, one aspect of this application proposes a method for processing legal documents.
[0008] The method for processing legal documents is used in an electronic device and includes the following steps: receiving a document to be processed and obtaining its metadata, wherein the document to be processed is multi-source heterogeneous data related to legal matters; converting the document to be processed into a pixel matrix type document, establishing a bidirectional mapping index between the document to be processed and the pixel matrix type document, and preprocessing it to obtain an operable text sequence; extracting the directory of the text sequence through a large model, and dividing the text sequence into blocks based on the directory to form legal knowledge blocks corresponding to the directory, wherein the legal knowledge blocks include text blocks and directory hierarchy relationships; further dividing text blocks whose text length exceeds a preset value based on semantic boundaries; and saving the legal knowledge blocks, the directory hierarchy relationships, and the metadata of the document to be processed.
[0009] Optionally, the documents to be processed include case documents and contract documents, wherein the metadata of the case documents includes the case version, the case creation time, and the last revision time, and the metadata of the contract documents includes the application time, the person in charge, the handling unit, and the responsible party.
[0010] Optionally, it also includes adding a digital watermark to the pixel matrix document.
[0011] Optionally, the preprocessing includes: calling a visual text understanding model to decode and extract the image regions of the pixel matrix document into a text sequence; correcting skewed text in the text sequence; removing noise content in the text sequence using regular expressions or keyword matching; and initially classifying the text sequence into multiple discourse regions for subsequent targeted processing.
[0012] Optionally, the noise content includes: the time watermark, court signature, court seal, and noise information unrelated to the main content of the case document; and noise information unrelated to the main content of the contract document.
[0013] Optionally, the step of extracting the directory of the text sequence through a large model and dividing the text sequence into blocks based on the directory to form legal knowledge blocks corresponding to the directory includes: inputting system prompt words for extracting the directory of the text sequence, the main text of the text sequence, and the attachment content into the large model; wherein, the system prompt words are used to indicate that the directory level extracted by the large model is less than or equal to three levels; determining the index position of each directory through a retrieval matching algorithm, and dividing the text sequence into multiple text blocks based on the index position, and establishing an association between each text block and the corresponding directory; and for each directory, merging the directory, the hierarchical relationship of the directory, and the text block corresponding to the directory to form the legal knowledge block.
[0014] Optionally, the step of segmenting text blocks whose length exceeds a preset value based on semantic boundaries and forming new text blocks from texts with semantic similarity less than a threshold includes: decomposing the text block into sentence units using delimiters, and creating a context window for each sentence unit, the context window including the sentence unit and its context; vectorizing the context window to obtain semantic feature vectors, calculating the drift between adjacent semantic feature vectors, the drift indicating the magnitude of the difference between the semantic feature vectors; for multiple consecutive context windows, when the number of context windows with drift less than a preset semantic difference tolerance is less than a threshold, marking the positions exceeding the semantic difference tolerance as semantic boundary points, and reorganizing the sentence units between the semantic boundary points into semantically calibrated text blocks.
[0015] Another aspect of this application provides an electronic device comprising a memory storing computer-executable instructions and a processor; when the instructions are executed by the processor, the device causes the device to perform the processing method according to any of the preceding claims.
[0016] Another aspect of this application proposes a computer-readable storage medium storing one or more programs that can be executed by one or more processors to implement the processing method described in any of the preceding claims.
[0017] Another aspect of this application proposes a computer program product comprising a computer program that, when executed by a processor, implements the processing method as described in any of the preceding claims.
[0018] This application, through its method for processing legal documents, achieves automated and structured processing of legal documents: receiving documents to be processed and obtaining metadata, converting multi-source heterogeneous data into a unified format and preprocessing it, using a large model to form hierarchical legal knowledge blocks by directory segmentation, and further segmenting long texts by semantics before saving. This brings significant benefits: processing efficiency is greatly improved, completely eliminating the reliance on manual document-by-document processing; standardized cleaning and unified directory parsing ensure high accuracy and consistency of data structuring, avoiding subjective human error; the constructed hierarchical directory and semantic association system supports precise retrieval and document linkage analysis, fully releasing the intelligent application value of the information system; document revision traceability is achieved based on version and time information in the metadata, combined with security control to meet the compliance requirements of legal work; at the same time, the structured knowledge blocks support the reuse of similar information across cases and can flexibly adjust parameters to adapt to new business types, possessing strong data reusability and business scalability, continuously supporting the digital upgrade of legal work. Attached Figure Description
[0019] Figure 1 This is a flowchart of a legal document processing method according to an embodiment of this application.
[0020] Figure 2 This is a schematic diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0021] The following section, in conjunction with the accompanying drawings, explains the method for processing the legal documents submitted in this application.
[0022] Figure 1 This is a flowchart of a legal document processing method according to an embodiment of this application.
[0023] In step 101, the document to be processed is received and its metadata is obtained. The document to be processed is multi-source heterogeneous data related to legal affairs. In this embodiment, the document to be processed includes case documents such as litigation, arbitration, and labor disputes, and legal documents such as contract documents for procurement, sales, and services. At the same time, the business metadata of the document to be processed is collected, including the version identifier, creation timestamp, and last revision time of case documents, and the application time, handler, handling unit, and responsible party of contract documents, as the basic attribute dimensions for subsequent structured processing.
[0024] In step 102, the documents to be processed are uniformly converted into pixel matrix type documents, a bidirectional mapping index between the documents to be processed and the pixel matrix type documents is established, and preprocessing is performed to obtain an operable text sequence.
[0025] Since the original documents to be processed are multi-source heterogeneous data, containing various multimodal documents in different formats, it is necessary to convert them into a unified format before proceeding with other subsequent processing steps to facilitate the processing flow. In this embodiment, the documents to be processed are first stored in object storage, such as Amazon S3 (Amazon SimpleStorage Service). The LibreOffice document conversion engine then converts all input multimodal documents into pixel matrix documents and establishes a bidirectional mapping index relationship between the documents to be processed and the pixel matrix documents. It is understood that other types of storage services, storage methods, document conversion engines, and document conversion methods can also be used.
[0026] Furthermore, to enhance document data security, tools such as the ReportLab library can be used as needed to automatically add digital watermarks specific to legal documents to pixel matrix-type documents, thereby improving document security and traceability. Simultaneously, a strict multi-level approval mechanism is implemented within the document conversion and distribution process to ensure that all operations throughout the document's lifecycle—from warehousing, normalization, and hardening to distribution—are controlled and auditable.
[0027] Then, the pixel matrix-type documents are preprocessed. Specifically, for the converted pixel matrix-type documents, a visual text understanding model (such as PaddleOCR) is used to decode the image regions and extract them into text sequences. After extracting the text sequences, a unified text cleaning process is initiated: correcting skewed text in the text sequences, using regular expressions or keyword matching to remove noise content in the text sequences, and initially classifying the text sequences into multiple discourse areas for subsequent targeted processing. The noise content removed includes: time watermarks, court signatures, court seals, and noise information unrelated to the main content of the case documents (such as specific texts like "this document is identical to the original"); and noise information unrelated to the main content of the contract documents, such as logos, signatures, and total page numbers. The multiple discourse areas initially classified for case content include three parts: "Case Text," "Case Attachments," and "Case Forms." The multiple discourse areas initially classified for contracts include "Contract Text," "Contract Forms," "Signature Page," "Contract Attachments," and "Contract Storage Confidentiality Agreement," etc. The information in the forms, the contract signing page, the contract storage confidentiality agreement, and similar parts can be treated as a separate legal knowledge block, while the main text and attachments need to be further divided into multiple knowledge blocks.
[0028] It is understood that in other embodiments, the documents to be processed can also be uniformly converted into PDF format. For legal documents uniformly converted to PDF format, different processing methods can be adopted for different content formats, namely text format (e.g., usually converted from electronic documents) and image format (e.g., usually scanned copies). In one embodiment, for text format legal documents, the internal text stream can be directly read using a layout structure parser (e.g., pdfplumber); for image format legal documents, text imaging can be performed using a visual text understanding model to convert it into text data. After obtaining the text stream and text data, the above preprocessing methods can be applied for preprocessing.
[0029] After preprocessing, the case documents can be initially divided into three parts: "Case Text," "Case Attachments," and "Case Forms." The contract documents can be initially divided into several parts, including "Contract Text," "Contract Forms," "Signature Page," "Contract Attachments," and "Contract Storage Confidentiality Agreement." Among these, modules with fixed formats and standardized content, such as form information, contract signature pages, and contract storage confidentiality agreements, can be treated as independent knowledge blocks. The main text and attachments, however, need to be further divided into multiple knowledge blocks.
[0030] In step 103, the table of contents of the text sequence is extracted using a large model, and the text sequence is then divided into blocks based on the table of contents to form legal knowledge blocks corresponding to the table of contents. These legal knowledge blocks include the hierarchical relationship between the text blocks and the table of contents. This block division prevents key information from being overwhelmed by a large amount of irrelevant information when using legal document data for applications such as knowledge-based question answering, thereby enabling more accurate retrieval and comparison.
[0031] In one embodiment, an agent is first constructed for different types of cases and contracts. This agent can parse and extract information from a directory of documents to be processed. The agent can be built based on various artificial intelligence models, such as Deepseek-R1, Qwen3.5-Plus, and GLM-5. During use, system prompts, the main text and attachments included in the text sequence are input into the large model (i.e., the agent). The main text of the text sequence refers to the substantive clauses or statements necessary to achieve the core legal purpose of the case litigation or contract signing, forming the basis of the legal act. Attachments are physically or logically subordinate to the main text, and their content is directly related to the main purpose of the case or contract, supplementing, refining, or proving specific matters in the main text. The system prompts are used to instruct the large model to extract the directory of the text sequence. Since the extracted directory hierarchy is too deep, the subsequent segmented text blocks will lose meaning and information. Therefore, it is preferable to instruct the large model to extract a directory hierarchy of less than or equal to three levels. More preferably, it is preferable to instruct the large model to extract a three-level directory structure to ensure a balance between the information content of the text blocks and the granularity of the blocks.
[0032] Upon receiving the above input, the agent first extracts a directory structure of no more than three levels from the text sequence based on system prompts. For example, the extracted three-level directory structure might look like this: I. Rights and Obligations -> (I) Rights and Obligations of Party A -> 1. Details of Rights, where "I.", "(I)", and "1." represent different numbering formats for different levels. Then, through retrieval and matching, the agent locates the precise index position of the extracted directory titles in the original text. Based on these index positions, the main text and attachments are then segmented into multiple text blocks, and each text block is associated with its corresponding directory. Next, the agent parses the hierarchical relationships of each directory, constructs a tree-like hierarchical structure outline, and merges the directories, their hierarchical relationships (from the root directory to the current directory title), and the corresponding text blocks to form a legal knowledge block with a complete logical path. Therefore, each legal knowledge block includes a text block and its complete directory path, from which the complete semantics of each text block can be clearly determined, facilitating subsequent semantic segmentation; and improving the accuracy and recall of retrieval when applying legal knowledge blocks to knowledge question answering.
[0033] However, the text length of the text blocks obtained based on directory segmentation may still be too long (e.g., exceeding 2000 characters), so further segmentation is required. In step 104, text blocks whose text length exceeds a preset value are segmented based on semantic boundaries, and texts with semantic similarity less than a threshold are grouped into new text blocks.
[0034] Semantic segmentation builds upon the large model segmentation by further segmenting text blocks exceeding a preset length within the resulting legal knowledge block. In one embodiment, the text block is first broken down into sentence units using delimiters (e.g., periods, question marks, semicolons), each relatively short and independent. Then, a context window is created for each sentence unit, including the sentence unit itself and its preceding and following sentence units. The context window is then vectorized to obtain semantic feature vectors, for example, using a vectorization model like bge-m3 to convert it into a vector representation. Next, the drift (vector space angle) between adjacent semantic feature vectors is calculated. The drift indicates the magnitude of the difference between semantic feature vectors; a larger drift indicates a smaller semantic similarity between the semantic feature vectors and a greater core semantic difference between adjacent context windows.
[0035] Then, a semantic difference tolerance is preset to indicate the upper limit of context window drift. If the drift exceeds the tolerance, the semantic difference is considered large enough to warrant segmentation. For multiple consecutive context windows, if the number of context windows with drift less than the preset tolerance exceeds a threshold (e.g., 75% of context windows have drift less than the tolerance), the semantic difference between these consecutive context windows is considered small, and they are semantically continuous, belonging to the same text block. If the number of context windows with drift less than the tolerance is less than the threshold, the semantic difference between these consecutive context windows is considered large, and these context windows should be semantically divided into different text blocks. In one embodiment, locations exceeding the tolerance are marked as semantic boundary points, indicating a semantic shift. Since the semantic meaning of texts between different semantic boundary points is similar, sentence units between these boundary points are reorganized into semantically calibrated text blocks. The semantically calibrated text blocks inherit the complete hierarchical directory path of the original legal knowledge blocks, ensuring that each calibrated text block is both semantically cohesive and retains its original logical structure.
[0036] By using semantic segmentation, text blocks can be further segmented while improving the semantic similarity of text content within the same text block, thereby further improving the accuracy of retrieval and question answering.
[0037] In step 105, the legal knowledge block, the hierarchical relationship of the directory, and the metadata of the documents to be processed are saved.
[0038] For legal knowledge blocks after semantic segmentation, the text blocks within them can be further reviewed. If a text block is too short, it can be merged into an adjacent text block to avoid low information density and missing effective information. When saving data, it can be included in different ways according to data type. In one embodiment, the metadata, directory, and segmentation results of the document to be processed are integrated into JSON data format and saved to Amazon S3; information expected to be used for retrieval, such as directories with multi-level relationships, contextual segmentation results, and case summaries or contract summaries extracted from large models, is formed into sparse and dense vector data through the bge-m3 vectorization model; and basic data such as document type and name are stored in the Milvus vector database.
[0039] The above process allows for the complete processing and storage of documents, enabling the processed data to be applied to various downstream tasks.
[0040] Furthermore, the above-described process can be executed through a system. This system will set up scheduled tasks to periodically and continuously perform structured processing and backup of relevant legal document databases (such as case document databases and contract document databases) using the document processing methods described above. For urgent or immediate data, documents and metadata can be provided semi-automatically or automatically to directly call the system to update the database. For example, when modifying cases or contracts, such as revising case descriptions or contract clauses, the updated data will not directly overwrite the original data, but will instead create a new version of the legal knowledge block. Creating a new version of the legal knowledge block specifically includes steps such as regenerating structured text, adding corresponding update records, and creating associations with the old version, so that historical revision records of cases or contracts can be queried in chronological order when needed.
[0041] The aforementioned methods and systems can be uniformly integrated into a legal affairs platform, enabling seamless collaboration and efficient linkage between the new information-based legal document processing methods and the existing manual document processing methods. Continuous monitoring of the system's operational status ensures the long-term reliable operation of the new information-based methods and their compatibility with the existing manual methods, thereby providing users with a smooth and high-quality user experience.
[0042] Now for reference Figure 2 The diagram shown is a block diagram of an electronic device 200 according to an embodiment of the present application. The electronic device 200 may include one or more processors 202, system control logic 208 connected to at least one of the processors 202, system memory 204 connected to the system control logic 208, non-volatile memory (NVM) 206 connected to the system control logic 208, and network interface 210 connected to the system control logic 208.
[0043] Processor 202 may include one or more single-core or multi-core processors. Processor 202 may include any combination of general-purpose processors and special-purpose processors (e.g., graphics processors, application processors, baseband processors, etc.). In embodiments herein, processor 202 may be configured to perform one or more embodiments of various embodiments of the legal document processing methods proposed in this application.
[0044] In some embodiments, system control logic 208 may include any suitable interface controller to provide any suitable interface to at least one of the processors 202 and / or any suitable device or component communicating with system control logic 208.
[0045] In some embodiments, system control logic 208 may include one or more memory controllers to provide an interface to system memory 204. System memory 204 may be used to load and store data and / or instructions. In some embodiments, system memory 204 of electronic device 200 may include any suitable volatile memory, such as suitable dynamic random access memory (DRAM).
[0046] The nonvolatile memory 206 may include one or more tangible, non-transitory computer-readable storage media for storing data and / or instructions. In some embodiments, the nonvolatile memory 206 may include any suitable nonvolatile memory such as flash memory and / or any suitable nonvolatile storage device, such as at least one of HDD (Hard Disk Drive), CD (Compact Disc) drive, and DVD (Digital Versatile Disc) drive.
[0047] The non-volatile memory 206 may include a portion of the storage resources installed on the device of the electronic device 200, or it may be accessible by the device, but is not necessarily part of the device. For example, the non-volatile memory 206 may be accessed over a network via the network interface 210.
[0048] Specifically, system memory 204 and non-volatile memory 206 may each include a temporary copy and a permanent copy of instructions 220. Instructions 220 may include instructions that, when executed by at least one of processors 202, cause electronic device 200 to implement the methods provided in this application. In some embodiments, instructions 220, hardware, firmware, and / or their software components may additionally / alternatively reside in system control logic 208, network interface 210, and / or processor 202.
[0049] In some embodiments, network interface 210 may be integrated into other components of electronic device 200. For example, network interface 210 may be integrated into at least one of processor 202, system memory 204, non-volatile memory 206, and firmware device (not shown) having instructions, which, when executed by at least one of the processor 202, enable electronic device 200 to implement one or more embodiments of the various embodiments described herein. Network interface 210 may further include any suitable hardware and / or firmware to provide a multiple-input multiple-output radio interface.
[0050] In one embodiment, at least one of the processors 202 may be packaged together with the logic of one or more controllers for system control logic 208 to form a system package (SiP). In another embodiment, at least one of the processors 202 may be integrated on the same die with the logic of one or more controllers for system control logic 208 to form a system on chip (SoC).
[0051] Electronic device 200 may further include: input / output (I / O) device 212. Input / output (I / O) device 212 may include a user interface that enables a user to interact with electronic device 200; the design of peripheral component interfaces enables peripheral components to also interact with electronic device 200.
[0052] In some embodiments, the user interface may include, but is not limited to, a display (e.g., a liquid crystal display, a touch screen display, etc.), a speaker, a microphone, one or more cameras (e.g., a still image camera and / or a video camera), a flashlight (e.g., a light-emitting diode flash), and a keyboard.
[0053] In some embodiments, the peripheral component interface may include, but is not limited to, a non-volatile memory port, an audio jack, and a power interface.
[0054] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 200. In other embodiments of this application, the electronic device 200 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0055] Program code can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, the processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application-specific integrated circuit (ASIC), or a microprocessor.
[0056] The program code can be implemented using a high-level procedural language or an object-oriented programming language to communicate with the processing system. Assembly language or machine language can also be used when needed. In fact, the mechanisms described in this paper are not limited to any particular programming language. In either case, the language can be a compiled language or an interpreted language.
[0057] One or more aspects of at least one embodiment can be implemented by representational instructions stored on a computer-readable storage medium, the instructions representing various logics in a processor, which, when read by a machine, cause the machine to create logic for performing the techniques described herein. These representations, referred to as “IP cores,” can be stored on a tangible computer-readable storage medium and provided to multiple customers or production facilities for loading into manufacturing machines that actually manufacture the logic or processor.
[0058] One embodiment of this application discloses a computer-readable storage medium storing one or more programs executable by one or more processors to implement the legal document processing method of this application.
[0059] One embodiment of this application discloses a computer program product, including a computer program that, when executed by a processor, implements the method for processing legal documents of this application.
[0060] The specific embodiments described above illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. Although the description of this application is presented in conjunction with preferred embodiments, this does not mean that the features of this invention are limited to these embodiments. Furthermore, to avoid confusion or obscuring the focus of this application, some specific details will be omitted in the description. It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other.
[0061] Furthermore, the various operations will be described as multiple discrete operations in a manner most conducive to understanding the illustrative embodiments; however, the order of description should not be construed as implying that these operations must depend on the order. In particular, these operations do not need to be performed in the order presented.
[0062] Unless the context otherwise specifies, the terms “contains,” “has,” and “includes” are synonyms. The phrase “A / B” means “A or B.” The phrase “A and / or B” means “(A and B) or (A or B).”
[0063] As used herein, the terms “module” or “unit” may refer to, be, or include: application-specific integrated circuits (ASICs), electronic circuits, (shared, dedicated, or group) processors and / or memories that execute one or more software or firmware programs, combinational logic circuits, and / or other suitable components that provide the described functionality.
[0064] In the accompanying drawings, certain structural or methodological features are shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be necessary. In some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Furthermore, the inclusion of structural or methodological features in a particular figure does not imply that such features are required in all embodiments, and in some embodiments, these features may be omitted or may be combined with other features.
Claims
1. A method for processing legal documents, used in electronic devices, characterized in that, include: Receive the document to be processed and obtain the metadata of the document to be processed, wherein the document to be processed is multi-source heterogeneous data related to legal affairs; The documents to be processed are uniformly converted into pixel matrix type documents, a bidirectional mapping index between the documents to be processed and the pixel matrix type documents is established, and preprocessing is performed to obtain an operable text sequence. The text sequence is extracted using a large model, and the text sequence is divided into blocks based on the directory to form legal knowledge blocks corresponding to the directory. The legal knowledge blocks include text blocks and directory hierarchy relationships. as well as Based on semantic boundaries, text blocks whose length exceeds a preset value are further divided into blocks; Save the legal knowledge block, the hierarchical relationship of the directory, and the metadata of the document to be processed.
2. The processing method according to claim 1, characterized in that, The documents to be processed include case documents and contract documents. The metadata of the case documents includes the case version, the case creation time, and the last revision time. The metadata of the contract documents includes the application time, the person in charge, the handling unit, and the responsible party.
3. The processing method according to claim 2, characterized in that, It also includes adding a digital watermark to the pixel matrix document.
4. The processing method according to claim 3, characterized in that, The preprocessing includes: The visual text understanding model is invoked to decode and extract the image regions of the pixel matrix document into a text sequence. Correcting slanted text in the text sequence by using regular expressions or keyword matching to remove noise content from the text sequence; and The text sequence is initially categorized into multiple discourse regions for subsequent targeted processing.
5. The processing method according to claim 4, characterized in that, The noise content includes: The case documents contain time watermarks, court signatures, court seals, and noise information unrelated to the main content of the case. The contract document contains noise information unrelated to the main content of the contract.
6. The processing method according to claim 1, characterized in that, The step of extracting a directory from the text sequence using a large model and then dividing the text sequence into blocks based on the directory to form legal knowledge blocks corresponding to the directory includes: The system prompt word for the directory used to extract the text sequence, the main text of the text sequence, and the attachment content are input into the large model; wherein, the system prompt word is used to indicate that the large model extracts a directory level of three or less when extracting the directory; By using a retrieval matching algorithm, the index position of each directory is determined, and the text sequence is segmented into multiple text blocks based on the index positions. Each text block is then associated with its corresponding directory. For each directory, the directory, the hierarchical relationship of the directory, and the text block corresponding to the directory are merged to form the legal knowledge block.
7. The processing method according to claim 6, characterized in that, The step of segmenting text blocks whose length exceeds a preset value based on semantic boundaries, and forming new text blocks from texts with semantic similarity less than a threshold, includes: The text block is broken down into sentence units using delimiters, and a context window is created for each sentence unit, the context window including the sentence unit and the context of the sentence unit; The context window is vectorized to obtain semantic feature vectors, and the drift between adjacent semantic feature vectors is calculated, the drift indicating the magnitude of the difference between the semantic feature vectors; For multiple consecutive context windows, when the number of context windows with drift amounts less than a preset semantic difference tolerance is less than a threshold, the positions exceeding the semantic difference tolerance are marked as semantic boundary points, and the sentence units between the semantic boundary points are reorganized into semantically calibrated text blocks.
8. An electronic device, characterized in that, The device includes a memory storing computer-executable instructions and a processor; when the instructions are executed by the processor, the device performs the processing method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, which can be executed by one or more processors to implement the processing method according to any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the processing method as described in any one of claims 1 to 7.