Judicial document vectorization segmentation and paragraph labeling method based on rule base

By building a rule base and using regular expressions to match various parts of judicial documents, combined with vectorization processing, the problems of low efficiency, poor accuracy and lack of standardization in judicial documents are solved, efficient and accurate judicial documents are achieved, and the ability of the large language model to understand judicial documents is improved.

CN120197623APending Publication Date: 2025-06-24CHONGQING XINZHI SMOOTH TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510272959.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing technology has problems such as low efficiency, poor accuracy and lack of standardization in the annotation of judicial documents. Especially when large language models assist in the generation of judicial documents, how to accurately understand the structure, rules and judgment logic of judicial documents is also a difficult problem.

Method used

The vectorized segmentation and paragraph annotation method of judicial documents based on the rule base is adopted. By constructing a rule base containing segmentation rules and paragraph annotation rules, using regular expressions to match various parts in the judicial documents, and vectorized them to improve the accuracy and efficiency of the labeling.

Benefits of technology

This method can significantly improve the accuracy and efficiency of judicial instrument labeling, ensure that each paragraph is correctly classified and marked, reduce labor burden, improve work efficiency, and improve the problem of inconsistency in annotation caused by differences in personal understanding, and improve the ability of large language models to understand judicial instruments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197623A_ABST
    Figure CN120197623A_ABST
Patent Text Reader

Abstract

The invention discloses a judicial document vectorization segmentation and paragraph labeling method based on a rule base. The method comprises the steps that S1, the rule base containing segmentation rules and paragraph labeling rules is constructed; s2, obtaining a judicial document to be labeled; s3, matching a corresponding segmentation rule and a paragraph labeling rule for each document part in the preprocessed judicial document through a regular expression; s4, segmenting the judicial document according to a matching result of the segmentation rule to obtain a plurality of document paragraphs; s5, performing paragraph labeling on each document paragraph according to a matching result of the paragraph labeling rule to obtain labeling information of each document paragraph; and S6, vectorizing all the document paragraphs with the annotation information, and outputting the vectorized document paragraphs as annotation results of the judicial document. According to the judicial document vectorized by the method, the understanding of the large language model on the meaning and judicial judgment logic can be improved, and the accuracy and quality of generating the judicial document under the assistance of the large language model artificial intelligence technology are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence-assisted generation of judicial documents, and in particular to a rule-based judicial document vectorization segmentation and paragraph annotation method. Background Art

[0002] In the judicial field, judicial documents, as an important carrier of legal procedures and legal decision-making, bear important responsibilities such as recording case facts, explaining legal reasons, and making legal decisions. With the rapid development of information technology, the electronic and information processing of judicial documents has become an important means to improve judicial efficiency and ensure judicial fairness. However, how to efficiently and accurately annotate judicial documents (such as the cause of the case, trial area, case category, number of trials, litigation requests, judgment results, etc.) has always been a difficult problem that has plagued judicial organs and legal researchers.

[0003] Judicial documents usually contain a large amount of information, such as case facts, legal clauses, and rulings, which are often presented in different paragraphs. Segmenting and paragraph annotation of judicial documents is the basis for extracting, analyzing, and utilizing this information. Through segmentation, judicial documents can be divided into different logical units, each of which corresponds to a specific information point or legal concept, which helps users quickly grasp the main content of the document. Paragraph annotation further clarifies the specific attributes and categories of each paragraph, such as the cause of the case, case summary, trial area, case category, number of trials, etc., providing rich label information for subsequent text analysis, data mining, etc.

[0004] In addition, segmenting and annotating judicial documents can also help improve judicial transparency and credibility. Through open and transparent segmentation and annotation results, the public can more clearly understand the content and structure of judicial documents, thereby enhancing their trust and recognition of judicial decisions.

[0005] Although the annotation of judicial documents is of great significance, there are still many problems with the existing annotation methods. First of all, the manual annotation method is time-consuming and laborious, with low efficiency. Due to the large number of judicial documents and the different content and structure of each document, manual annotation requires a large amount of time and effort. This not only increases the workload of judicial organs but also may lead to inconsistencies and subjectivity in the annotation results. Secondly, most of the existing automated annotation methods are based on keyword matching or simple text classification algorithms, making it difficult to accurately identify the complex structure and semantic information in judicial documents. In addition, the existing annotation methods lack unified standards and norms. Different judicial organs or research institutions may adopt different annotation rules and methods, resulting in inconsistencies in the annotation results and difficulty in comparison. This not only affects the accuracy and reliability of the annotation results but also hinders the sharing and exchange of judicial document information. Therefore, how to design a method that can improve the accuracy, efficiency, and standardization of judicial document annotation is a technical problem that urgently needs to be solved.

[0006] Especially at present when large language models are widely used to assist in generating various types of judicial documents, how to enable large language models to accurately understand the structure, rules, meanings, and judicial judgment logic of electronic judicial documents, so as to improve the accuracy and quality of using large language model artificial intelligence technology to assist in generating judicial documents, is a technical problem that urgently needs to be solved. Summary of the Invention

[0007] In view of the above deficiencies of the prior art, the technical problem to be solved by the present invention is: how to provide a method for vectorizing and segmenting judicial documents and paragraph annotation based on a rule library, segmenting judicial documents through the segmentation rules in the rule library to make the structure of the entire judicial document clearer, and at the same time enabling each paragraph to carry corresponding annotation information through the paragraph annotation rules in the rule library, so as to improve the accuracy, efficiency, and standardization of judicial document annotation, facilitate the understanding of the meaning and judicial judgment logic of judicial documents vectorized by this method by large language models, and improve the accuracy and quality of using large language model artificial intelligence technology to assist in generating judicial documents.

[0008] To solve the above technical problems, the present invention adopts the following technical solutions:

[0009] A method for vectorizing and segmenting judicial documents and paragraph annotation based on a rule library, comprising:

[0010] S1: Construct a rule library containing segmentation rules and paragraph annotation rules;

[0011] S2: Obtain the judicial document to be annotated and preprocess the judicial document;

[0012] S3: Match the corresponding segmentation rules and paragraph annotation rules for each part of the preprocessed judicial document through regular expressions;

[0013] S4: Segment the judicial document according to the matching result of the segmentation rule to obtain a number of document paragraphs;

[0014] S5: Perform paragraph annotation on each document paragraph according to the matching result of the paragraph annotation rule to obtain the annotation information of each document paragraph;

[0015] S6: After vectorizing all the document paragraphs with annotation information, output them as the annotation result of the judicial document.

[0016] Preferably, in step S1, the rule categories of the segmentation rule and the paragraph annotation rule in the rule library include one or more of: structural rules, semantic rules, and format rules;

[0017] (1) Structural rules: used to determine the basic segmentation structure of the judicial document and the logical relationship between its various parts;

[0018] (2) Semantic rules: used to understand the information content in each paragraph or sentence and associate it with specific legal concepts or entities;

[0019] (3) Format rules: used to ensure that the format of the judicial document complies with specific legal or institutional requirements.

[0020] Preferably, in step S1, the rule categories of the segmentation rule and the paragraph annotation rule in the rule library include the synonym rule: used to map one or more synonyms to standard vocabulary;

[0021] Construct the synonym rule through the following steps:

[0022] S101: Obtain the pre-trained BERT model on a large-scale corpus;

[0023] S102: Convert the vocabulary in the judicial document sample into word vectors through the BERT model;

[0024] S103: Calculate the cosine similarity between word vectors, and take the vocabulary corresponding to the word vectors with cosine similarity exceeding the threshold as synonyms;

[0025] S104: Store all the synonyms in the synonym library;

[0026] S105: Generate the corresponding synonym rule for each standard vocabulary according to the synonym library.

[0027] Preferably, in step S1, construct a decision tree for identifying document categories and determining the priority of rules;

[0028] Construct the decision tree through the following steps:

[0029] S111: Build a corresponding rule library for each type of legal document in advance, including segmentation rules and paragraph annotation rules;

[0030] S112: Obtain and preprocess the labeled legal document samples;

[0031] S113: Extract features from the preprocessed legal document samples as the input of the decision tree; Train the decision tree through machine learning algorithms;

[0032] S114: Identify the document type of the legal document through the trained decision tree;

[0033] S115: Select the corresponding rule library for the legal document according to the identified document type to perform the rule matching work in step S3;

[0034] S116: In steps S4 and S5, when there are multiple rules applicable to the same situation, the decision tree determines the execution order of each rule through a preset priority logic.

[0035] Preferably, in step S2, the preprocessing of the legal document includes one or more of data cleaning, text normalization, text segmentation, stop word removal, and data desensitization processing.

[0036] Preferably, in step S2, the preprocessing includes using a synonym rule for synonym replacement: If the current word is a synonym in the synonym library, replace it with the corresponding standard word.

[0037] Preferably, in step S3, use regular expressions to identify the parts of the legal document that conform to the patterns in the rule library, and then match the corresponding segmentation rules and paragraph annotation rules for each document part in the legal document;

[0038] The specific processing steps are as follows:

[0039] S301: Determine the patterns and labels of each rule in the rule library;

[0040] S302: Write the corresponding regular expressions for the patterns of each rule in the rule library;

[0041] S303: Traverse the regular expressions of each rule in the rule library, use the regular expression matching function to match the legal document. If the match is successful, record the matched document part and the corresponding rule label;

[0042] S304: Establish a mapping relationship between each document part in the legal document and the corresponding rule according to the matching result, and associate each text part with the corresponding rule.

[0043] Preferably, in step S3, the regular expression is trained and optimized through a machine learning algorithm;

[0044] The specific processing steps are as follows:

[0045] S311: Use the regular expression to extract the text part from the judicial document samples; divide the judicial document samples with the extraction results into a training set and a test set;

[0046] S312: Select a machine learning model as the optimization model;

[0047] S313: Convert the judicial document samples and regular expressions in the training set into feature vectors as the model input, and use the extraction results as the model output to train and optimize the model; test the performance of the optimized model through the test set;

[0048] S314: Use the trained optimized model to predict the judicial documents and output the corresponding predicted extraction results;

[0049] S315: Obtain the actual extraction results of the regular expression for the judicial documents; compare the predicted extraction results with the actual extraction results, and optimize the regular expression based on the differences in the comparison results.

[0050] Preferably, in step S6, the processing steps for vectorizing the document paragraphs are as follows:

[0051] S601: Obtain a vector model and fine-tune the vector model through the text data in the judicial field;

[0052] S602: After preprocessing and tokenizing the document paragraphs with annotation information, convert the document paragraphs into an input format that can be understood by the vector model;

[0053] S603: Input the converted document paragraphs into the fine-tuned vector model for vectorization and output the document paragraph vectors;

[0054] S604: Post-process the document paragraph vectors to make them meet the preset standards and requirements;

[0055] S605: Store and output the post-processed document paragraph vectors and their corresponding annotation information.

[0056] Preferably, in step S5, the annotation information of the document paragraphs includes one or more of the case cause, case summary, trial area, case category, number of trials, litigation request, factual reasons, defendant's defense, determined facts, court's opinion, legal basis, and judgment result.

[0057] Compared with the prior art, the method for vectorizing and segmenting judicial documents and paragraph annotation based on a rule base in the present invention has the following

[0058] Beneficial effects:

[0059] By constructing a rule library containing detailed segmentation rules and paragraph annotation rules, the present invention can accurately match each part in judicial documents, such as the cause of action, case summary, etc., avoiding errors that may occur during manual annotation. At the same time, the use of regular expressions further enhances the accuracy of rule matching, ensuring that each paragraph can be correctly classified and annotated, thereby improving the accuracy of judicial document annotation. And compared with the traditional manual annotation method, the automatic annotation method of the present invention can quickly process a large number of judicial documents, reduce the manual burden, and improve work efficiency.

[0060] During the annotation process of the present invention, the judicial document is first segmented according to the segmentation rules, making the structure of the entire judicial document clearer, facilitating users to quickly grasp the main content of the document, and each paragraph having a clear theme, such as the cause of action, case summary, etc., helping readers to form a well-organized cognitive framework when reading. At the same time, after the judicial document is segmented, each paragraph is provided with corresponding annotation information, such as case category, number of trials, etc., which are very important reference bases for users and can help users find the required information faster, thereby improving the efficiency of reading judicial documents. In addition, by using a unified rule library and regular expressions for segmentation and annotation, the present invention can ensure the unity and standardization of judicial document processing, improve the problem of inconsistent annotation caused by individual understanding differences, and thus improve the standardization of judicial document annotation. The large language model can more accurately understand the structure, rules, meanings, and judicial judgment logic of the electronic judicial documents vectorized by the present invention, thereby improving the accuracy and quality of the judicial documents assisted by the large language model artificial intelligence technology. Brief description of the drawings

[0061] In order to make the objectives, technical solutions, and advantages of the invention clearer, the present invention will be further described in detail below with reference to the drawings, where:

[0062] Figure 1 It is a logic block diagram of a method for vectorized segmentation and paragraph annotation of judicial documents based on a rule library. Detailed implementation manners

[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention generally described and illustrated in the figures herein can be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0064] It should be noted that like reference numerals and letters denote like items in the following figures, and thus, once an item is defined in one figure, it need not be further defined and explained in subsequent figures. In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the inventive product is customarily placed during use. This is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention. In addition, the terms "first", "second", "third", etc. are only used for descriptive distinction and should not be construed as indicating or implying relative importance. In addition, terms such as "horizontal" and "vertical" do not mean that the components are required to be absolutely horizontal or hanging, but can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but can be slightly inclined. In the description of the present invention, it should also be noted that unless otherwise clearly specified and limited, the terms "set", "installed", "connected", "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0065] The following is a more detailed description through specific embodiments:

[0066] Embodiment:

[0067] In this embodiment, a method for vectorizing and segmenting judicial documents and paragraph annotation based on a rule base is disclosed.

[0068] AsFigure 1 As shown in the figure, the rule-based vectorized segmentation and paragraph annotation method of judicial documents includes:

[0069] S1: Build a rule base containing segmentation rules and paragraph annotation rules;

[0070] S2: Obtain the judicial documents to be annotated and pre-process them;

[0071] S3: Matching corresponding segmentation rules and paragraph marking rules for each document part in the preprocessed judicial document through regular expressions;

[0072] In this embodiment, regular expressions are used to identify and process specific text paragraphs, such as extracting a case summary. For example, the regular expression reference example is: litigation request (.*) trial and investigation. (.*) will match as few arbitrary character sequences as possible, starting from "litigation request" until the first "trial and investigation" is encountered.

[0073] The segmentation rules include:

[0074] 1. Rules based on specific characters or symbols

[0075] 1.1 Newline and blank line:

[0076] Rule: Divide paragraphs between consecutive blank lines (usually two or more line breaks). In legal documents, paragraphs are often separated by blank lines.

[0077] 1.2 Paragraph mark symbols:

[0078] Rule: Use specific markings (such as paragraph numbers, separators, etc.) to divide paragraphs. In some formatted judicial documents, each paragraph may be preceded by a number or a specific separator.

[0079] 2. Rules based on text structure and format

[0080] 2.1 Title and text:

[0081] Rule: Divide the paragraphs based on the formatting differences between the title and the body (such as font, size, bold, etc.). In judicial documents, titles are often highlighted by using different fonts, sizes, or bold. For example, "I. Basic Information of the Case" is used as a title, and the body content that follows it can be regarded as a paragraph.

[0082] 2.2 Lists and paragraphs:

[0083] Rule: Divide paragraphs by formatting differences between list items and paragraphs (e.g., indents, bullets, etc.). List formatting may be used when listing facts or evidence. Each list item can be considered a subparagraph, while the content outside the list belongs in another paragraph.

[0084] 3. Rules Based on Semantics and Content

[0085] 3.1 Logical Structure and Topic Sentences:

[0086] Rule: Determine paragraph boundaries based on the logical structure and topic sentences of the text. For example, when a new topic sentence appears, it usually means the start of a new paragraph. In judicial documents, each paragraph usually has a clear topic sentence to summarize the main content of that paragraph. For example, a sentence like "The court holds that, according to..." usually marks the start of a new paragraph.

[0087] 3.2 Keywords and Phrases:

[0088] Rule: Determine paragraph boundaries by identifying specific keywords or phrases. These keywords or phrases may indicate transitions or separations between paragraphs. In judicial documents, words such as "In summary", "Accordingly", etc. may indicate the end of one paragraph and the start of another.

[0089] The definition of the paragraph annotation rule is similar to the paragraph segmentation rule.

[0090] S4: Segment the judicial document according to the matching result of the segmentation rule to obtain several document paragraphs;

[0091] S5: Perform paragraph annotation on each document paragraph according to the matching result of the paragraph annotation rule to obtain the annotation information for each document paragraph;

[0092] In this embodiment, the annotation information of the document paragraph includes one or more of the case cause, case summary, trial area, case category, number of trials, litigation request, factual reasons, defendant's defense, found facts, the court's opinion, legal basis, and judgment result. After obtaining the annotation information for each document paragraph, attribute information can be further added to each document paragraph; the attribute information includes one or more of the author, release date, and citation times.

[0093] S6: After vectorizing all the document paragraphs with annotation information, output them as the annotation result of the judicial document.

[0094] By constructing a rule library containing detailed segmentation rules and paragraph annotation rules, the present invention can accurately match each part in the judicial document, such as the case cause, case summary, etc., avoiding errors that may occur during manual annotation; at the same time, the use of regular expressions further enhances the accuracy of rule matching, ensuring that each paragraph can be correctly classified and annotated, thereby improving the accuracy of judicial document annotation. And compared with the traditional manual annotation method, the automatic annotation method of the present invention can quickly process a large number of judicial documents, reduce the manual burden, and improve work efficiency.

[0095] In the annotation process of the present invention, the judicial document is segmented according to the segmentation rules first, making the structure of the entire judicial document clearer, facilitating users to quickly grasp the main content of the document, and each paragraph has a clear theme, such as the case cause, case summary, etc., which helps readers form a well-organized cognitive framework when reading. At the same time, after the judicial document is segmented, each paragraph is provided with corresponding annotation information through the paragraph annotation rules, such as the case category, number of trials, etc. These information are very important reference bases for users, which can help users find the required information faster, thereby improving the efficiency of reading judicial documents. In addition, the present invention performs segmentation and annotation through a unified rule library and regular expressions, which can ensure the unity and standardization of judicial document processing, and can improve the problem of inconsistent annotation caused by individual understanding differences, thereby improving the standardization of judicial document annotation.

[0096] To better introduce the technical solution of the present invention, this embodiment is described through the following several parts.

[0097] I. Rule Library

[0098] In this embodiment, the design principles of the rule library are as follows:

[0099] Domain analysis: Analyze the characteristics of the target text domain, including language style, structural layout, common terms, etc.

[0100] Rule classification: According to the text characteristics, the rules are classified into categories such as structural rules, semantic rules, and format rules.

[0101] Rule priority: Determine the applicable priority between different rules to solve the problem of rule conflicts.

[0102] Specifically, the rule categories of the segmentation rules and paragraph annotation rules in the rule library include one or more of structural rules, semantic rules, and format rules;

[0103] (1) Structural rules: Used to determine the basic segmentation structure of the judicial document and the logical relationship between each part;

[0104] The construction steps of the structural rules are as follows:

[0105] 1) Analyze the common structure of the judicial document and determine several document structures;

[0106] In this embodiment, the document structure includes a title, an introduction, a body (which may contain multiple sub-paragraphs), a list, a conclusion, and an appendix, etc.

[0107] 2) Define a unique identifier for each document structure;

[0108] In this embodiment, the title is represented by "TITLE", the first part of the text is represented by "BODY_PART1", and so on.

[0109] 3) Determine the sequential relationship between each document structure; for example, the title comes before the introduction, and the text comes before the conclusion, etc.

[0110] 4) Record the structural rules in a format that is easy to program and query (such as XML, JSON, or a database table).

[0111] (2) Semantic rules: Used to understand the information content in each paragraph or sentence and associate it with specific legal concepts or entities;

[0112] The construction steps of the semantic rules are as follows:

[0113] 1) Identify common legal concepts or entities in judicial documents; including charges, legal provisions, names of parties, dates, etc.

[0114] 2) Define tags for each legal concept or entity; for example, "charge" can be represented by "CHARGE", and "legal provision" can be represented by "LEGAL_PROVISION".

[0115] 3) Define rules to match legal concepts or entities in the text; this may involve regular expressions, named entity recognition (NER) models, or custom parsers.

[0116] 4) Integrate the semantic rules with the structural rules so that semantic annotation can be performed while segmenting the text;

[0117] (3) Formatting rules: Used to ensure that the format of judicial documents complies with specific legal or institutional requirements;

[0118] 1) Determine the formatting requirements for judicial documents; including font, font size, paragraph spacing, title format, etc.

[0119] 2) Define rules to check whether the format of the document complies with the requirements; involving regular expressions, document parsing libraries, or custom format validators.

[0120] 3) If the format does not meet the requirements, define corrective measures; this may include automatically adjusting the format or prompting the user to modify it manually.

[0121] 4) Integrate the formatting rules with the structural rules and semantic rules so that format issues can be verified and corrected simultaneously during the annotation process.

[0122] (4) Integrate the rule library through the following steps:

[0123] 1) Integrate structural rules, semantic rules, and formatting rules into a unified rule base, which can be a database, configuration file, or code library.

[0124] 2) Provide an easy-to-use API or interface for the rule base to facilitate access and query of rules during the annotation process.

[0125] 3) Test the rule base to ensure it can correctly process and annotate various legal documents.

[0126] (5) Optimize the rule base through the following steps:

[0127] 1) Continuously update the rule base as new types of legal documents and legal concepts emerge.

[0128] 2) Collect user feedback and fine-tune the rule base to improve the accuracy and efficiency of annotation.

[0129] 3) Regularly evaluate the performance of the rule base to ensure it can adapt to the changing legal environment and user needs.

[0130] Specifically, the rule categories of the segmentation rules and paragraph annotation rules in the rule base include synonym rules: used to map one or more synonyms to standard vocabulary;

[0131] Construct synonym rules through the following steps:

[0132] S101: Obtain a pre-trained BERT (Bidirectional Encoder Representations from Transformers) model on a large-scale corpus;

[0133] S102: Convert the vocabulary in the legal document samples into word vectors through the BERT model;

[0134] S103: Calculate the cosine similarity between word vectors, and take the vocabulary corresponding to the word vectors with cosine similarity exceeding the threshold as synonyms;

[0135] S104: Store all synonyms in the synonym library;

[0136] S105: Generate corresponding synonym rules for each standard vocabulary according to the synonym library.

[0137] In this embodiment, a DBMS is used to store and manage the rule base to ensure fast retrieval and update of rules. Software engineering practices such as modular design and version control are adopted to maintain the stability and scalability of the rule base.

[0138] II. Decision Tree

[0139] In this embodiment, a decision tree for identifying document categories and determining the priority of judgment rules is constructed;

[0140] The decision tree is constructed through the following steps:

[0141] S111: (By analyzing the characteristics of rules and document categories) For each type of document, a corresponding rule library is constructed in advance, including segmentation rules and paragraph annotation rules;

[0142] S112: Obtain the labeled judicial document samples and perform preprocessing;

[0143] In this embodiment, the labeled judicial document samples need to cover all possible document categories. The preprocessing includes word segmentation, stop word removal, stemming, etc.

[0144] S113: Extract features (including word frequency, TF-IDF value, word vector) from the preprocessed judicial document samples, select the key features with high discrimination for document categories as the input of the decision tree; train the decision tree through machine learning algorithms (such as CART, ID3, C4.5, etc.);

[0145] S114: Identify the document category of the judicial document through the trained decision tree;

[0146] S115: Select the corresponding rule library for the judicial document according to the identified document category to perform the rule matching work in step S3;

[0147] S116: In steps S4 and S5, when there are multiple rules applicable to the same situation, the decision tree determines the execution order of each rule through a preset priority logic.

[0148] Through the constructed decision tree, the present invention can perform accurate document classification based on the features of the document, thereby ensuring that the document is correctly classified and improving the accuracy of subsequent rule selection. Through the branch structure of the decision tree, the rules applicable to specific document categories can be automatically selected, reducing the time and cost of manual selection and judgment. At the same time, the priority determination mechanism in the decision tree can effectively solve the conflicts between rules and ensure the selection of the optimal rule when multiple rules are applicable.

[0149] III. Preprocessing

[0150] In this embodiment, preprocessing is a key step in text analysis, and its purpose is to convert the original text into a format suitable for subsequent processing. The preprocessing of judicial documents includes one or more of data cleaning, text normalization, text segmentation, stop word removal, and data desensitization processing;

[0151] Data cleaning: including noise removal and normalization; noise removal means removing irrelevant information from legal documents, such as advertisements, copyright notices, meaningless characters, etc.; normalization means unifying the text encoding format, such as converting all text to UTF-8 encoding to ensure consistency in subsequent processing.

[0152] Text normalization: including case conversion and punctuation handling; case conversion means converting all English characters to lowercase or uppercase as needed to eliminate differences caused by case; punctuation handling means standardizing the use of punctuation marks and removing or replacing non-standard punctuation marks.

[0153] Text segmentation: including sentence boundary recognition and word segmentation; sentence boundary recognition means using sentence segmentation algorithms to identify sentence boundaries to prepare for subsequent paragraph annotation; word segmentation means performing word segmentation on the text, especially for languages without obvious word boundaries such as Chinese.

[0154] Stop word removal: Establish a stop word list according to language and domain characteristics, and remove stop words in legal documents based on the stop word list. Stop words are words that frequently appear in the text but contribute little to understanding the meaning of the text, such as "of", "in", "and", etc. Removing these words can help reduce data noise and improve the efficiency of subsequent processing.

[0155] Text standardization: including data desensitization and stemming; data desensitization means replacing the last three digits of consecutive digital strings of personal identity information (name, ID number, mobile phone number, address, etc.) in the text with asterisks for data desensitization; stemming means using stemming algorithms (such as Porter Stemmer) for languages like English to extract the root words and reduce the impact of word form changes.

[0156] Specifically, the preprocessing includes synonym replacement using synonym rules: if the current word is a synonym in the synonym library, it is replaced with the corresponding standard word.

[0157] IV. Rule Matching

[0158] In this embodiment, regular expressions are used to identify the parts in the legal document that conform to the patterns in the rule library, and then corresponding segmentation rules and paragraph annotation rules are matched for each document part in the legal document;

[0159] The specific processing steps are as follows:

[0160] S301: Determine the pattern and label of each rule in the rule library;

[0161] S302: Write the corresponding regular expression for the pattern of each rule in the rule library;

[0162] Illustrative example:

[0163] Rule 1: Identify the titles in the text. The title starts with "Title:" followed by any characters.

[0164] The regular expression pattern corresponding to Rule 1: ^Title:.* (matches the lines starting with "Title:").

[0165] Rule 2: Identify the dates in the text. The date format is "YYYY-MM-DD".

[0166] The regular expression pattern corresponding to Rule 2: \d{4}-\d{2}-\d{2} (matches the dates in the "YYYY-MM-DD" format).

[0167] In this embodiment, if some rules cannot be accurately expressed by regular expressions, a custom string is defined for the rule, and then the matching is completed through a string matching algorithm.

[0168] S303: Traverse the regular expressions of each rule in the rule library, and use a regular expression matching function (such as re.search() or re.match() in Python) to match the judicial documents. If the match is successful, record the matched document part (including the position and length) and the corresponding rule label.

[0169] S304: Establish a mapping relationship between each document part in the judicial document and the corresponding rule according to the matching results, and associate each text part with the corresponding rule.

[0170] V. Optimization of Regular Expressions

[0171] In this embodiment, the regular expressions are trained and optimized through machine learning algorithms.

[0172] The specific processing steps are as follows:

[0173] S311: Use regular expressions to extract text parts from the judicial document samples, and verify the accuracy of the extraction results (through manual detection); divide the judicial document samples with extraction results into a training set and a test set.

[0174] S312: Select a machine learning model as the optimization model; such as logistic regression, support vector machine, random forest or deep learning model.

[0175] S313: Convert the judicial document samples and regular expressions in the training set into feature vectors as the model input, and use the extraction results as the model output to train and optimize the model; test the performance of the optimized model through the test set.

[0176] S314: Use the trained optimization model to predict judicial documents and output the corresponding predicted extraction results;

[0177] S315: Obtain the actual extraction results of the regular expression for judicial documents; compare the predicted extraction results with the actual extraction results, and optimize the regular expression based on the differences in the comparison results.

[0178] In this embodiment, the predicted extraction results and the actual extraction results can be compared and the regular expression can be optimized by means of manual inspection and expert evaluation. Regularly evaluate the model with new data, and through continuous monitoring, evaluation, training and optimization, form a positive feedback loop to continuously improve the prediction ability of the model and the accuracy of the regular expression.

[0179] The present invention can automatically adjust and optimize the parameters of the regular expression through machine learning algorithms, making it more accurately match text features, thereby improving the accuracy of annotation. And the machine learning algorithm can learn the features of different document types, making the regular expression more generalizable and adaptable to more variable text environments. Through continuous training and optimization, the regular expression can identify key information faster and improve the annotation speed.

[0180] VI. Vectorize and output the annotation results

[0181] In this embodiment, the BGE-M3 vector model is used to vectorize the document paragraphs. The BGE-M3 vector model is a multi-functional, multi-language, multi-granularity text embedding model that can simultaneously perform three common retrieval functions of the embedding model, namely dense retrieval, multi-vector retrieval and sparse retrieval. The structure of the BGE-M3 vector model is based on advanced deep learning technologies and has shown excellent performance in multiple natural language processing benchmark tests, such as leading results in benchmark tests such as MTEB, C-MTEB, Miracl and Mk QA, providing a high-quality basis for the vectorization of text paragraphs.

[0182] Specifically, the processing steps for vectorizing and outputting the annotation results are as follows:

[0183] S601: Obtain the vector model and fine-tune the vector model with text data (corpus) in the judicial field;

[0184] In this embodiment, in order to further optimize the application effect of the vector model in the judicial field, it is fine-tuned using a large-scale corpus in the judicial field. The specific fine-tuning steps are as follows:

[0185] S6011: Collect text data in the judicial field, including various judicial documents, laws and regulations, judicial interpretations, case analyses, etc., and construct a corpus in the judicial field. These data cover different types of judicial cases, court documents at different levels, and application scenarios of various legal provisions, ensuring the comprehensiveness and representativeness of the corpus.

[0186] S6012: Preprocess the collected text data in the judicial field, including operations such as text cleaning, word segmentation, and stop word removal. For specific terms and legal concepts in judicial documents, we conduct special identification and marking so that the model can better learn these key information. For example, for professional terms in legal provisions, such as "liability for breach of contract", "tortious act", "legal representative", etc., we conduct unified annotation and normalization processing to make their representations in the text consistent, facilitating the model to learn and understand their semantics.

[0187] S6013: Design fine-tuning tasks. For example, the judicial document classification task can be used as the target task for fine-tuning. We annotate judicial documents according to different categories, such as civil judgments, criminal rulings, administrative complaints, etc., enabling the model to learn the characteristics and differences of different types of judicial documents, thus better adapting to the text characteristics and semantic requirements of the judicial field.

[0188] S6014: Conduct fine-tuning, use appropriate optimizers (such as the Adam optimizer) and loss functions (such as the cross-entropy loss function), and gradually adjust the model's parameters according to the model's performance on the fine-tuning task, so that the model can achieve better performance in the text vectorization task in the judicial field. Through multiple iterative trainings, continuously optimize the model's weights, enabling it to more accurately capture the semantic information and legal logical relationships in judicial documents, and improving the quality and accuracy of the vectorized representation.

[0189] S602: After preprocessing and word segmentation of the document paragraphs with annotation information, perform text conversion to convert the document paragraphs into an input format that the vector model can understand;

[0190] In this embodiment, for each paragraph of the document, its annotation information is collected, including paragraph type (such as factual description, legal basis, etc.), related legal concepts, entity names, etc. The preprocessing includes reading and parsing the annotated text paragraphs, removing unnecessary punctuation marks, stop words, etc., to ensure that the input text is clean and meets the model input requirements. Tokenization means using the tokenizer of BGE-M3 to divide the text paragraphs into tokens. Text conversion means converting the annotation information into a form that the model can understand. For example, if the model accepts words as input, the annotation information can be converted into keywords or phrases. For legal concepts and entity names, a professional legal term dictionary or named entity recognition technology is used for accurate extraction and standardization processing to ensure that the model can accurately identify and understand the meaning and role of these key information in judicial documents.

[0191] S603: Input the converted document paragraphs into the fine-tuned vector model for vectorization processing, and output the corresponding document paragraph vectors;

[0192] In this embodiment, the vector model is used to perform vectorization processing on each annotated text paragraph. The vector model is based on the Transformer architecture and uses the self-attention mechanism to encode the relationship of each word in the context. The self-attention mechanism enables the model to automatically focus on the degree of association between different words in the text, thereby better capturing the semantic information of the paragraph. For example, when processing a paragraph containing multiple legal clause references and factual descriptions, the self-attention mechanism focuses on the words and legal clauses related to the core facts of the case to accurately understand the semantic focus and logical structure of the paragraph.

[0193] The processing steps of the vector model include:

[0194] S60301: Encode the words into vectors in a high-dimensional space through the forward propagation of the vector model. In this process, the model maps each word to a high-dimensional vector space according to the semantic representation method it has learned, so that words with similar semantics are closer in the vector space, while words with different semantics are farther apart, thus realizing the vectorized representation of the text.

[0195] S60302: Apply a pooling strategy (such as average pooling, max pooling, attention pooling) to combine the word-level vectors to form a paragraph-level vector representation. Average pooling simply averages the word vectors and can retain the overall semantic information of the paragraph. In the vectorization of judicial documents, the average pooling strategy can be adopted. For example, for the legal clause retrieval task that needs to highlight the key information of the paragraph, the max pooling or attention pooling strategy may be selected; while for the document classification task that needs to comprehensively understand the paragraph semantics, the average pooling strategy is more appropriate.

[0196] S60303: After the vectorization process and pooling operation of the vector model, the vector representation of each document paragraph is obtained. These vectors have a fixed dimension, and the dimension size depends on the parameters set during the model initialization. For example, it can be 768 dimensions, 1024 dimensions, etc. The choice of the specific dimension usually requires a trade-off between model performance and computing resources.

[0197] S604: Post-process the document paragraph vectors to make them meet the preset standards and requirements;

[0198] In this embodiment, the post-processing includes standardization to ensure that the vectors are within the same numerical range. Standardization means subtracting the mean of each dimension of the vector and dividing it by its standard deviation, making the numerical distribution of the vector more stable and facilitating subsequent calculations and analyses. It also includes dimensionality reduction, which reduces the dimension of the vector through techniques such as PCA (Principal Component Analysis) and t-SNE (t-Distributed Stochastic Neighbor Embedding), while retaining the original information as much as possible. PCA projects the high-dimensional data into a low-dimensional space through a linear transformation, so that the projected vector can retain the variance information of the original data to the greatest extent, thereby reducing information loss while reducing the dimension; t-SNE is more adept at maintaining the local structure and distribution characteristics of the data in the low-dimensional space, and can map the high-dimensional vector into a two-dimensional or three-dimensional space for visual analysis and further exploratory data analysis, helping users more intuitively understand the relationships and distribution laws between judicial document vectors.

[0199] S605: Store and output the post-processed document paragraph vectors and their corresponding annotation information.

[0200] In this embodiment, the vectorization results of each paragraph are stored in an appropriate data structure, such as an array, a list, or a DataFrame. According to the requirements of the actual application, select a suitable data storage method to facilitate the management and subsequent processing of the vectorized data. At the same time, the vectorized data structure can also be output to files and databases. When outputting the data to a file, a common text format (JSON) can be selected to facilitate data storage and transmission; if the data is stored in a database, appropriate table structures and storage methods need to be designed to ensure the integrity and consistency of the data.

[0201] S606: Verification and evaluation: Verify and evaluate the output document paragraph vectors to ensure that they accurately reflect the content of the document paragraphs and meet our quality standards.

[0202] In this embodiment, a set of test data is used to verify whether the vectorized query database of the model meets the expectations. The test data should have similar distributions and characteristics to the training data and the judicial documents in the actual application scenarios, covering different types of judicial documents, different legal fields, and text contents of various complexities. By inputting the test data into the model for vectorization processing, comparing the obtained similar document query results with the expected text, it is checked whether the model can accurately capture the semantic information and annotation features in the judicial documents.

[0203] Meanwhile, evaluate the performance of the vectorized data in downstream tasks, such as clustering, classification, or relevance retrieval. In the clustering task, observe whether the vectorized judicial document paragraphs can be naturally clustered into different categories according to their semantic and annotation information. For example, paragraphs of the same type of judicial documents (such as evidence paragraphs in civil cases, paragraphs of criminal facts in criminal cases, etc.) are grouped together, and there is an obvious distinction between different categories of paragraphs; in the classification task, use classifiers such as support vector machines (SVM), decision trees, and deep learning classification models to classify the vectorized judicial documents, and evaluate classification accuracy, recall rate, F1 value and other metrics to measure the model's ability to identify the types of judicial documents; in the relevance retrieval task, by constructing a retrieval system, inputting query statements and retrieving relevant judicial document paragraphs, evaluate the accuracy of the retrieval results and the rationality of the relevance ranking. For example, use metrics such as mean average precision (MAP) and normalized discounted cumulative gain (NDCG) to measure the performance of the retrieval system, and judge whether the vectorized data can effectively support the rapid retrieval and accurate matching of judicial documents to meet the information retrieval needs in actual judicial work.

[0204] The present invention realizes the vectorization of the annotated text paragraphs through the BGE-M3 vector model. First, the vectorized text paragraphs are convenient for rapid retrieval and matching, thus improving the efficiency of judicial document processing. Secondly, through vectorization, the model can better capture the semantic information of the text paragraphs, which helps the understanding and application of legal knowledge. Finally, the vectorization of text paragraphs can provide a basis for the automatic classification, abstract generation, intelligent recommendation, etc. of judicial documents, which is conducive to improving the standardization of judicial document annotation and the intelligence of judicial work. The judicial documents vectorized by the present invention are more convenient for large language models to accurately understand the structure, rules, meanings, and judicial judgment logics of these judicial documents, so as to improve the accuracy and quality of the judicial documents assisted by large language model artificial intelligence technology.

[0205] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Those of ordinary skill in the art should understand that any modifications or equivalent replacements to the technical solutions of the present invention without departing from the purpose and scope of the present technical solution shall be covered by the scope of the claims of the present invention.

Claims

1. A method for vectorized segmentation and paragraph annotation of judicial documents based on a rule base, characterized in that: include: S1: Build a rule base containing segmentation rules and paragraph annotation rules; S2: Obtain the judicial documents to be annotated and pre-process them; S3: Matching corresponding segmentation rules and paragraph marking rules for each document part in the preprocessed judicial document through regular expressions; S4: segmenting the judicial document according to the matching results of the segmentation rules to obtain a number of document paragraphs; S5: marking each document paragraph according to the matching result of the paragraph marking rule to obtain the marking information of each document paragraph; S6: After all document paragraphs with annotation information are vectorized, they are output as annotation results of judicial documents.

2. The method for vectorized segmentation and paragraph annotation of judicial documents based on a rule base as claimed in claim 1, characterized in that: In step S1, the rule categories of segmentation rules and paragraph marking rules in the rule base include: one or more of structural rules, semantic rules and format rules; (1) Structural rules: used to determine the basic segmentation structure of judicial documents and the logical relationship between the various parts; (2) Semantic rules: used to understand the information content of each paragraph or sentence and associate it with specific legal concepts or entities; (3) Format rules: used to ensure that the format of judicial documents complies with specific legal or institutional requirements.

3. The rule-based vectorized segmentation and paragraph marking method for judicial documents according to claim 2, characterized in that: In step S1, the rule categories of segmentation rules and paragraph annotation rules in the rule base include synonym rules: used to map one or more synonyms to standard vocabulary; The synonym rules are constructed by following the steps below: S101: Get the BERT model pre-trained on a large-scale corpus; S102: Convert the vocabulary in the judicial document sample into word vectors through the BERT model; S103: Calculate the cosine similarity between word vectors, and take the words corresponding to the word vectors whose cosine similarity exceeds a threshold as synonyms; S104: storing all synonyms in a synonym database; S105: Generate corresponding synonym rules for each standard word according to the synonym database.

4. The rule-based vectorized segmentation and paragraph marking method for judicial documents according to claim 1, characterized in that: In step S1, a decision tree is constructed for identifying document categories and determining rule priorities; The decision tree is constructed by the following steps: S111: Pre-build a corresponding rule base for each document type, including segmentation rules and paragraph marking rules; S112: Obtain annotated judicial document samples and perform pre-processing; S113: Extract features from the preprocessed judicial document samples as input to the decision tree; train the decision tree using a machine learning algorithm; S114: Identify the document category of judicial documents through the trained decision tree; S115: Selecting a corresponding rule base for the judicial document according to the identified document category to perform the rule matching work in step S3; S116: In steps S4 and S5, when there are multiple rules applicable to the same situation, the decision tree determines the order of execution of each rule through a preset priority logic.

5. The rule-based judicial document vectorization segmentation and paragraph marking method according to claim 1, characterized in that: In step S2, the preprocessing of the judicial documents includes: one or more of data cleaning, text normalization, text segmentation, removal of stop words and data desensitization.

6. The rule-based vectorized segmentation and paragraph marking method for judicial documents according to claim 2, characterized in that: In step S2, the preprocessing includes performing synonym replacement using synonym rules: if the current word is a synonym in the synonym library, it is replaced with the corresponding standard word.

7. The rule-based vectorized segmentation and paragraph marking method for judicial documents according to claim 1, characterized in that: In step S3, regular expressions are used to identify parts of the judicial document that conform to the patterns in the rule library, and then the corresponding segmentation rules and paragraph marking rules are matched for each document part in the judicial document; The specific processing steps are as follows: S301: Determine the mode and label of each rule in the rule base; S302: Write a corresponding regular expression for the pattern of each rule in the rule base; S303: traverse the regular expression of each rule in the rule base, use the regular expression matching function to match the judicial document, and if the match is successful, record the matched document part and the corresponding rule label; S304: Establish a mapping relationship between each document part in the judicial document and the corresponding rule according to the matching result, and associate each text part with the corresponding rule.

8. The rule-based judicial document vectorization segmentation and paragraph marking method according to claim 7, characterized in that: In step S3, the regular expression is trained and optimized by a machine learning algorithm; The specific processing steps are as follows: S311: extracting text parts from judicial document samples using regular expressions; dividing the judicial document samples with extraction results into a training set and a test set; S312: Select a machine learning model as an optimization model; S313: converting the judicial document samples and regular expressions in the training set into feature vectors as model inputs, and using the extraction results as model outputs to train the optimization model; and testing the performance of the optimization model through the test set; S314: Use the trained optimization model to predict the judicial document and output the corresponding prediction extraction result; S315: Obtain the actual extraction result of the regular expression for the judicial document; compare the predicted extraction result with the actual extraction result, and optimize the regular expression based on the difference in the comparison results.

9. The rule-based vectorized segmentation and paragraph marking method for judicial documents according to claim 1, characterized in that: In step S6, the steps for vectorizing the document paragraph are as follows: S601: Obtain a vector model, and fine-tune the vector model using text data in the judicial field; S602: After preprocessing and word segmenting the document paragraph with the annotated information, the document paragraph is converted into an input format that can be understood by the vector model; S603: input the converted document paragraph into the fine-tuned vector model for vectorization, and output the document paragraph vector; S604: Post-processing the document paragraph vector to make it meet the preset standards and requirements; S605: Store and output the post-processed document paragraph vector and its corresponding annotation information.

10. The rule-based vectorized segmentation and paragraph marking method for judicial documents according to claim 1, characterized in that: In step S5, the annotation information of the document paragraph includes one or more of the case cause, case summary, trial area, case category, number of trials, litigation request, factual reasons, defendant's defense, determined facts, this court's opinion, legal basis and judgment result.

Citation Information

Cited By

  • Deep learning-based judgment document key entity extraction method

    CN121031598A

  • Construction scheme intelligent auditing method and system

    CN121435921A