Judicial text entity recognition method based on ternary instruction fine tuning and VCoT verification
Through the judicial text entity recognition method of ternary instruction fine-tuning and VCoT verification, the problem of recognition of long expressions and nested entities in judicial text is solved, and efficient and accurate identification of key entities in judicial text is achieved, and the efficiency and accuracy of judicial text processing is improved.
Patent Information
- Application Number
- CN202510403229.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-01
AI Technical Summary
There are problems in judicial texts with long expressions, nested entities and fine-grained recognition. It is difficult for the existing technology to efficiently and accurately identify key entities in judicial texts, especially in complex contexts and multi-entity interleaving, and the recognition effect is poor.
The judicial text entity recognition method based on ternary instruction fine-tuning and VCoT verification is adopted. Through the enhanced instruction fine-tuning and VCoT verification mechanism of ternary understanding, the model's understanding of judicial text is improved and the accuracy and consistency of entity recognition is ensured. The specific steps include: inputting judicial documents and instructions, initial recognition through the judicial entity recognition model, optimizing using the ternary understanding enhancement module, combining the specification module, knowledge guidance module and comparison learning module, and finally gradual verification and correction through the VCoT verification mechanism.
It significantly improves the accuracy and consistency of entity recognition in judicial text, reduces misidentification, improves the recognition performance of the model in long expressions and nested entities, solves the problems of fuzzy entity boundaries and difficulty in distinguishing multiple entity types in judicial text, and ensures that the recognition results are in line with the semantic logic and actual situation of judicial text.
Smart Images

Figure CN120337927A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of smart judicial technology, and in particular to a judicial text entity recognition method based on ternary instruction fine-tuning and VCoT verification. Background Art
[0002] In recent years, the rapid development of artificial intelligence and big data technology has spawned a series of research directions with broad application value, such as smart justice and smart medical care. Among them, legal artificial intelligence has gradually become an important part of judicial practice. In actual judicial work, judicial personnel often need to quickly and accurately extract useful information from massive judicial texts, and the basic carrier of this information is named entities. How to efficiently and accurately process massive judicial texts and accurately identify key entities in judicial texts has become a key issue that needs to be urgently solved in the current judicial field.
[0003] In legal practice, named entity recognition plays a vital role in multiple application scenarios such as legal information retrieval, judicial knowledge graph construction, and legal judgment prediction. It can not only greatly improve the work efficiency of judicial personnel, but also effectively alleviate the work pressure of judges and promote the further development of smart justice.
[0004] In addition, with the continuous advancement of large language models, named entity recognition technology based on large language models has become a hot topic in current research. By training on a wide range of corpus knowledge bases, large models have strong natural language understanding capabilities and can accurately identify complex entities in rich contexts.
[0005] However, the task of named entity recognition in the field of judicial texts still faces many challenges. First, long expression entities composed of multiple nouns or phrases are common in judicial texts. These entities are often difficult to accurately segment using traditional word segmentation methods, and at the same time, higher requirements are placed on sequence modeling. Secondly, in order to more accurately describe the case, nested named entities are often used in judicial texts to express complex legal facts, which results in the entity boundaries often overlapping, making entity extraction extremely difficult. In addition, compared with named entities in general fields (such as names of people, places, and institutions), the judicial field pays more attention to the fine-grained distinction of entities, requiring that characters be further distinguished into specific judicial roles such as suspects and victims, and the time, place, and related items of the case be accurately identified. This highly refined entity recognition requires not only a deeper division of entity categories, but also the problem of data sparsity. These challenges together increase the complexity of named entity recognition in judicial texts and place higher requirements on existing technologies. Summary of the invention
[0006] The object of the present invention is: aiming at the deficiencies in the above-mentioned background technology, to provide a method for instruction fine-tuning based on an improved large language model and a scheme for introducing a VCoT verification mechanism, aiming to improve the entity recognition performance in judicial texts and solve the problems of long expressions, nested entities, and fine-grained recognition.
[0007] To achieve the above object, the present invention provides a method for judicial text entity recognition based on triple instruction fine-tuning and VCoT verification, including the following steps:
[0008] S1. The user inputs the judicial document to be recognized and instructions.
[0009] S2. An initial response is obtained through a large model for judicial entity recognition.
[0010] The large model for judicial entity recognition is trained and optimized through triple understanding-enhanced instruction fine-tuning to improve the understanding ability of judicial texts and accurately recognize different types of entities.
[0011] Instruction fine-tuning guides the large model for judicial entity recognition to recognize various key entity information from judicial documents through explicit instruction design and injection of domain knowledge.
[0012] Triple understanding enhancement improves the entity recognition performance in judicial texts through in-depth semantic understanding and context adaptation ability.
[0013] S3. The initial response is gradually inferred and verified through the VCoT verification mechanism, and the verification result is corrected and optimized to generate a final entity recognition result after verification.
[0014] The VCoT verification mechanism gradually checks whether each recognized entity is consistent with the context in the original text and ensures that there are no omissions or inconsistencies in the annotations generated by the model.
[0015] The VCoT verification mechanism gradually generates a series of inference chains through multiple rounds of inference verification, so that the entities generated by the model are not only correctly recognized but also conform to the actual context and logical relationship of the context.
[0016] Further, in S1, the input module receives the judicial document to be recognized and instructions provided by the user.
[0017] Further, the instruction fine-tuning includes:
[0018] By using context-aware data re-representation, the context around the entity is regarded as a key non-entity sample, so that the model can learn how to more accurately distinguish entity and non-entity information through samples.
[0019] The dataset is re-annotated using the label prefix annotation method, adding a unique prefix label for each entity type so that the model can accurately distinguish various fine-grained entities when generating output.
[0020] Furthermore, the instruction fine-tuning also includes:
[0021] A task mode of generating and predicting simultaneously, so that the model will predict in real time whether the currently generated word belongs to a specific entity category when generating text, and thus decide whether to add an entity label. For ordinary words that do not belong to any entity, the model will keep their original form without adding any labels.
[0022] Furthermore, the triple understanding enhancement includes: a specification module, a knowledge guidance module, and a contrastive learning module;
[0023] The specification module is used to standardize the generated content to ensure the accuracy, integrity, and consistency of the output;
[0024] The knowledge guidance module is used to provide a heuristic list containing the definitions of each entity type and related feature words to help the model understand and distinguish different entity types; the heuristic list contains high-level rules or strategies for inferring specific tasks, and the feature word list contains the feature words associated with each entity type;
[0025] The contrastive learning module is used to select examples with highly similar semantics, TF-IDF, and dependency relationships from the transformed training set, and combine them with the corresponding pre-transformed instances to construct a pair of strongly contrastive learning samples. In the way of contrastive learning, it helps the model deeply understand the annotation rules combining context-aware data re-representation method and label prefix method, so that the model can not only understand the semantics and context information of entity types during the learning process, but also master how to output normalized results with label prefixes according to instructions.
[0026] Furthermore, when training the large model for judicial entity recognition in S2, the input F of the model input consists of the following three parts:
[0027] Judicial text: X = {x1, x2, …, x n};
[0028] Instruction: I;
[0029] Triple understanding enhancement: F enhanced , which includes the specification module text, the knowledge guidance module text, and the contrastive learning module text, that is
[0030] F enhanced = X formal + X knowledge + X contrast
[0031] Among them, X formal is the text of the specification module, X knowledge is the text of the knowledge guidance module, X contrast is the text of the contrastive learning module;
[0032] F input = [X; I; F enhanced
[0033] During model training, first perform task definition and loss function setting; then perform input vectorization; then perform encoder processing; then perform decoder generation.
[0034] Furthermore, in S3, the self-verification module is used to ensure the accuracy and consistency of entity recognition results, and the VCoT verification mechanism is nested in the self-verification module.
[0035] Furthermore, the content of the inference chain includes:
[0036] Entity type consistency check: whether it is a predefined entity type; context consistency check: ensure that the generated content has no omissions, additions, or deviations from the original content; logical check: check whether the context collocation of entities conforms to logic.
[0037] Furthermore, the VCoT verification mechanism designs prompt words so that the final output not only meets the requirements of the instructions but also can handle potential ambiguities and complex situations in judicial texts, providing secondary correction.
[0038] The above solution of the present invention has the following beneficial effects:
[0039] The judicial text entity recognition method based on triple instruction fine-tuning and VCoT verification provided by the present invention, through the context-aware data re-representation method of instruction fine-tuning, introduces non-entity samples as context information during the training process, significantly enhances the model's understanding ability of complex contexts, improves the model's discrimination ability for non-entity texts, and reduces the phenomenon of misrecognition; on this basis, through a unified label prefix annotation method, not only introduces exclusive label prefixes for each entity type, solves the problems of fuzzy entity boundaries and difficult distinction of multiple entity types in judicial texts, but also simultaneously annotates multiple entity types with a single output, avoiding the complexity of multi-round prompt interaction in traditional methods;
[0040] Based on instruction fine-tuning, the present invention effectively solves the problem of insufficient understanding of complex contexts in judicial texts through a ternary understanding enhancement module composed of a specification module, a knowledge-guided module, and a contrastive learning module. In particular, significant breakthroughs have been made in the accurate segmentation of long-expression entities, the boundary distinction of nested entities, and the recognition performance of fine-grained entities. Through multi-dimensional semantic reinforcement, the model can more accurately understand the instruction requirements and comprehensively improve the recognition and classification capabilities of complex entities in judicial documents.
[0041] Based on traditional chain-of-thought reasoning, the present invention has developed a new verification chain checking mechanism named VCoT. Through generating a question list, context consistency checking, and logical relationship verification, it realizes step-by-step self-correction and answer verification, significantly improving the reliability and consistency of the entity recognition results of the model in the context of long texts, effectively avoiding the hallucination problem in generative models, and ensuring that the recognition results conform to the semantic logic and actual situation of judicial texts.
[0042] Other beneficial effects of the present invention will be described in detail in the subsequent specific implementation section. Brief Description of the Drawings
[0043] Figure 1 is the step flow chart of the present invention;
[0044] Figure 2 is the system schematic diagram of the present invention;
[0045] Figure 3 is the schematic diagram of the specification module of the present invention. Detailed Description of the Invention
[0046] The following specific examples illustrate the implementation manners of the present disclosure. Those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts fall within the scope of protection of the present disclosure.
[0047] It should be noted that the following description relates to various aspects of embodiments within the scope of the appended claims. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of the aspects set forth herein can be used to implement a device and / or practice a method. Additionally, this device and / or method can be implemented using other structures and / or functionality in addition to one or more of the aspects set forth herein.
[0048] It should also be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present disclosure schematically. The diagrams only show the components related to the present disclosure and are not drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex. Additionally, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the aspects described can be practiced without these specific details.
[0049] As Figure 1 shown, an embodiment of the present invention provides a judicial text entity recognition method based on ternary instruction fine-tuning and VCoT verification, including the following steps:
[0050] S1, the user inputs the judicial document and instruction to be recognized.
[0051] In this embodiment, the input module is used to receive the judicial document and instruction provided by the user, and can pass them to the subsequent entity information extraction module. Refer to Figure 2 . Among them, the input module has a user input sub-module and a user instruction sub-module built in to complete the above functions respectively. Since the entity information extraction module can infer the correct task requirements from the concise instructions after training and fine-tuning, the user can use concise or colloquial instructions in the input module to express the request for the entity recognition task and guide the entity recognition.
[0052] S2, obtain an initial response through the judicial entity recognition large model.
[0053] In this embodiment, key information entities are accurately identified and extracted from judicial documents through the entity information extraction module of the built-in judicial entity recognition large model. It should be noted that the judicial entity recognition large model in this embodiment is based on the advanced pre-trained language model Flan-T5, and is optimized and trained through fine-tuning of instructions enhanced by ternary understanding, aiming to cope with the recognition challenges of long expression entities, nested entities and fine-grained entities in judicial documents. Among them, Flan-T5 is a pre-trained language model fine-tuned for different tasks, and its powerful natural language understanding capabilities enable it to handle complex text structures. In judicial texts, the recognition of long expression entities and nested entities is particularly difficult because these entities often span multiple words and are intertwined and overlapped with other entities.
[0054] To meet this challenge, the Judicial Entity Recognition Large Model improves its understanding of judicial texts through instruction fine-tuning enhanced by ternary understanding, ensuring that different types of entities can be accurately identified. Specifically, instruction fine-tuning is used to optimize the model training. During the training process, through clear instruction design and the injection of rich domain knowledge, the model is guided to identify various key entity information such as suspects, victims, and case occurrence time from judicial documents to better adapt to the named entity recognition task in the judicial field. At the same time, in order to cope with the complex semantic structure of judicial texts, instruction fine-tuning further improves the model's understanding of judicial entities and instruction requirements by combining ternary understanding enhancement, thereby more accurately extracting complex entity information and significantly improving the model's task adaptability and recognition capabilities.
[0055] Instruction fine-tuning requires a data set for model optimization training. Considering that there is currently no data set related to named entity recognition in the judicial field, the training data set used in this embodiment comes from the information extraction track of the 2021 China Legal Intelligence Technology Evaluation (CAIL 2021), focusing on analyzing the "theft" case involved in Article 264 of the Criminal Law, as an example. The data set contains 5247 data, covering ten different entity types, namely NHCS (suspect), NASI (stolen items), NS (location), NHVI (victim), NT (time), NCGV (item value), NCSM (stolen currency), NO (institution), NATS (tools for committing crimes), NCSP (theft profit), with a total of 343,640 characters and 25,466 entities. In order to improve the performance of the model in information extraction tasks, the CAIL 2021 data set is deeply transformed, and its annotation method is converted into a context-aware data re-representation method to enhance the model's ability to recognize and discriminate entities.
[0056] The context-aware data re-representation method specifically includes: regarding the context around the entity (i.e., the part excluding the entity) as key non-entity samples, through which the model can learn how to more accurately distinguish entity and non-entity information. For example, the part of the text that does not contain any entity is labeled as a context sample, which helps the model understand the surrounding context and thus more accurately determine the entity category of each word. Through this dataset construction method, the model can more correctly understand the entity type of each word according to the surrounding context, helping the model learn how to effectively distinguish entity and non-entity information, making it more accurate when processing complex texts. It can not only correctly identify the entities in the text but also more accurately distinguish ordinary texts without entities, reducing the situations of mislabeling and over-labeling, thereby greatly improving the performance and generalization ability when processing complex judicial texts.
[0057] On this basis, the dataset is re-labeled using the label prefix annotation method. Specifically, a unique prefix label is added to each entity type, so that the model can more accurately distinguish various fine-grained entities when generating the output. The improved annotation examples are as follows:
[0058] {"id":"xxx",
[0059] "context":"Around 13:00 on a certain day around August 10, 2017, the defendant Huang Moumou stole a 'lovme' brand mobile phone worth 285 yuan from the storage box of an electric bicycle parked by Guo Moumou on the first floor of the dormitory of Zhejiang ** Necktie and Garment Co., Ltd.",
[0060] "entities":"
NT:Around 13:00 on a certain day around August 10, 2017
NHCS:Huang Moumou
NS:the first floor of the dormitory of Zhejiang ** Necktie and Garment Co., Ltd.
NASI:a 'lovme' brand mobile phone
NHVI:Guo Moumou
NCGV:285 yuan in RMB
[0061] Therefore, through the unified label prefix method for annotation, each entity type has its exclusive prefix, solving the problems of difficult entity boundary definition and fine-grained entities in judicial texts, and multiple entity types can be marked simultaneously in one output, reducing the number of calls required by the model during generation and avoiding the complexity of multi-round instruction prompts. Thus, the context-aware data re-representation method adopted in this embodiment enables the model to retain the original sentence's context structure while annotating entities, so that the generated output not only contains entity information but also can reflect the integrity of the sentence, especially performing more excellently in the distinction of non-entity texts. The model can make more accurate judgments on entity types according to the context around the entity.
[0062] It should be noted that, in order to ensure the generalization ability of the model and its performance in real scenarios, in this embodiment, the transformed dataset is scientifically divided to enable more effective training and evaluation. Specifically, 5,247 pieces of data are divided into a training set, a validation set, and a test set according to the ratio of 8:1:1. Through the in-depth transformation of the CAIL 2021 dataset and the scientific dataset division strategy, the information extraction ability of the model for theft case texts has been significantly improved, providing more accurate technical support for the intelligent analysis of judicial texts.
[0063] Given that the generative large language model adopts a token-by-token generation paradigm, in this embodiment, a task mode of generating while predicting is further set up to maximize the performance of the model in the fine-grained entity recognition task. That is, while generating text, the model will real-time predict whether the currently generated word belongs to a specific entity category, and then decide whether to add an entity label to it. For ordinary words (i.e., context samples) that do not belong to any entity, the model will keep their original form without adding any labels. This design ensures the naturalness and readability of the text, while reducing the cases of mislabeling and over-labeling. For example: "Please analyze the provided sentence and identify the type of each entity token by token. In the output, please add a label prefix for each entity type. If a word does not belong to any entity category, do not add a label to it." During the token-by-token output process of the generative large language model, the model can adjust and predict in real-time based on the context information, which helps the model better understand the context and enables it to make more accurate judgments when facing polysemous words and complex sentence structures. By gradually understanding the text while generating an output with fine-grained entity labels, the accuracy and efficiency of entity information extraction have been significantly improved, providing a more efficient and accurate solution for the information processing of judicial texts.
[0064] It can be understood that the role of instruction fine-tuning is to make the model more precisely understand the user's needs and accurately execute tasks. However, simple instruction fine-tuning often cannot solve the polysemy and context understanding problems in complex judicial texts. To address this issue, triple understanding enhancement is incorporated into the instruction fine-tuning process to further improve the entity recognition performance in judicial documents through deeper semantic understanding and context adaptation capabilities.
[0065] Among them, triple understanding enhancement includes three modules: a specification module, a knowledge guidance module, and a contrast learning module. These three modules work together to ensure that the model not only understands the instruction requirements when performing tasks, but also can perform in-depth semantic reasoning based on complex judicial document texts to improve the recognition effect.
[0066] For the normalization module, since generative large language models have a certain degree of flexibility and uncertainty when outputting legal texts, inconsistencies or formatting irregularities may occur in entity annotation and entity types. Therefore, through the normalization module, the generated content is specifically standardized to ensure the accuracy, integrity, and consistency of the output, as Figure 3 shown. This module can ensure that the model can only recognize and output predefined fixed entity types and only uses the tag prefix method for annotation, avoiding the generation of information unrelated to the task.
[0067] Therefore, by setting strict entity recognition rules and annotation formats, the normalization module makes the behavior of the model meet expectations. For example, for entities such as criminal suspects and victims, the normalization module clearly requires the model to only recognize and output entities of these specific types, while avoiding the extraction of other irrelevant entities or information.
[0068] For the knowledge guidance module, it consists of a heuristic list and a feature vocabulary list. This module helps the model understand and distinguish different entity types by providing a heuristic list containing the definitions of various entity types and related feature vocabulary. In legal documents, the semantics and forms of entities may vary due to various factors such as case types and description forms. The knowledge guidance module enables the model to better understand this information. With the guidance of this prior knowledge, the model can better understand the definitions of entities and correctly identify relevant information from the text, thus correctly extracting entities from complex legal texts and improving the accuracy of entity recognition. For example, criminal suspects are usually the subject or object of a sentence and may be related to words such as "suspect" and "defendant".
[0069] The heuristic list is shown in Table 1. Heuristics are defined as high-level rules or strategies for inferring specific tasks and play a crucial role in human cognition, which is usually more accurate in judgment than complex methods. Therefore, by setting the heuristic list shown in Table 1, which contains detailed strategies and methods for finding various entity types, it helps the model understand the manifestation forms of various entity categories in legal documents.
[0070] Table 1 Example of Heuristic List
[0071]
[0072]
[0073] The feature vocabulary list is shown in Table 2 and contains the feature vocabulary associated with each entity type. The model can recognize these feature vocabulary during the training process to help the model more accurately identify specific entity types, thereby further improving the accuracy of entity recognition.
[0074] Table 2 Example of Feature Vocabulary List
[0075]
[0076]
[0077] The core idea of the contrastive learning module is to select examples with highly similar semantics, TF-IDF, and dependency relationships from the transformed training set, and combine them with their corresponding pre-transformation instances to construct a pair of strongly contrastive learning samples. In the way of contrastive learning, it helps the model deeply understand the annotation rules combining context-aware data re-representation method and label prefix method, so that the model can not only understand the semantics and context information of entity types during the learning process, but also master how to output normalized results with label prefixes according to instructions, helping the model to more deeply understand the requirements of instructions and strengthening its accurate recognition of entity categories. The specific process is as follows:
[0078] Calculate the semantic similarity between the current text and the training examples: When calculating the semantic similarity, let the input text be T, and the i-th example in the training set be T i . First, pre-encode each example T in the training set i , and use the pre-trained language model RoBERTa to extract semantic embeddings e(T) and e(T i ). The semantic similarity between the two can be calculated by the cosine similarity formula:
[0079]
[0080] This formula ensures that the directional similarity of high-dimensional embedding vectors is captured, thus reflecting the semantic proximity of the two texts.
[0081] Calculate the TF-IDF similarity between the current text and the training examples: Adopt the TF-IDF weighted vector representation method, use the jieba library to perform Chinese word segmentation on all examples in the training set, and construct a bag-of-words model based on word frequencies. Specifically, let the vocabulary be V, and the TF-IDF representation of the input text be v(T) = [v1, v2, v j , …, v |V| , where v j represents the TF-IDF weight of the j-th word in the vocabulary. Similarly, the representation of the i-th example in the training set is v(T i ). The similarity calculation formula is:
[0082]
[0083] This formula effectively measures the overlap degree of texts in feature vocabulary and is suitable for capturing the relationship between keywords in the input text and the training examples.
[0084] Calculate the dependency similarity between the current text and the training examples: This is achieved by calculating the matching degree of the text dependency tree structure. Specifically, through the dependency parsing tool Spacy, extract the dependency tree structures D(T) and D(T i ) of the input text and the i-th example in the training set, and use the Graph Edit Distance (GED) to calculate the similarity between the two dependency trees. The formula is as follows:
[0085]
[0086] where GED represents the edit distance between the dependency trees, and |D(T)| and |D(T i )| are the number of nodes of the dependency trees respectively. This metric reflects the similarity in syntactic structure between the input text and the candidate examples.
[0087] Perform a weighted sum of the three metrics of semantic similarity, TF-IDF similarity, and dependency similarity to obtain the comprehensive similarity:
[0088] Sim total (T,T i ) = w1·Sim semantic (T,T i ) + w2·Sim TF-IDF (T,T i )
[0089] + w3·Sim dependency (T,T i )
[0090] where w1, w2, and w3 are the weights corresponding to Sim semantic , Sim TF-IDF and Sim dependency respectively, and satisfy w1 + w2 + w3 = 1. The example with the highest comprehensive similarity is selected as the contrast learning example T best :
[0091]
[0092] Match the example before transformation according to the example with the highest comprehensive similarity, and select the best contrast example: According to the example T best with the highest comprehensive similarity selected from the transformed training set, find the example T best-pre in the training set before transformation corresponding to its ID, and form the input of the contrast learning module with this pair of examples.
[0093] When training the large model for judicial entity recognition based on the above training dataset, the input F input of the model mainly consists of the following three parts:
[0094] Judicial text: X = {x1, x2, …, x n};
[0095] Instruction: I;
[0096] Triple understanding enhancement: F enhanced , which includes the specification module text, the knowledge guidance module text, and the contrast learning module text, that is
[0097] F enhanced = X formal + X knowledge + X contrast
[0098] Among them, X formal is the specification module text, X knowledge is the knowledge guidance module text, X contrast is the contrast learning module text.
[0099] F input = [X; I; F enhanced
[0100] As can be seen from the above, these three parts are concatenated into a semantic instruction template F input , which is used as input for the model to learn and understand. By including instructions in the specific context of the judicial field, the model can obtain richer context information from it, thereby enhancing its ability to understand complex language structures in judicial documents. This integration method effectively improves the model's accurate recognition and classification performance of entity categories, especially when facing long expressions and nested entities.
[0101] During model training, first, task definition and loss function setting are carried out: Given a judicial text X = {x1, x2, …, x n}, where x n represents the nth word of the input, and a user instruction I (for example: Please analyze the provided sentence and identify the type of each entity one by one. Please add a label prefix for each entity type in the output. If a word does not belong to any entity category, do not add a label for it.), and the triple understanding enhancement F enhanced , the goal is to output a sequence Y = {y1, y2, …, y m} with entity label prefixes. Each y m contains the text content and the corresponding label prefix. The goal of the model learning is the conditional probability distribution P(Y|F input ), which is generated word by word by the generative large model:
[0102]
[0103] Among them, y <t Denotes all the words generated before the t-th time step, P(Y|F input ) is the conditional probability distribution.
[0104] To optimize the gap between the sequence generated by the model and the target sequence, the negative log-likelihood loss function is used as the basic generation loss:
[0105]
[0106] where, -logP(y t |y <t ,F input ) is the prediction loss of the model at time step t, and the total loss of the entire output sequence is obtained by summation.
[0107] In addition, a constraint loss is set here to penalize predictions that violate entity specifications. Assume the model output sequence
[0108]
[0109] is the result generated by the decoder. The goal is to impose a penalty on the part that does not meet the rules. The general formula for the constraint loss is: t where λ is the penalty weight at time step t, which can be dynamically adjusted according to the rule importance, is the indicator function, which is 1 when the output violates the constraint and 0 otherwise.
[0110] is a conditional function that defines the situation of violating the rules. The constraint rules and violation judgments are defined as follows: The set of entity types C = {NHCS, NASI, NS, NHVI, NT, NCGV, NCSM, NO, NATS, NCSP} defines the entity categories that may appear in the task. Each prediction of the output must belong to one of the categories in this set. If the output
[0111]
[0112] Combining the above parts, the complete loss function can be obtained:
[0113]
[0114] That is,
[0115]
[0116] Then perform input vectorization: Use the pre-trained model Flan-T5 as the basic language model, denoted as H, and utilize the pre-trained language model BERT to map the input text F input to the embedding space to obtain the embedding representation:
[0117] E input = Embedding(F input )
[0118] where, n is the length of the input sequence, d is the embedding dimension, and Embedding(F input ) is to convert the input text F input into a high-dimensional vector.
[0119] Then perform encoder processing: Extract the context semantic representation from the embedding vector through the encoder (Encoder):
[0120] H enc = Encoder(E input )
[0121] where, H enc = {h1, h2,..., h T} is the output sequence representation of the encoder, containing the deep representation of the input text in the semantic space.
[0122] Then perform decoder generation: Generate the probability distribution of the next word according to the context and historical output at the current time step:
[0123] P(y t |y <t , F input ) = Decoder(H enc , y <t )
[0124] where, P(y t |y <t , F input ) represents the probability distribution that the decoder generates the current output y <t under the conditions of the generated historical output y input and the input feature F <t , and Decoder(H enc , y <t ) represents the probability distribution of generating the output at the current moment. The goal of the decoder is to maximize the probability generated at each step and finally generate a complete output sequence that meets the input features and task requirements.
[0125] After completing the model training, the user inputs the text to be recognized into the fine-tuned large model for judicial entity recognition, and the large model for judicial entity recognition can output the initial response.
[0126] S3. Gradually reason and verify the initial response through the VCoT verification mechanism, correct and optimize the verification results, and generate the finally verified entity recognition results.
[0127] In this embodiment, the self-verification module is used to ensure the accuracy and consistency of the entity recognition results. It can be understood that although the entity information extraction module has achieved high-precision entity recognition through instruction fine-tuning and triple understanding enhancement, due to the often complex grammatical structures in judicial documents and the inherent uncertainties in the generation of large language models, the prediction length of the model increases with the integration of the context-aware data re-representation strategy and the label prefix annotation method. Longer generation sequences may pose challenges to large language models, and there may be certain misrecognition or annotation problems, including word omission, addition, and replacement. With the gradual popularity of the Chain of Thought (COT) technology, in this embodiment, based on COT, the Verification with Chain of Thought (VCoT) verification mechanism is innovatively designed to gradually verify the entity recognition results to discover and correct potential errors.
[0128] Among them, the VCoT verification mechanism gradually checks whether each recognized entity is consistent with the context in the original text and ensures that there are no omissions or inconsistencies in the annotations generated by the model. Through this self-correction mechanism, the model can better process complex judicial documents, avoid misrecognition, and accurately capture all key entities in the text.
[0129] In this embodiment, the VCoT verification mechanism gradually generates a series of inference chains through multiple rounds of inference verification to ensure that the entities generated by the model are not only correctly recognized but also conform to the actual context and logical relationship of the context. For example, after the model recognizes "Zhang San" as a suspect, it will infer whether this entity matches the specific circumstances of the case to confirm whether there is a mislabel.
[0130] Optionally, the content of the inference chain includes entity type consistency check: whether it is a predefined entity type; context consistency check: ensuring that the generated content has no omissions, additions, and deviations from the original content; logical check: checking whether the context collocation of the entity conforms to logic (for example, a location cannot be labeled as "person"). Through gradual verification, the model can timely discover errors such as omission, addition, and replacement, especially with better effects in complex sentences. And through gradual reasoning, it helps the model verify whether the entity tags conform to the sentence logic, thus avoiding mislabeling.
[0131] The detailed process of the VCoT verification mechanism is further described through the following specific process: The user inputs the judicial text to be recognized into the fine-tuned large model for judicial entity recognition, and conducts a series of verifications based on the initial response obtained from the large model for judicial entity recognition; checks the integrity of the generated content and the original sentence to ensure that there are no missing, redundant, or replaced words; lists the entity types and entities in each tag prefix based on the obtained initial response; generates a series of context verification questions based on the obtained tag list, which helps to self-analyze whether there are any errors in the original response; answers each verification question in turn, and then checks the answer against the initial response to check for inconsistencies or errors, further improving the recognition performance of the model; generates the final verification response. As can be seen from the above, the VCoT verification mechanism performs self-correction based on reasoning to ensure that the results of entity recognition are both accurate and in line with the actual case situation. In view of the discovered inconsistencies (if any), a revised response containing the verification results is generated to ensure that the final output is optimal.
[0132] As a preferred implementation manner, in this embodiment, through the rational design of the prompt in the VCoT verification mechanism, it is ensured that the final output not only meets the requirements of the instruction, but also can handle potential ambiguities and complex situations in judicial texts, providing secondary correction, thereby improving the reliability of the recognition results.
[0133] Specifically, the following design can be carried out: Please verify the integrity of the following generated sentence and the original sentence to ensure that there are no missing, redundant, or replaced words, and gradually check the accuracy of each tag and ensure its consistency with the context:
[0134] First, list all entity tags and their corresponding entity types in the generated response. For example, NHCS: Zhang XX, NASI: Samsung mobile phone. For each entity type, check whether it conforms to the predefined types (NHCS, NASI, NS, NHVI, NT, NCGV, NCSM, NO, NATS, NCSP).
[0135] Then, verify the following content in turn:
[0136] For the suspect (NHCS), confirm its reasonable use in the context of describing criminal acts or investigations; for the stolen item (NASI), ensure its appearance in the context of describing loss or theft; for the location (NS), check its reasonableness in the context of describing the place where the event occurred; for the victim (NHVI), confirm whether the role appears in the context related to the criminal act, investigation or legal procedure of the case; for the time (NT), check whether it appears in the context of describing the time of the event; for the value of the item (NCGV), verify its context in describing the value of the property; for the stolen currency (NCSM), confirm its appearance in the context involving monetary loss; for the organization (NO), check whether it is in the context of describing an organization or institution; for the tool used in the crime (NATS), ensure its context in describing the tool used in the crime; for the illegal proceeds (NCSP), verify whether it appears in the context of describing criminal profits, illegal gains or illegal funds; answer each verification question one by one and check for inconsistencies or errors; according to the verification questions executed, if any inconsistencies or errors are found, provide revision suggestions and generate a final verification response. If all verifications are correct, generate the following response: "The entity annotation conforms to the context, verification passed, no revision required.", if errors or inconsistencies are found, generate a revised version.
[0137] Sentence: <Original sentence>
[0138] Generated annotation result: <Annotation generated by the model>
[0139] In summary, the judicial text entity recognition method based on triple instruction fine-tuning and VCoT verification provided in this embodiment improves the model's ability to distinguish entity and non-entity information through triple understanding enhanced combined instruction fine-tuning on the basis of instruction fine-tuning and Chain of Thought (CoT) technology, thereby reducing the phenomenon of mis-identification. Through the verification chain reasoning based on VCoT, by generating a problem list and a self-verification mechanism, the recognition performance of the model is further improved.
[0140] The following further illustrates the role and effect of this method through a specific complete case. Input instruction: Help me extract all entities from the following sentence. Sentence: Zhang Moumou smashed the glass with a hammer in Room XX, Unit XX, XX City and stole a Samsung mobile phone and 600 yuan in cash.
[0141] (a) Generate the initial response:
[0142]
NHCS: Zhang Moumou
NATS: hammer
NS: Room XX, Unit XX, XX City
NASI: Samsung mobile phone
NCSP: 600 yuan
[0143] (b) Sentence integrity check:
[0144] The text content of the generated sentences is consistent with the original text.
[0145] (c) Generate a list of tags:
[0146] NHCS: Zhang XX; NS: Room XX, Unit XX, XX City; NATS: hammer; NASI: Samsung mobile phone; NCSP: 600 yuan.
[0147] (d) Entity type consistency check:
[0148] Entity types conform to predefined types.
[0149] (e) Design contextual verification questions:
[0150] Confirm whether "Zhang XX" appears in the context of reasonable use related to the criminal behavior, investigation or legal proceedings of the case; confirm whether "Room XX, Unit XX, City XX" is in the context of describing the place where the incident occurred; confirm whether "hammer" appears in the context of describing the tool used in the crime; confirm whether "Samsung mobile phone" appears in the context of describing loss or theft; confirm whether "600 yuan" is in the context of describing criminal profits, illegal gains or illegal funds.
[0151] (f) Execution Verification:
[0152] The entity annotation is consistent with the context, the verification is passed, and no revision is required; the entity annotation is consistent with the context, the verification is passed, and no revision is required; the entity annotation is consistent with the context, the verification is passed, and no revision is required; the entity annotation is consistent with the context, the verification is passed, and no revision is required; the entity annotation does not conform to the context and needs to be revised.
[0153] (g) Generate the final verification response:
[0154] NCSP (theft proceeds) has been revised to NCSM (stolen money) because "600 yuan" should be marked as stolen cash rather than theft proceeds. No inconsistencies were found in other parts and no revision is required.
[0155] The final labeling results are as follows: NHCS: Zhang XX; NS: Room XX, Unit XX, City XX; NATS: hammer; NASI: Samsung mobile phone; NCSM: 600 yuan.
[0156] Please refer again Figure 2, based on the same inventive concept, this embodiment also provides a judicial text entity recognition system based on ternary instruction fine-tuning and VCoT verification, including an input module, an entity information extraction module, and a self-verification module. Among them, the input module includes a user instruction sub-module and a user input sub-module; the entity information extraction module includes a large judicial entity recognition model, which optimizes the model by combining ternary understanding enhancement with instruction fine-tuning; the self-verification module has a VCoT verification mechanism. The functions of each module and sub-module have been given above and will not be elaborated here.
[0157] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0158] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A judicial text entity recognition method based on ternary instruction fine-tuning and VCoT verification, characterized in that It includes the following steps: S1. The user inputs the judicial document to be recognized and instructions; S2. An initial response is obtained through the large model for judicial entity recognition; The large model for judicial entity recognition is trained and optimized through instruction fine-tuning enhanced by triple understanding to improve the understanding ability of judicial texts and accurately identify different types of entities; Instruction fine-tuning guides the large model for judicial entity recognition to identify various key entity information from judicial documents through explicit instruction design and the injection of domain knowledge; Triple understanding enhancement improves the entity recognition performance in judicial documents through deep semantic understanding and context adaptation ability; S3. The initial response is gradually inferred and verified through the VCoT verification mechanism, and the verification result is corrected and optimized to generate the final verified entity recognition result; The VCoT verification mechanism gradually checks whether each recognized entity is consistent with the context in the original text and ensures that the annotations generated by the model are not missing or inconsistent; The VCoT verification mechanism gradually generates a series of inference chains through multiple rounds of inference verification, so that the entities generated by the model are not only correctly recognized, but also conform to the actual context and logical relationship of the context.
2. The judicial text entity recognition method based on ternary instruction fine-tuning and VCoT verification according to claim 1, characterized in that, In S1, the input module receives the judicial document to be recognized and instructions provided by the user.
3. The method for identifying judicial text entities based on ternary instruction fine-tuning and VCoT verification according to claim 1, wherein, Instruction fine-tuning includes: By using context-aware data re-representation, the context around the entity is regarded as a key non-entity sample, so that the model can learn how to more accurately distinguish entity and non-entity information through samples; The dataset is re-annotated using the label prefix annotation method, and a unique prefix label is added to each entity type, so that the model can accurately distinguish various fine-grained entities when generating output.
4. The method for identifying judicial text entities based on ternary instruction fine-tuning and VCoT verification according to claim 3, wherein Instruction fine-tuning also includes: A task mode of generating while predicting, so that the model will predict in real time whether the word it is currently generating belongs to a specific entity category when generating text, so as to decide whether to add an entity label to it. For ordinary words that do not belong to any entity, the model will keep its original form without adding any labels.
5. The judicial text entity recognition method based on ternary instruction fine-tuning and VCoT verification according to claim 4, wherein Triple understanding enhancement includes: a specification module, a knowledge guidance module, and a contrast learning module; The specification module is used to standardize the generated content to ensure the accuracy, integrity, and consistency of the output; The knowledge guidance module is used to provide a heuristic list containing the definitions of each entity type and related feature words to help the model understand and distinguish different entity types; the heuristic list contains high-level rules or strategies for inferring specific tasks, and the feature word list contains feature words associated with each entity type; The contrast learning module is used to select examples with highly similar semantics, TF-IDF, and dependency relationships from the transformed training set, and combine them with the corresponding pre-transformation instances to construct a pair of strongly contrasting learning samples, and use contrast learning to help the model deeply understand the annotation rules combining context-aware data re-representation and label prefix method, so that the model can not only understand the semantics and context information of entity types during the learning process, but also master how to output normalized results with label prefixes according to instructions.
6. The method for identifying judicial text entities based on ternary instruction fine-tuning and VCoT verification according to claim 5, wherein When training the judicial entity recognition large model in S2, the input F of the model input consists of the following three parts: Legal text: X = {x1, x2, …, x n}; Instruction: I; Three-way understanding enhancement: F enhanced , which includes the specification module text, the knowledge guidance module text, and the contrastive learning module text, namely F enhanced = X formal + X knowledge + X contrast Among them, X formal is the standardized module text, X knowledge is the knowledge guidance module text, X contrast is the contrastive learning module text; F input = [X; I; F enhanced During model training, task definition and loss function setting are performed first; then input vectorization is carried out; then encoder processing is performed; and finally decoder generation is carried out.
7. The judicial text entity recognition method based on ternary instruction fine-tuning and VCoT verification according to claim 1, wherein In S3, the self-verification module ensures the accuracy and consistency of entity recognition results, and the VCoT verification mechanism is nested in the self-verification module.
8. The method for identifying judicial text entities based on ternary instruction fine-tuning and VCoT verification according to claim 7, wherein, The content of the inference chain includes: Entity type consistency check: whether it is a predefined entity type; context consistency check: ensuring that the generated content has no omissions, additions, or deviations from the original content; logical check: checking whether the context collocation of entities is logical.
9. The method for identifying judicial text entities based on ternary instruction fine-tuning and VCoT verification according to claim 1, wherein The VCoT verification mechanism designs prompt words so that the final output not only meets the requirements of the instructions but also can handle potential ambiguities and complex situations in judicial texts, providing secondary correction.
Citation Information
Patent Citations
Judicial named entity recognition method based on natural language processing
CN115238697A
Entity recognition model training method, entity recognition method, device and equipment
CN118798197A
Information processing method and device, electronic equipment and storage medium
CN119338001A