Method for evaluating Chinese law text language profession degree generated by large language model

By comparing real documents and simulated documents and training the knowledge base, the problem of difficulty in evaluating the professionalism of Chinese legal texts in the existing technology is solved, and efficient evaluation without reference texts is achieved, improving the accuracy and pertinence of the evaluation.

CN119990880AInactive Publication Date: 2025-05-13马义然
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510073722.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to effectively evaluate the language expertise of Chinese legal texts generated by large language models, especially in terms of logic, accuracy of professional term usage, and domain-specific expression ability.

Method used

By simplifying real documents, generating simulated documents, comparing real documents with simulated documents, training the knowledge base, and finally, based on training knowledge base, directly assessing the language professionalism of the document without reference to the text.

Benefits of technology

It has achieved the ability to generate quality evaluation without reference texts, adapt to more actual scenario needs, improved the ability to evaluate legal language style, professional term accuracy and logical structure rigor, and provided fine-grained scores and clear evaluations and suggestions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990880A_ABST
    Figure CN119990880A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of Chinese legal text generation, and discloses a method for evaluating the language profession degree of Chinese legal text generated by a large language model, which simplifies a real document, converts the real document into a key content abstract, only reserves core information in a judgment document, removes the legal language property of the real document, and improves the language profession degree of the Chinese legal text generated by the large language model. Then generating a complete simulation document according to the key content abstract and a preset prompt word, training and learning non-parameterized knowledge base dynamic training through comparison, evaluating the Chinese legal text language professional degree of the document based on the trained knowledge base, and finally combining shallow linguistics feature analysis and legal term library verification. Through combination of non-reference evaluation, a dynamic knowledge base, shallow linguistic feature analysis and external term library verification, multiple limitations in language profession evaluation of the generated text in the prior art are overcome, an efficient, accurate and explanatory evaluation result can be provided, and a solid technical support is provided for improvement of Chinese legal text generation quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of Chinese legal text generation, and specifically relates to a method for evaluating the language expertise of Chinese legal text generated by a large language model. Background Art

[0002] In recent years, the rapid development of Large Language Models (LLMs) has promoted the widespread application of Natural Language Generation (NLG) technology. In this context, it is particularly important to effectively evaluate the quality of generated text. However, existing evaluation methods have many limitations and it is difficult to fully and accurately reflect the language expertise and semantic quality of generated text.

[0003] In the field of Chinese legal text generation, it is particularly important to evaluate the language expertise of the generated text. Most of the existing evaluation methods are aimed at general texts, and lack special optimization for professional languages ​​in the legal field. For example, methods such as BLEU and ROUGE cannot reflect the logic of legal texts and the accuracy of the use of professional terms; although the embedding method can capture some semantic information, it has weak expression capabilities for professionalism and field-specificity; and the evaluation method based on large language models has not yet been customized for the special field of Chinese legal texts. In addition, the above methods all require reference texts when evaluating the performance of model generation, but in real-life scenarios where the legal language expertise of the model needs to be evaluated, reference texts are scarce or difficult to obtain.

[0004] Therefore, existing technologies cannot meet the needs of efficient and accurate evaluation of the quality of Chinese legal text generation, especially in the evaluation of language expertise. These problems urgently need to be solved through innovative methods to improve the pertinence and reliability of the evaluation, so as to better support the application of large language models in the legal field. Summary of the invention

[0005] In order to solve the above problems, the present invention proposes a method for evaluating the language expertise of Chinese legal texts generated by a large language model, which is intended to solve the problems mentioned in the above background documents. The present invention is implemented in the following ways:

[0006] A method for evaluating the language expertise of Chinese legal texts generated by a large language model, comprising the following steps:

[0007] S1. Simplify the real document;

[0008] S2. Generate simulation documents;

[0009] S3. Compare the real documents and simulated documents to train the knowledge base;

[0010] S4. Evaluate the language expertise of Chinese legal texts based on the trained knowledge base.

[0011] As a preferred embodiment of the present invention, in S1, the judgment document is input as a high-quality text source based on the CAIL 2021 public judgment document data; then the MiniCPM-2B model and preset prompt words are used to convert the real document into a summary of its key content, retaining only the core information in the judgment document and removing the legal language of the real document.

[0012] As a preferred embodiment of the present invention, in S2, a complete simulated document is generated by using the MiniCPM-2B model on the key content summary of the real document and using preset prompt words.

[0013] As a preferred embodiment of the present invention, in S3, the Qwen-2.5-72B model is used to compare real documents with simulated documents, focusing on and summarizing empirical knowledge on language use. On the one hand, the model summarizes the high-quality features of real documents under the guidance of preset prompt words; on the other hand, the model can produce experience based on problems and deviations found in the comparison process.

[0014] As a preferred embodiment of the present invention, in S3, the Qwen-2.5-72B model extracts key operational knowledge; all the extracted knowledge will be stored in the knowledge base file outside the model in the form of "if...then..." rules, and another Qwen-2.5-72B model is set up for rule management such as redundancy removal and validity verification. These knowledge are stored in the knowledge base in a non-parametric manner, which is easy for human experts to read and review, and can be continuously learned, dynamically updated and expanded through different input data to meet the needs of legal text generation in different fields.

[0015] As a preferred embodiment of the present invention, in S4, the Qwen-2.5-72B model directly calls the rule knowledge stored in the knowledge base without referring to the text to perform quality scoring on the generated document.

[0016] As a preferred implementation of the present invention, the quality scoring step includes: scoring in units of sentences, retrieving rule knowledge available in the knowledge base according to the sentence content, and giving:

[0017] 1) A brief evaluation of each sentence;

[0018] 2) Sentence-level professionalism score: 0-100 points;

[0019] 3) The paragraph-level professionalism score after averaging all the sentences in the article: 0-100 points.

[0020] As a preferred embodiment of the present invention, in S4, the depth and accuracy of the evaluation are improved by introducing an auxiliary module, and the auxiliary module includes a shallow linguistic feature module and an external terminology library integration module;

[0021] The shallow linguistic feature module uses a multi-layer perceptron and an attention mechanism to build a neural network, introduces non-legal document negative samples for training, and the input layer receives a number of numerical features and outputs a shallow linguistic feature professional score between [0,1]. This feature professional score is input into the Qwen-2.5-72B model during training or evaluation as a reference.

[0022] The external terminology library integration module uses open source legal terminology libraries to directly retrieve keywords during training or evaluation by calling model tools, thereby verifying whether the use of legal terminology in the generated documents complies with the regulations.

[0023] Compared with the prior art, the present invention has the following beneficial effects:

[0024] (1) The present invention constructs a knowledge base and uses rules extracted from the comparison between real documents and generated documents to achieve the generation quality assessment capability without reference text, which can meet the needs of more practical scenarios.

[0025] (2) The present invention introduces a non-parametric knowledge base dynamic training and continuous learning mechanism, which continuously updates the rules through online learning to adapt to different fields and text generation needs. At the same time, the rule form of the knowledge base is highly scalable and can quickly adapt to new tasks or new fields of text generation quality assessment.

[0026] (3) During the evaluation process, the model can provide fine-grained sentence-level and paragraph-level scores, and generate clear evaluations and specific suggestions for each sentence based on knowledge base rules. For example, the model can point out the specific location of misused terms or sentences with logical breaks, thereby providing clear and specific guidance for optimizing and adjusting the generation model.

[0027] (4) The present invention combines shallow linguistic feature analysis with legal terminology database verification, significantly improving the ability to evaluate legal language style, professional terminology accuracy, and logical structure rigor. In particular, in the specific field of Chinese legal text generation, the evaluation results are more accurate and meet the needs of the field.

[0028] (5) The knowledge base is stored in the form of "if...then..." rules, which is convenient for human experts to review and expand. At the same time, the modular design of the present invention allows it to be flexibly embedded in the existing legal text generation or evaluation system, making it convenient for users to adjust and optimize according to actual needs.

[0029] (6) The knowledge base construction process and evaluation mechanism of the present invention are not limited to the legal field. By introducing new domain data and terminology libraries, it can be extended to other professional text generation tasks, such as medicine, finance and other fields, to achieve broader application value.

[0030] The specific implementation modes of the present invention are further described in detail below in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In the attached picture:

[0032] Figure 1 A flow chart of a method for evaluating the language expertise of Chinese legal texts generated by a large language model according to the present invention. DETAILED DESCRIPTION

[0033] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. The following embodiments are used to illustrate the present invention.

[0034] like Figure 1 As shown, a method for evaluating the language expertise of Chinese legal texts generated by a large language model of the present invention comprises the following steps:

[0035] S1. Simplify the real document;

[0036] In this embodiment, the CAIL 2021 public judgment document data is used as a reference document as a high-quality text source, and then the real judgment document is input;

[0037] Then, the MiniCPM-2B model and specific prompt words are used to convert the real document into a summary of its key content, retaining only the core information in the judgment document and removing the legal language. The MiniCPM-2B model with a small number of parameters is used here to control the language style to focus on the facts of the case.

[0038] S2. Generate simulation documents;

[0039] After converting the real document into a summary of its key content, the MiniCPM-2B model was used again to generate a complete simulated document through specific prompt words. The generated simulated document is a representation of the generation capability of the small parameter model, which can show potential language problems, including style, logic and terminology.

[0040] S3. Compare the real documents and simulated documents to train the knowledge base;

[0041] Then, the Qwen-2.5-72B model is used to focus on and summarize the empirical knowledge on language use by comparing the real documents with the generated documents. The Qwen-2.5-72B model has a larger number of parameters and stronger language ability, so it is suitable for capturing complex language features.

[0042] On the one hand, the Qwen-2.5-72B model summarizes the high-quality features of real documents under the guidance of prompt words; on the other hand, the Qwen-2.5-72B model can produce experience based on the problems and deviations found in the comparison process, such as non-standard language expression and misuse of legal terms.

[0043] Based on observations and reflections from both aspects, the model extracts some key operational knowledge. All extracted knowledge will be stored in the knowledge base file outside the model in the form of "if...then..." rules, and a Qwen-2.5-72B model will be set up for rule management such as redundancy removal and validity verification.

[0044] The above process is called non-parametric training of the knowledge base. This knowledge is stored in the knowledge base in a non-parametric way, which is easy for human experts to read and review, and can be continuously learned, dynamically updated and expanded through different input data to meet the needs of legal text generation in different fields.

[0045] S4. Evaluate the language expertise of Chinese legal texts based on the trained knowledge base;

[0046] The document to be evaluated is evaluated based on the trained knowledge base. During the evaluation, the Qwen-2.5-72B model does not need to refer to the text, but directly calls the rule knowledge stored in the knowledge base to score the quality of the document.

[0047] Specifically, after the model receives a new document, it scores it sentence by sentence, retrieves the rule knowledge available in the knowledge base based on the sentence content, and gives:

[0048] 1) A brief evaluation of each sentence;

[0049] 2) Sentence-level professionalism score: 0-100 points;

[0050] 3) The paragraph-level professionalism score after averaging all the sentences in the article: 0-100 points.

[0051] The knowledge base content provides explainability for the evaluation process and can point out deviations in specific language features or errors in terminology usage in a fine-grained manner, which is more consistent with the evaluation of human experts.

[0052] Finally, based on the core process, the present invention introduces two important auxiliary modules to improve the depth and accuracy of the evaluation.

[0053] Shallow linguistic features: On a large number of real legal documents, we use natural language processing tools to extract basic linguistic numerical features, such as part of speech, word frequency distribution, proportion of commonly used characters, sentence length, etc. We combine these features to quantify the professionalism of the legal language style, train the regression model, and obtain the professionalism of shallow linguistic features, which are input into the evaluation model for reference. Specifically, we use a multi-layer perceptron and attention mechanism to build a neural network, introduce negative samples of non-legal documents for training, and the input layer receives several numerical features and outputs a shallow linguistic feature professionalism score between [0,1]. This score is input into the Qwen-2.5-72B model during training or evaluation as a reference.

[0054] External terminology library integration: Utilize open source legal terminology libraries to directly retrieve keywords during training or evaluation by calling model tools, thereby verifying whether the use of legal terminology in the generated documents complies with regulations.

[0055] The above is only a preferred implementation case of the present disclosure. Although the present disclosure is described in conjunction with the accompanying drawings, the purpose is not to limit the present disclosure. For those skilled in the art, the present disclosure may have various changes and modifications. Any modification, replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A method for evaluating the language expertise of Chinese legal texts generated by a large language model, characterized in that: The following steps are involved: S1. Simplify the real document; S2. Generate simulation documents; S3. Compare the real documents and simulated documents to train the knowledge base; S4. Evaluate the language expertise of Chinese legal texts based on the trained knowledge base.

2. According to claim 1, a method for evaluating the language expertise of Chinese legal texts generated by a large language model is characterized in that: In S1, the judgment document is input as a high-quality text source based on the CAIL 2021 public judgment document data; then the MiniCPM-2B model and preset prompt words are used to convert the real document into a summary of its key content, retaining only the core information in the judgment document and removing the legal language of the real document.

3. The method for evaluating the language expertise of Chinese legal texts generated by a large language model according to claim 1, characterized in that: In S2, a complete simulated document is generated by using the MiniCPM-2B model on the key content summary of the real document and using preset prompt words.

4. According to claim 1, a method for evaluating the language expertise of Chinese legal texts generated by a large language model is characterized in that: In S3, the Qwen-2.5-72B model is used to compare real documents with simulated documents, focusing on and summarizing empirical knowledge on language use. On the one hand, the model summarizes the high-quality features of real documents under the guidance of preset prompt words; on the other hand, the model can produce experience based on the problems and deviations found in the comparison process.

5. The method for evaluating the language expertise of Chinese legal texts generated by a large language model according to claim 1, characterized in that: In S3, the Qwen-2.5-72B model extracts key operational knowledge; all the extracted knowledge will be stored in the knowledge base file outside the model in the form of "if...then..." rules, and another Qwen-2.5-72B model is set up for rule management such as removing redundancy and validity verification. This knowledge is stored in the knowledge base in a non-parametric manner, which is easy for human experts to read and review, and can continue to learn, dynamically update and expand through different input data to meet the needs of legal text generation in different fields.

6. The method for evaluating the language expertise of Chinese legal texts generated by a large language model according to claim 1, characterized in that: In S4, the Qwen-2.5-72B model directly calls the rule knowledge stored in the knowledge base without referring to the text to perform quality scoring on the generated documents.

7. A method for evaluating the language expertise of Chinese legal texts generated by a large language model according to claim 6, characterized in that: The quality scoring step includes: scoring in units of sentences, retrieving rule knowledge available in the knowledge base according to the sentence content, and giving: 1) A brief evaluation of each sentence; 2) Sentence-level professionalism score: 0-100 points; 3) The paragraph-level professionalism score after averaging all the sentences in the article: 0-100 points.

8. The method for evaluating the language expertise of Chinese legal texts generated by a large language model according to claim 1, characterized in that: In S4, the depth and accuracy of the evaluation are improved by introducing auxiliary modules, which include a shallow linguistic feature module and an external terminology library integration module; The shallow linguistic feature module uses a multi-layer perceptron and an attention mechanism to build a neural network, introduces non-legal document negative samples for training, and the input layer receives a number of numerical features and outputs a shallow linguistic feature professional score between [0,1]. This feature professional score is input into the Qwen-2.5-72B model during training or evaluation as a reference. The external terminology library integration module uses open source legal terminology libraries to directly retrieve keywords during training or evaluation by calling model tools, thereby verifying whether the use of legal terminology in the generated documents complies with the regulations.