Automatic case classification method and system based on large model context learning and medium

By building an automatic case classification system based on large models and using large models to study and predict case documents in context, the problems of low case classification accuracy and system complexity in the existing technology are solved, and high accuracy and low cost case classification are achieved.

CN120031024APending Publication Date: 2025-05-23SHANGHAI DATA GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311539276.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-17
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing referee prediction system is complex in construction and use, the model algorithm is low in accuracy, and it is difficult to achieve detailed case classification prediction, and it requires frequent retraining to adapt to new discretionary benchmarks.

Method used

The automatic case classification method based on the context learning of the big model is adopted. By obtaining the case documents to be classified, a crime classification model and a case database are constructed, a large model is used for classification prediction, and a combination of manual deviation correction is used to improve the accuracy of the model.

Benefits of technology

High accuracy of case classification is achieved, complexity of system construction and use is reduced, labor cost requirements are reduced, and model adaptability and performance are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031024A_ABST
    Figure CN120031024A_ABST
Patent Text Reader

Abstract

The invention relates to an automatic case classification method and system based on large model context learning, and a medium. The method comprises the following steps: obtaining a to-be-classified case document; inputting the to-be-classified case documents into a pre-constructed crime name classification model to obtain a crime name classification reference; compressing the to-be-classified case documents into vector representation, taking the vector representation and the crime name classification reference as an index, and retrieving approximate case documents from a pre-constructed class case library; and constructing prompts based on the to-be-classified case documents, the crime name classification benchmark and the approximate case documents, and submitting the prompts to a pre-trained large model for classification prediction to complete an automatic case classification process. Compared with the prior art, the method has the advantages of being high in classification accuracy, providing more accurate case trial suggestions for the judge and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, system and medium for automatic case classification based on large model context learning. Background Art

[0002] Discretion refers to making judgments and handling according to one's own understanding. In the legal sense, administrative discretion refers to the way, method or form in which administrative subjects and their staff (hereinafter referred to as "administrative subjects") make judgments and handle according to their own understanding, based on the scope, limits and even standards or principles set by legal norms. The administrative penalty discretion benchmark (hereinafter referred to as "discretion benchmark") refers to the basis for administrative subjects to refine and quantify the circumstances of specific illegal acts within the scope of discretion prescribed by laws, regulations and rules to determine whether to impose penalties, what kind of penalties to impose and what range of penalties to impose. In the specific implementation process, the discretion benchmark can be regarded as a penalty discretion table. After clarifying the facts and charges of violations of laws and regulations, and fully considering their causes, nature, circumstances, consequences, charges and other factors, the corresponding penalty consequences are selected.

[0003] Existing legal judgment prediction models usually break down tasks into three categories: legal article prediction, crime prediction, and sentence prediction. There is no task form that completely overlaps with discretion. Compared with crime prediction, discretion has a finer granularity, and there may be multiple discretionary benchmarks to choose from under any crime. Compared with sentence prediction, the content of the discretionary benchmark is not limited to the length of time and only gives the range of possible penalties, while sentence prediction is usually regarded as a regression problem. To some extent, discretion is an independent intermediate link between crime and sentence in legal judgments. For a discretionary prediction system, given the legal document corresponding to a case, the prediction system gives the applicable punishment consequences.

[0004] At present, the "judgment prediction system" combined with artificial intelligence technology has been applied in some local courts to help judges make sentencing decisions. The existing "judgment prediction system" generally adopts a combination of knowledge graphs and first-order predicate logic, as well as a method based on deep learning. After the judge imports legal documents such as indictments and trial records into the system at the sentencing stage, the system uses internal functions or neural network operations to provide reference suggestions for sentencing. However, these problems exist, such as complex construction and use, and low accuracy of model algorithms. For example, the criminal case intelligent assistance system of a certain intermediate people's court includes 26 major categories and 88 sub-functions, and the system construction and use are complex. At the same time, the process of customizing corresponding certification standards and knowledge graphs for different situations has a low degree of automation and high professional level requirements, resulting in huge labor cost requirements. In addition, the low accuracy of early natural language processing algorithms and poor model reasoning capabilities are also a major factor restricting the final system performance. At the same time, with the development of legal system construction, earlier models need to be retrained to adapt to new discretionary benchmarks.

[0005] At present, large-scale language models represented by ChatGPT have demonstrated high language comprehension capabilities, and have greatly improved the "contextual" reasoning capabilities on zero-sample complex tasks compared to previous deep learning models. It provides a theoretical basis and technical guidance for building a higher-performance "case classification prediction system". In particular, the GPT-4 model released in March 2023 reached a level close to or exceeding that of humans in the US Bar Examination; large models in the Chinese judicial vertical field represented by LawGPT expand the legal field's proprietary vocabulary and large-scale Chinese legal corpus pre-training on the basis of the general Chinese base model to enhance the model's basic semantic understanding capabilities in the legal field, but the above models are mainly aimed at the writing of legal documents or common test questions in judicial examinations, and cannot be used for the new task form of detailed case classification prediction. Summary of the invention

[0006] The purpose of the present invention is to provide a method, system and medium for automatic case classification based on large model context learning with high case classification accuracy.

[0007] The purpose of the present invention can be achieved by the following technical solutions:

[0008] A case automatic classification method based on large model context learning includes the following steps:

[0009] Obtain case documents to be classified;

[0010] Inputting the case documents to be classified into a pre-built crime classification model to obtain a crime classification benchmark;

[0011] Compress the case document to be classified into a vector representation, use it together with the crime classification standard as an index, and retrieve similar case documents from a pre-built similar case library;

[0012] Based on the case documents to be classified, the crime classification criteria and similar case document construction prompts, and handed over to the pre-trained large model for classification prediction, the automatic case classification process is completed.

[0013] Furthermore, the steps for constructing the crime classification model are specifically as follows:

[0014] Obtain crime prediction dataset;

[0015] The text classification model is called and fine-tuned on the crime prediction dataset to obtain a crime classification model.

[0016] Furthermore, the text classification model is a LongFormer model.

[0017] Furthermore, the steps of constructing the case library are specifically as follows:

[0018] Access to a large number of existing case documents;

[0019] Scanning and digitally extracting the case documents using an optical character recognition method and a text vectorization method to obtain the case documents represented by character codes;

[0020] The case details in the case document represented by character code are compressed into a vector form through a large model to form a vectorized case library.

[0021] Furthermore, the ChatGLM model is used to compress the case documents to be classified into vector representation.

[0022] Furthermore, it also includes: manually correcting the crime classification criteria and similar case documents.

[0023] Furthermore, the specific steps of the correction include:

[0024] Manually screen or add the crime classification criteria and similar case documents;

[0025] Based on the case documents to be classified, crime classification benchmarks and similar case documents are obtained again;

[0026] The above process is repeated until the final crime classification criteria and approximate case documents are obtained.

[0027] Furthermore, the large model is a GPT-3 series model.

[0028] The present invention also provides a case automatic classification system based on large model context learning, comprising:

[0029] Acquisition module: Acquisition of case documents to be classified;

[0030] Classification module: inputting the case documents to be classified into a pre-built crime classification model to obtain a crime classification benchmark;

[0031] Retrieval module: used to compress the case documents to be classified into vector representation, use it together with the crime classification standard as an index, and retrieve similar case documents from a pre-built similar case library;

[0032] Classification module: used to construct prompts based on the case documents to be classified, the crime classification criteria and similar case documents, and submit them to the pre-trained large model for classification prediction to complete the automatic case classification process.

[0033] The present invention also provides a computer-readable storage medium, comprising one or more programs for execution by one or more processors of an electronic device, wherein the one or more programs include instructions for executing the method for automatic case classification based on large model context learning as described in any one of claims 1-8.

[0034] Compared with the prior art, the present invention has the following beneficial effects:

[0035] (1) The present invention makes full use of existing cases. By constructing a case database, it can retrieve case documents similar to the case to be classified. At the same time, it fully utilizes the advantages of the big model to learn the context of the case, improves the understanding ability of the big model by constructing prompts, and outputs more accurate and detailed classification results.

[0036] (2) The present invention can modify the case classification criteria and similar case texts through manual correction, further improving the accuracy of model classification predictions and providing more accurate reference suggestions for judges to make corresponding case decisions.

[0037] (3) The crime classification model of the present invention is obtained through a text classification model, and the large model is predicted through sample classification. There is no need to spend a lot of time on model training, model parameter modification and other steps. It has low computing power requirements, fast system iteration, and is simple to implement. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 A flowchart of a method implementation of an embodiment of the present invention;

[0039] Figure 2 This is a template prompt structure diagram of an embodiment of the present invention;

[0040] Figure 31 is an overall architecture diagram of the system in an embodiment of the present invention. DETAILED DESCRIPTION

[0041] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0042] Example 1

[0043] This embodiment provides a case automatic classification method based on large model context learning, such as Figure 1 As shown, the method comprises the following steps:

[0044] S1. Obtain case documents to be classified.

[0045] S2. Input the case documents to be classified into a pre-built crime classification model to obtain a crime classification benchmark.

[0046] This step requires building a crime classification model in advance. The model is obtained by calling the text classification model and fine-tuning it on the crime prediction dataset. The text classification model is built through the LongFormer model. The LongFormer model is good at processing long text sequences and is helpful for processing case texts.

[0047] The case documents to be classified are input into the above-mentioned crime classification model, and the possible crime classification criteria of the case documents to be classified are classified and predicted based on the case plot paragraphs as the main judgment basis, so as to assist in the retrieval of similar case documents.

[0048] S3. Compress the case documents to be classified into a vector representation, use it together with the crime classification standard as an index, and retrieve similar case documents from a pre-constructed similar case library.

[0049] This step requires the construction of a case library in advance. The specific construction steps are: first use OCR (optical character recognition technology) and text vectorization technology to scan and digitize a large number of legal documents into computer-processable texts, and then use large models (such as GPT-3, etc.) to compress the case details (causes, nature, circumstances, consequences) in the legal files into vector form to build a vectorized case library.

[0050] The case documents to be classified are compressed into vector representations using the ChatGLM model, which are used to calculate the similarity of cases in the database using vector similarity algorithms such as cosine distance. Using the crime name and the compressed vector of the text as the index, similar case documents, also known as similar cases, are retrieved by using the distance between the crime names of two cases and their vector representations.

[0051] The charges and similar cases obtained in step S2 and step S3 can be corrected and confirmed manually. For inappropriate results, the judge can manually specify charges, delete or add similar case documents, and search the similar case database for related similar cases again through the manually specified charges and confirmed similar case documents, and obtain the results again for the judge to review. Repeat this step until the judge confirms that the charges are appropriate for the retrieved similar cases.

[0052] S4. Based on the case documents to be classified, the crime classification criteria and similar case documents, prompts are constructed and submitted to a pre-trained large model for classification prediction to complete the automatic case classification process.

[0053] This step serializes the text of the current case document, the classification benchmark of the crime, and similar cases, builds prompts through templates, and submits them to the big model for classification prediction in the form of zero-shot classification to provide the judge with corresponding discretionary suggestions. Among them, the big model uses the GPT-3 series model, which performs classification prediction based on case documents, the classification benchmark of the crime, and similar cases, and has strong zero-shot learning and reasoning capabilities.

[0054] Among them, Figure 2 The template prompts shown in the figure can effectively improve the large model's ability to understand specific tasks and limit the model's output format. This template is composed of five parts: task description, case documents, crime and classification benchmark table, similar case documents, and output format restrictions.

[0055] In addition, the cases in this embodiment are automatically added to the case library after the classification is completed, and can be used for the next classification.

[0056] Example 2

[0057] This embodiment provides a case automatic classification system based on large model context learning, the system comprising:

[0058] Acquisition module: Acquisition of case documents to be classified;

[0059] Classification module: inputting the case documents to be classified into a pre-built crime classification model to obtain a crime classification benchmark;

[0060] Retrieval module: used to compress the case documents to be classified into vector representation, use it together with the crime classification standard as an index, and retrieve similar case documents from a pre-built similar case library;

[0061] Classification module: used to construct prompts based on the case documents to be classified, the crime classification criteria and similar case documents, and submit them to the pre-trained large model for classification prediction to complete the automatic case classification process.

[0062] Through the various modules of the above system, this embodiment adopts the browser / server (B / S) mode to build an automatic classification system, such as Figure 3 As shown in the figure, the system mainly consists of two parts, namely the local server and the external large model. Among them, the local server includes: a vectorized case library, a crime prediction model and a text vectorization model. The vectorized case library stores historical case texts, uses the crime and the compressed text vector as the index, and retrieves similar cases through the distance between the crime of two cases and their vector representations; the crime prediction model receives the document materials of the current case, takes the case plot paragraph as the main judgment basis, gives the crime that the current case meets, and assists in retrieving similar cases; the text vectorization model receives the case text, compresses it into a vector representation, and is used to calculate the similarity of cases in the database. The external large model takes the case documents given by the judge, the similar cases retrieved from the vectorized case library, the crime given by the crime prediction model, and the corresponding classification table as input, and gives classification predictions based on zero-shot context learning.

[0063] The specific steps for case classification using this system are:

[0064] (1) The judge accesses and logs into the front-end operation interface through a browser and uploads the case documents to be classified.

[0065] The platform used for front-end interaction in this embodiment is based on a web platform application and has good device compatibility.

[0066] (2) The front end sends a request to the back end to predict the crime and search for similar cases.

[0067] With the help of HTML and TCP protocol, the system interaction interface is constructed, and reliable data reception and push can be achieved through communication with the front-end and back-end.

[0068] (3) The backend returns the predicted crime and similar case texts to the frontend, and the judge confirms them. For inappropriate returned results, the judge manually specifies the crime, deletes or adds similar case documents through web page interaction. Based on the manually specified crime and confirmed similar case documents, the backend searches for related similar cases again and returns them to the frontend for the judge to review.

[0069] (4) Repeat steps (2)-(3) until the judge confirms that the crime is consistent with the searched case.

[0070] Steps (3)-(4) involve the judge manually correcting the model's prediction results, thereby improving the model's accuracy and flexibility in handling complex (or rare) cases through human intervention.

[0071] (5) The back-end receives and confirms the current case document, the classification criteria for the crime, and the text of similar cases, which are serialized and prompted through template construction. The large model performs classification prediction and returns the results.

[0072] (6) The front end displays the classification results, vectorizes the classification results and stores them in the case library, and returns to step (1).

[0073] After the current case is classified, it is automatically added to the case library. Through semi-supervised learning, the system cold start effect is improved and the data volume requirements are reduced.

[0074] The rest is the same as in Example 1.

[0075] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program code.

[0076] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes. The schemes in the embodiments of the present invention may be implemented in various computer languages, for example, object-oriented programming language Java and literal scripting language JavaScript, etc.

[0077] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1A device that provides the functions specified in a block or multiple blocks.

[0078] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0079] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0080] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0081] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A case automatic classification method based on large model context learning, It is characterized in that The following steps are involved: Obtain case documents to be classified; Inputting the case documents to be classified into a pre-built crime classification model to obtain a crime classification benchmark; Compress the case document to be classified into a vector representation, use it together with the crime classification standard as an index, and retrieve similar case documents from a pre-built similar case library; Based on the case documents to be classified, the crime classification standards and similar case document construction prompts, and handed over to the pre-trained large model for classification prediction, the automatic case classification process is completed.

2. According to claim 1, a case automatic classification method based on large model context learning, It is characterized in that The specific steps of constructing the crime classification model are as follows: Obtain crime prediction dataset; The text classification model is called and fine-tuned on the crime prediction dataset to obtain a crime classification model.

3. According to claim 2, a case automatic classification method based on large model context learning, It is characterized in that The text classification model is a LongFormer model.

4. According to claim 1, a case automatic classification method based on large model context learning, It is characterized in that The specific steps for constructing the case library are as follows: Access to a large number of existing case documents; Scanning and digitally extracting the case documents using an optical character recognition method and a text vectorization method to obtain the case documents represented by character codes; The case details in the case document represented by character code are compressed into a vector form through a large model to form a vectorized case library.

5. According to claim 1, a case automatic classification method based on large model context learning, It is characterized in that The ChatGLM model is used to compress the case documents to be classified into vector representation.

6. According to claim 1, a case automatic classification method based on large model context learning, It is characterized in that Also includes: Manual corrections are made to the crime classification standards and similar case documents.

7. According to claim 6, a case automatic classification method based on large model context learning, It is characterized in that The specific steps of the deviation correction include: Manually screen or add the crime classification criteria and similar case documents; Based on the case documents to be classified, crime classification benchmarks and similar case documents are obtained again; The above process is repeated until the final crime classification criteria and approximate case documents are obtained.

8. According to claim 1, a case automatic classification method based on large model context learning, It is characterized in that The large model is the GPT-3 series model.

9. An automatic case classification system based on large model context learning, It is characterized in that include: Acquisition module: Acquisition of case documents to be classified; Classification module: inputting the case documents to be classified into a pre-built crime classification model to obtain a crime classification benchmark; Retrieval module: used to compress the case documents to be classified into vector representation, use it together with the crime classification standard as an index, and retrieve similar case documents from a pre-built similar case library; Classification module: used to construct prompts based on the case documents to be classified, the crime classification criteria and similar case documents, and submit them to the pre-trained large model for classification prediction to complete the automatic case classification process.

10. A computer-readable storage medium, It is characterized in that It includes one or more programs for execution by one or more processors of an electronic device, and the one or more programs include instructions for executing the automatic case classification method based on large model context learning as described in any one of claims 1-8.