Multi-source heterogeneous knowledge injection type prompt learning method for legal criminal name prediction

Through the multi-source heterogeneous knowledge injection prompt learning method, combined with RoBERTa and dialogue-based large language model, the factual elements are extracted, and the THUOCL_Law knowledge base is used to match legal provisions, which solves the problem of insufficient input length and domain knowledge in the prediction of legal crimes, and achieves efficient and accurate prediction of legal crimes.

CN120471171APending Publication Date: 2025-08-12NORTHEAST FORESTRY UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510576602.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The prior art has problems such as input length limitations, insufficient domain knowledge and insufficient utilization of external structured knowledge in the prediction of legal crimes, which leads to the inability of the model to fully model long texts and accurately understand legal texts.

Method used

Multi-source heterogeneous knowledge injection-based prompt learning method is used to construct a joint semantic space between case description and legal provisions, use the RoBERTa model for comparison training, combine the dialogue-based large language model to extract factual elements, and use the THUOCL_Law knowledge base to match relevant knowledge fragments, inject legal language model for reasoning, and finally determine the crime category through Jaccard similarity.

Benefits of technology

It significantly improves the accuracy of legal offence prediction, solves input length limitations, enhances domain adaptability and data efficiency, improves model interpretability, and supports multitasking learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471171A_ABST
    Figure CN120471171A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source heterogeneous knowledge injection type prompt learning method for legal criminal name prediction. The method comprises the following steps: step 1, obtaining and coding fact elements; 2, knowledge matching and prompt injection; 3, reasoning the legal language model; and 4, legal crime name category mapping. The method has the advantages that prediction accuracy is improved; the problem of input length limitation is solved; the field adaptability is enhanced; the data efficiency is improved; the interpretability of the model is enhanced; multi-task learning is supported; and intelligent development of laws is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and legal technology, and in particular to a multi-source heterogeneous knowledge injection prompt learning method for legal crime prediction. Background Art

[0002] Crime prediction is a key task in the field of legal artificial intelligence. It aims to automatically predict the corresponding crime label by analyzing case descriptions. This task relies on leveraging technologies such as machine learning, deep learning, and natural language processing to extract key information from unstructured legal texts and map it to specific crime categories. Crime prediction not only helps legal professionals handle cases more efficiently and reduce human error, but also improves the consistency and fairness of judgments. It also promotes public legal awareness by disseminating legal knowledge.

[0003] Traditional methods for predicting legal charges typically treat it as a text classification problem, employing general text classification models. For example, methods based on convolutional neural networks (CNNs) achieve classification by capturing local features; methods based on bidirectional long short-term memory (BiLSTM) networks leverage sequence modeling capabilities to process long texts; and methods based on gated recurrent units (GRUs) improve computational efficiency by simplifying the network structure. Furthermore, attention mechanisms have been introduced to enhance the model's ability to focus on key information. While these methods have achieved some success, their performance remains limited due to the unique nature of legal texts.

[0004] The particularity of legal texts is mainly reflected in the following aspects:

[0005] Domain expertise: Legal texts contain a large amount of professional terminology and specific expressions, and general text classification models find it difficult to accurately understand their semantics.

[0006] Importance of factual information: Legal charge prediction focuses more on the factual information in the case description (such as time, place, behavior, etc.) rather than the theme or emotion of the text.

[0007] Long text length: Case descriptions often contain a lot of detailed information, and their length far exceeds the input limit of common pre-trained language models (such as BERT) (usually 512 words).

[0008] To address these issues, researchers have proposed various improvements. For example, language models pre-trained in the legal field (such as Lawformer) enhance understanding of legal texts by pre-training on large-scale legal corpora. However, these models are still limited by input length and cannot fully model long texts. Furthermore, some studies have introduced hierarchical attention mechanisms to capture key information in case descriptions through hierarchical modeling, but this still does not fully utilize external structured knowledge.

[0009] In recent years, the development of large-scale language models (such as ChatGPT and ChatGLM) has provided new insights into legal guilt prediction through their capabilities in zero-shot learning and contextual reasoning. However, these models still suffer from the "hallucination" problem in legal applications, where the generated content may not reflect the facts. Therefore, combining legal expertise with the reasoning capabilities of large-scale language models has become a hot research topic.

[0010] In summary, existing technologies face the following major challenges in the task of legal charge prediction:

[0011] Input length limitation: General pre-trained models cannot fully model long texts, resulting in information loss.

[0012] Insufficient domain knowledge: General models lack a deep understanding of legal terminology and factual information.

[0013] Insufficient utilization of external knowledge: Existing methods do not fully utilize external structured legal knowledge (such as legal provisions, case libraries, etc.). Summary of the Invention

[0014] In order to solve the above problems, especially to address the deficiencies in the existing technology, the present invention provides a multi-source heterogeneous knowledge injection prompt learning method for legal crime prediction that can solve the above problems.

[0015] To achieve the above objectives, the present invention adopts the following technical means:

[0016] The multi-source heterogeneous knowledge injection prompt learning method for legal crime prediction includes the following steps:

[0017] Step 1: Fact element acquisition and encoding

[0018] 1.1 Constructing a joint semantic space of case description and legal provisions:

[0019] Use the RoBERTa model for comparative training and optimize the model through the comparative loss function to make the semantic vectors of relevant case descriptions and legal provisions closer;

[0020] The CAIL-2018 dataset is used to automatically construct positive and negative sample pairs, where the positive sample is a case description and its related legal provisions, and the negative sample is a case description and unrelated legal provisions;

[0021] 1.2 Search for the N legal provisions most relevant to the case description:

[0022] Encode the case description and candidate legal provisions into semantic vectors, calculate the correlation through vector inner product, and select the N legal provisions with the highest correlation;

[0023] 1.3 Extracting Factual Elements Using a Conversational Large Language Model:

[0024] Build question templates and use a conversational large language model to extract key factual elements from case descriptions;

[0025] 1.4 Encoding fact elements into semantic vectors:

[0026] The extracted fact elements are concatenated into sequences, fed into the BiGRU encoder, and semantic vectors are generated and injected into the hint learning framework;

[0027] Step 2: Knowledge matching and prompt injection

[0028] 2.1 Matching case description with legal knowledge base:

[0029] Use the THUOCL_Law knowledge base to match the content in the case description with the legal terms and provisions in the knowledge base;

[0030] 2.2 Extract relevant knowledge fragments:

[0031] Extract the most relevant knowledge fragments to the case description from the knowledge base as prompt information;

[0032] 2.3 Injecting knowledge fragments into the model:

[0033] Splice the matching knowledge fragments into the original input as prompt information to enhance the model's reasoning ability; Step 3: Legal language model reasoning

[0034] 3.1 Constructing legal language model input:

[0035] The input includes soft prompt tokens, manually constructed template text, masked vocabulary, case descriptions, and knowledge snippets;

[0036] 3.2 Perform masked vocabulary prediction:

[0037] Use the legal language model to jointly reason over the input content and predict the words at the masked positions;

[0038] 3.3 Generate inference results:

[0039] Generate legal guilt reasoning results related to the case description through language modeling tasks;

[0040] Step 4: Legal Crime Category Mapping

[0041] 4.1 Calculate the similarity between the predicted words and the crime label:

[0042] Use Jaccard similarity to calculate the similarity between the predicted words and the legal crime label text;

[0043] 4.2 Determine the final crime category:

[0044] The crime label with the highest similarity is selected as the final prediction result.

[0045] A further solution of the present invention is that in step 1.1, the construction of the joint semantic space is achieved through contrastive training, which specifically includes:

[0046] Use the RoBERTa model to encode case descriptions and legal provisions;

[0047] The model is optimized by contrasting the loss function, so that the semantic vectors of positive sample pairs are closer and the semantic vectors of negative sample pairs are farther apart.

[0048] A further solution of the present invention is that in step 1.3, the use of the conversational large language model includes:

[0049] Build question templates and query the conversational language model for key factual elements in the case description;

[0050] Use ChatGLM or ChatGPT as a conversational large language model to extract a list of factual elements.

[0051] A further solution of the present invention is that in step 1.4, the specific process of encoding the fact elements includes:

[0052] Concatenate the extracted fact elements into a sequence and input it into the BiGRU encoder;

[0053] The sequences are processed separately by forward GRU and backward GRU to generate semantic vectors and inject them into the hint learning framework.

[0054] A further solution of the present invention is that in step 2.1, the specific process of knowledge matching includes:

[0055] Use the THUOCL_Law knowledge base to match and extract legal terms and provisions related to the case description;

[0056] Inject the matching knowledge fragments into the model as prompt information.

[0057] A further solution of the present invention is that in step 2.2, the process of extracting knowledge fragments includes:

[0058] Through semantic matching algorithm, the most relevant knowledge fragments to the case description are extracted from the THUOCL_Law knowledge base;

[0059] The extracted knowledge fragments include legal terms, article interpretations and relevant cases.

[0060] A further solution of the present invention is that in step 2.3, the specific process of injecting the knowledge fragment includes:

[0061] Splice the matching knowledge fragments into the original input as prompt information;

[0062] Through the prompt learning framework, knowledge fragments are combined with case descriptions to enhance the model's reasoning ability.

[0063] A further solution of the present invention is that in step 3.2, the specific process of masked vocabulary prediction includes:

[0064] Jointly reason over the input using a legal language model;

[0065] Generate legal guilt reasoning results related to the case description through the masked word prediction task.

[0066] A further solution of the present invention is that in step 4.1, the calculation of the Jaccard similarity includes:

[0067] Calculate the intersection and union ratio of the predicted vocabulary and the legal crime label text;

[0068] The crime label with the highest similarity is selected as the final prediction result.

[0069] A further solution of the present invention is that in step 4.2, the specific process of determining the final crime category includes:

[0070] Traverse all crime labels and calculate the Jaccard similarity between them and the predicted words;

[0071] The crime label with the highest similarity is selected as the final prediction result.

[0072] Beneficial effects of the present invention:

[0073] 1. The present invention improves prediction accuracy.

[0074] By integrating multi-source heterogeneous knowledge (including case descriptions, external legal knowledge bases, and conversational large language models), the proposed method significantly enhances the model's ability to understand legal texts. Specifically:

[0075] Factual element extraction: The conversational large language model is used to extract key factual elements from case descriptions, enhancing the model's ability to capture case details.

[0076] Knowledge fragment injection: Utilize the legal knowledge base to match relevant knowledge fragments, providing additional domain knowledge support for the model.

[0077] Long text modeling: Using a legal language model that supports long text input (such as Lawformer) avoids information loss due to input length limitations.

[0078] 2. The present invention solves the problem of input length limitation.

[0079] Traditional pre-trained language models (such as BERT) have an input length limit of 512 words, while case descriptions often exceed 1,000 words, resulting in the loss of key information. The method of the present invention solves this problem by:

[0080] Using a legal language model that supports long text input (such as Lawformer) can fully model case descriptions and minimize information loss.

[0081] Through hierarchical modeling and attention mechanism, local key information in case descriptions is captured to further improve model performance.

[0082] 3. The present invention enhances field adaptability.

[0083] Legal texts contain a large number of specialized terms and specific expressions, making it difficult for general models to accurately understand their semantics. The proposed method enhances the model's domain adaptability through the following means:

[0084] Legal domain pre-training: Use language models pre-trained on large-scale legal corpora (such as Lawformer) to equip them with prior knowledge in the legal field.

[0085] External knowledge utilization: By matching knowledge fragments in legal knowledge bases (such as THUOCL_Law), additional domain knowledge support is provided to the model.

[0086] Factual element extraction: Using a conversational large language model to extract key factual elements from case descriptions enhances the model's ability to understand legal texts.

[0087] 4. The present invention improves data efficiency.

[0088] In data-scarce scenarios, the proposed method can still maintain high performance and is suitable for small sample learning. This advantage is mainly attributed to:

[0089] Multi-source knowledge integration: By integrating case descriptions, external knowledge bases, and conversational large language models, the model can learn more effective information from limited data.

[0090] Prompt learning framework: Prompt learning is used to transform the legal charge prediction task into a language modeling task, reducing the model's dependence on large-scale labeled data.

[0091] 5. The present invention enhances the interpretability of the model.

[0092] The method of the present invention enhances the interpretability of the model by:

[0093] Factual element extraction: Clearly extract key factual elements (such as time, place, behavior, etc.) in the case description to make the model's prediction process more transparent.

[0094] Knowledge fragment injection: By matching knowledge fragments in the legal knowledge base, additional explanation basis is provided for the model's prediction results.

[0095] Crime category mapping: This method clarifies the model's prediction logic by calculating the Jaccard similarity between the predicted words and the crime labels.

[0096] 6. The present invention supports multi-task learning.

[0097] The method proposed in this paper is not only applicable to the task of predicting legal charges but can also be extended to other legal AI tasks, such as recommending legal provisions and calculating case similarity. By adjusting the prompt learning framework and knowledge injection method, the model can adapt to different task requirements and has high versatility and scalability.

[0098] 7. This invention promotes the development of intelligent law.

[0099] The method of the present invention provides technical support for the development of legal intelligence by improving the accuracy and efficiency of legal charge prediction.

[0100] Assisting legal practitioners: Helping judges, lawyers and other legal practitioners handle cases efficiently, reduce human errors, and improve the consistency and fairness of judgments.

[0101] Support legal education: Promote public understanding and compliance with the law and enhance social legal awareness by disseminating legal knowledge.

[0102] Promote the application of legal AI: Provide a technical foundation for other tasks of legal artificial intelligence (such as legal question-answering, case retrieval, etc.), and promote the widespread application of legal AI. BRIEF DESCRIPTION OF THE DRAWINGS

[0103] Figure 1 is a flow chart of the method of the present invention;

[0104] Figure 2 It is a diagram showing the impact of the knowledge fragments and factual elements of the present invention on the method;

[0105] Figure 3 This is a graph showing the change in the F1 score of the model on the test set as the training data decreases. DETAILED DESCRIPTION

[0106] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0107] Example 1

[0108] like Figure 1 As shown in FIG, the multi-source heterogeneous knowledge injection prompt learning method for legal crime prediction includes the following steps:

[0109] Step 1: Fact element acquisition and encoding

[0110] 1.1 Constructing a joint semantic space of case description and legal provisions:

[0111] Use relevant legal provisions to extract key factual elements from the case description, as these provisions clearly define which factual elements are relevant to a specific legal offense. For example, the legal provision related to copyright infringement is "copying and distributing literary, audio-visual, or computer software works for profit without the permission of the copyright owner, without the consent of the copyright owner."

[0112] Publishing books that others have exclusive publishing rights to, reproducing audio-visual products produced by others without their permission, producing and exhibiting works of art that are falsely attributed to others, and the illegal gains are large or there are other serious circumstances.

[0113] From this, it can be inferred that factors such as whether the purpose is to make a profit, whether copyright permission has been obtained, and the amount of illegal income are all factual aspects worthy of attention.

[0114] In order to match relevant legal provisions with case descriptions, it is proposed to construct a joint semantic space containing case descriptions and legal provisions. The language model RoBERTa is trained comparatively to promote its learning of this joint semantic space. Compared with rule-based methods, language models can model deeper semantic connections between case descriptions and legal provisions, thereby obtaining better matching results. In addition, compared with traditional neural network-based semantic matching methods, comparative training emphasizes the relative relationship between positive and negative samples. Therefore, comparative training helps RoBERTa learn more unique discriminative features, which are crucial for determining the relevance between case descriptions and legal articles. The specific steps for RoBERTa to learn the joint semantic space are:

[0115] Step 1: Construct positive and negative pairs

[0116] CAIL-2018 is the largest legal crime prediction dataset, where each case description is annotated with the relevant legal provisions. We automatically construct contrasting positive and negative pairs from the entire training set. Given a training set containing pairs of case descriptions and relevant legal provisions, Use Algorithm 1 to automatically construct a set of positive and negative pairs Among them, X i Indicates that i in the training set th Case description of the (i-th) case, Indicates the relevant legal provisions.

[0117] Algorithm 1

[0118]

[0119] Step 2: Obtain a description of the case and the legal text

[0120] In step 1, the positive and negative pairs have been obtained Each pair in the set is either a related pair of case description and legal text, or an unrelated pair. In this step, RoBERTa is used to obtain the semantic vectors of the case description and legal text in each pair. This process is shown in formula (1).

[0121]

[0122] Here, P represents a matrix. The odd-numbered columns of P represent the semantic vectors of the case descriptions, and the even-numbered columns represent the semantic vectors of the legal provisions. Now that the semantic vectors for all positive and negative pairs have been obtained, we will use contrastive loss to train RoBERTa.

[0123] Step 3: Train RoBERTa with contrastive loss

[0124] During training, RoBERTa learns how to reduce the semantic distance between positive pairs of samples while increasing the semantic distance between negative pairs of samples. This goal is achieved through a contrastive loss function that quantifies the similarity between the semantic vectors in a pair of samples. th The loss calculation for the (i-th) case description is shown in formula (2).

[0125]

[0126] in, represents the semantic vector of the case description, and and Represents positive and negative alignment i th Semantic vector of the legal provision (Article i). In addition, Represents a vector and The cosine similarity between , τ is a temperature hyperparameter.

[0127] The trained RoBERTa model can encode case descriptions and legal provisions into a joint semantic space, where the representations of case descriptions and their corresponding legal provisions show closer semantic distances in this space.

[0128] 1.2 Search for the N legal provisions most relevant to the case description:

[0129] After obtaining the joint semantic space of case description and legal provisions through the above operations, this joint semantic space will be used to search for N legal provisions that are most relevant to the given case description. After that, we first use the trained RoBERTa model to encode them into the joint semantic space to obtain their respective semantic vectors. This process is shown in formula (3).

[0130]

[0131] in, Represents the semantic vector of case description X, Representing a collection in i th Semantic vector of the legal provision (Article i).

[0132] Subsequently, the correlation between the case description and each candidate legal provision is calculated by vector inner product, as shown in formula (4), so as to select the legal provision with the highest correlation.

[0133]

[0134] Use the N searched legal provisions and case description X to consult the dialogue language model to obtain the most valuable factual elements in X.

[0135] 1.3 Extracting factual elements using a large conversational language model:

[0136] Large conversational language models have a large number of parameters, strong contextual reasoning capabilities, and have learned a large amount of general world knowledge during pre-training. In addition, large conversational language models usually have strong zero-shot reasoning capabilities, so they can be used directly as a ready-made tool without additional fine-tuning. For these reasons, a large conversational language model is used to help obtain factual elements from X. The following question template is constructed and used to ask the large conversational language model a question to obtain a list of factual elements, which are expressed as F = [f1, f2, ..., f |F| ] These fact elements are then encoded into semantic vectors in 1.4. When the dataset is in Chinese, we use the GLM-130B Chinese Conversational Language Model developed by Zhipu Huazhang Technology Co., Ltd. When the dataset is in English, we directly call the ChatGPT API to implement this process.

[0137] 1.4 Encoding fact elements into semantic vectors:

[0138] Given the above obtained fact element list F=[f1,f2,…,f |F| ], we concatenate them and feed the combined sequence into the BiGRU encoder. A BiGRU consists of two GRU layers that process data in opposite directions: a forward GRU and a backward GRU. The forward GRU processes the data from f1 to f |F| The reverse GRU process starts from f |F| to f1. Each GRU updates its hidden state at each step of the sequence.

[0139] Will Let be the hidden state of the forward GRU at time step t, is the hidden state of the backward GRU at time step t. Their calculation formula is formula (5-6).

[0140]

[0141] The final semantic vector It is usually formed by connecting the last hidden state of the forward GRU and the first hidden state of the backward GRU, as shown in formula (7).

[0142]

[0143] The semantic vector obtained through this process captures information about all factual elements derived from the case description. This vector is then incorporated into the inference model to enhance its ability to predict legal charges.

[0144] Step 2: Knowledge matching and prompt injection

[0145] Compare the case description X with the given knowledge base Matching is performed. The matching knowledge fragments can serve as hints to enhance the reasoning model's ability to predict legal charges.

[0146] We use THUOCL_Law as our knowledge base. THUOCL_Law is a sub-database of the Tsinghua University Open Chinese Lexicon (THUOCL), a high-quality Chinese lexicon compiled and published by the Natural Language Processing and Social Humanities Computing Laboratory at Tsinghua University. All sub-databases have undergone multiple rounds of manual screening to ensure their accuracy.

[0147] A knowledge fragment is essentially a keyword. Keywords in case descriptions are crucial for predicting legal charges. Simply use regular expressions to match these keywords in case descriptions to obtain a keyword list K = [k1, k2, ..., k |K| ], also known as a list of knowledge fragments. The concatenation of the knowledge fragments in K will serve as a hint and, together with other components, as the input of the reasoning model.

[0148] Step 3: Legal Language Model Inference

[0149] A legal language model is used to infer the legal charge associated with a given case description X. Traditionally, predicting legal charges is considered a classification problem, where the model output is a probability distribution, with the index with the highest probability being the predicted label. However, this task is transformed into a language modeling (cloze test) task, forcing the model to predict masked words. The predictions for these masked positions are then mapped to the final class label.

[0150] In order to complete the language modeling task, hard prompt templates T1 and T2 are constructed as follows:

[0151] T1 = "He will be held criminally responsible for..."

[0152] T2 = "The keywords in the case description are as follows:"

[0153] These hard hints are used as part of the input to the inference model to guide the model to predict the mask words. In addition to the hard hints, two soft hints s1 and s2 are added to the input. The semantic vector obtained in step 1 will be merged with these soft hints to inject knowledge about factual elements into the inference model. In addition, the mask vocabulary M = [m1,m2,…,m |M| ] and case descriptions are also important components of the input. Finally, the concatenation of the knowledge fragments K obtained in step 2 is also part of the input.

[0154] The input of the reasoning model consists of case description X, hard prompt texts T1 and T2, soft prompt texts s1 and s2, and mask sequence M = [m1,m2,…,m |M| ] and the knowledge fragment K are connected in series, as shown in formula (8).

[0155]

[0156] Among them, t 1,i represents i in T1 th (i-th) word, t 2,i Represents i in T2 th vocabulary. In addition, m i Representative i th Mask, x i represents i in X th Vocabulary, k i represents i in K th vocabulary.

[0157] Next, X′ is input to the inference model, a pre-trained legal language model. The inference model consists of an embedding layer and an encoding layer. During the embedding layer phase, the embedding layer of the inference model embeds all words in X′ except for the soft prompts. Simultaneously, the soft prompts in X′ are also embedded using another trainable embedding matrix. This process is shown in Equation (9).

[0158]

[0159] in, is a trainable embedding matrix, soft idx is the index of the soft prompt vocabulary, d h is the embedding size of the model, so the embedding vector sequence E of all words in X′ (including soft prompt words) can be obtained by Eq.

[0160]

[0161] in, and Indicates i in T1, T2, M, X, and K th The embedding vectors of the vocabulary. In addition, and Embedding vectors representing the soft hint words s1 and s2.

[0162] In order to inject factual element information into the forward propagation calculation process of the inference model, the semantic vector obtained in step 1 is Add to Vector and This allows the fact element information to enrich the hint vector. This process is demonstrated in formulas (10) and (11).

[0163]

[0164] Then, use and Replaced the E and Obtained:

[0165]

[0166] Finally, we input E′ into the encoding layer of the inference model and obtain the hidden layer output of the model, which is the context representation of each word. This process is shown in Equation (12).

[0167]

[0168] The goal of the model is to predict the word at the mask position. Therefore, by Projecting into the vocabulary space, we can obtain the probability distribution of the predicted tokens. Subsequently, by selecting the index with the highest probability, the model can determine the vocabulary of the mask position. We will i th The pre-vocabulary representation of the mask position is

[0169] The next section describes the predicted vocabulary The process of mapping to crime category labels.

[0170] Step 4: Legal Crime Category Mapping

[0171] The results of the language modeling (cloze) task need to be mapped back to the classification task. To this end, a mapping is constructed from predicted words to legal charge categories. The Jaccard similarity between the predicted labels and the legal charge label text is calculated. For example, if the predicted label text has the highest Jaccard similarity to "manufacturing, trafficking, and disseminating obscene materials," the model predicts the legal charge for the current case description as "manufacturing, trafficking, and disseminating obscene materials."

[0172] Therefore, the final prediction of the model is shown in formula (13).

[0173]

[0174] in, is the final predicted label of the model, The predicted word representing the mask position, v y Represents the text category y.

[0175] Example 2

[0176] Theft Prediction

[0177] Case Description

[0178] "On May 1, 2023, Zhang stole the mobile phone from a shopping cart belonging to a customer named Li while she was not paying attention. Upon investigation, it was found that Zhang had previously been sentenced to one year in prison for theft."

[0179] Implementation steps

[0180] Fact element acquisition and encoding

[0181] Through joint semantic space learning, match the legal provisions related to the case description (such as Article 264 of the Criminal Law).

[0182] Use ChatGLM to extract factual elements: time (May 1, 2023), location (a shopping mall), perpetrator (Zhang), victim (Li), stolen items (mobile phone), and criminal record (previously sentenced for theft).

[0183] Factual elements are encoded into semantic vectors and injected into the hint learning framework.

[0184] Knowledge matching and prompt injection

[0185] Use the THUOCL_Law knowledge base to match legal terms in case descriptions (such as "theft" and "criminal record").

[0186] Extract relevant knowledge fragment: "The crime of theft refers to the act of secretly stealing public or private property for the purpose of illegal possession."

[0187] Legal Language Model Reasoning

[0188] The input includes soft prompt tokens, template text, mask vocabulary, case description and knowledge snippets.

[0189] The model predicts the masked word is "theft".

[0190] Legal crime category mapping

[0191] Calculate the Jaccard similarity between the predicted word "theft" and the crime label, and determine the final crime as "theft".

[0192] result

[0193] The model successfully predicted the crime as "theft," which is consistent with the case description.

[0194] Example 3

[0195] Prediction of intentional injury

[0196] Case Description

[0197] "On June 10, 2023, Wang and his neighbor Zhao had an argument over trivial matters. Wang stabbed Zhao with a knife, causing him a second-degree minor injury. Upon investigation, Wang had no criminal record."

[0198] Implementation steps

[0199] Fact element acquisition and encoding

[0200] Match relevant legal provisions (such as Article 234 of the Criminal Law).

[0201] ChatGLM was used to extract the factual elements: time (June 10, 2023), location (neighbor’s house), perpetrator (Wang), victim (Zhao), injury tool (knife), injury result (second-degree minor injury), and criminal record (none).

[0202] Factual elements are encoded into semantic vectors and injected into the hint learning framework.

[0203] Knowledge matching and prompt injection

[0204] Match legal terms (such as "intentional injury" and "minor injury").

[0205] Extract relevant knowledge fragment: "Intentional injury refers to the act of intentionally and illegally harming the physical health of others."

[0206] Legal Language Model Reasoning

[0207] The input includes soft prompt tokens, template text, mask vocabulary, case description and knowledge snippets.

[0208] The model predicts the masked word is "intentional injury".

[0209] Legal crime category mapping

[0210] Calculate the Jaccard similarity between the predicted word "intentional injury" and the crime label, and determine the final crime as "intentional injury".

[0211] result

[0212] The model successfully predicted the crime as "intentional injury," which is consistent with the case description.

[0213] Example 4

[0214] Contract Fraud Prediction

[0215] Case Description

[0216] "On July 15, 2023, Liu used a false contract to defraud a company of 500,000 yuan in payment for goods and then disappeared. Upon investigation, it was found that Liu had been sentenced to two years in prison for fraud."

[0217] Implementation steps

[0218] Fact element acquisition and encoding

[0219] Match relevant legal provisions (such as Article 224 of the Criminal Law).

[0220] Use ChatGLM to extract factual elements: time (July 15, 2023), perpetrator (Liu), victim (a company), amount of fraud (500,000 yuan), means of fraud (false contract), and criminal record (once sentenced for fraud).

[0221] Factual elements are encoded into semantic vectors and injected into the hint learning framework.

[0222] Knowledge matching and prompt injection

[0223] Match legal terms (such as "contract fraud" and "sham contract").

[0224] Extract relevant knowledge fragment: "Contract fraud refers to the act of defrauding the other party of property during the signing and performance of a contract for the purpose of illegal possession."

[0225] Legal Language Model Reasoning

[0226] The input includes soft prompt tokens, template text, mask vocabulary, case description and knowledge snippets.

[0227] The model predicts the masked word is "contract fraud".

[0228] Legal crime category mapping

[0229] The Jaccard similarity between the predicted word "contract fraud" and the crime label is calculated to determine the final crime as "contract fraud".

[0230] result

[0231] The model successfully predicted the crime as "contract fraud," which is consistent with the case description.

[0232] Ablation experiments

[0233] The contribution of the present invention is analyzed through ablation experiments, and the experimental results are shown in Table 1. The first row of the table represents the performance of the original method on the CAIL-2018 dataset.

[0234] Table 1 Ablation experiment results

[0235] P R F1 Ours 0.85 0.83 0.84 -knowledge snippets 0.82 0.81 0.82(-0.02) -factual elements 0.81 0.81 0.81(-0.03) -knowledge snippets and factual elements 0.79 0.77 0.78(-0.06) -Contrastive training 0.83 0.82 0.82(-0.02)

[0236] First, we removed the knowledge fragment module. The experimental results are shown in the second row of the table. We observe that the model's macro F1 score drops by 0.02, indicating that focusing on factual elements in case descriptions does enhance the predictive validity of legal charges. To thoroughly investigate the contribution of knowledge fragments to our method, we reduced the maximum number of matching knowledge fragments from 12 to 0. Figure 2 Subfigure (a) in [1] shows that the F1 score of our method is approximately proportional to the number of knowledge fragments.

[0237] Next, we retained the knowledge fragment module and removed the factual element acquisition and encoding module. As can be seen in the third row of the table, this resulted in a 0.03 drop in the model's macro F1 score. This demonstrates that factual elements in case descriptions also contribute to the legal charge prediction task. We evaluated the impact of different conversational large language models on our method, as these are crucial for acquiring case elements. Figure 2 Subgraph (b) in Figure 3 shows that using ChatGLM enables our model to achieve the highest F1 score. This is because ChatGLM is a large conversational language model developed specifically for Chinese, making it more proficient in handling Chinese tasks.

[0238] Finally, we removed both modules, and the experimental results are shown in the fourth row of the table. We can observe that the model's performance dropped significantly by 0.06. This significant drop demonstrates that integrating multi-source heterogeneous legal knowledge is crucial for the task of legal charge prediction.

[0239] We also verified the effectiveness of the contrastive training described in Step 1 in accurately retrieving relevant legal provisions. We eliminated the learning of the joint semantic space and directly used RoBERTa to encode case descriptions and legal provisions. We still determined the relevance between a given case description and each legal provision by calculating the inner product between their encoding vectors. Experimental results show that this operation reduced the model's final performance by 0.02, as shown in the last row of Table 1. This demonstrates the importance of training the joint semantic space of case descriptions and legal provisions through contrastive learning.

[0240] The impact of training data volume

[0241] Test the impact of using training data of different scales on the model and analyze the experimental results. Figure 3The curve of the change of the model's F1 score on the test set as the training data decreases is shown. It can be seen that as the amount of training data decreases, the macro F1 scores of all methods are decreasing (except ChatGLM, because it does not require training data), while the decrease in the macro F1 score of the method of the present invention is gradual. This shows that the method of the present invention has a stronger advantage in scenarios where data is scarce. In addition, even if the amount of training data is reduced to only 10% of the original size, the method of the present invention still obtains a macro F1 score that exceeds ChatGLM. This can be attributed to two factors. On the one hand, the CAIL-2018 dataset itself is quite large, which means that even 10% of the original training data volume consists of 80,000 training samples. On the other hand, this shows that conversational large language models still have no clear advantages in specific fields such as law and medicine.

[0242] The present invention is provided as an example, not as a limitation of the embodiments. Those skilled in the art will appreciate that other variations or modifications may be made based on the above description. It is not necessary and impossible to enumerate all the embodiments here, and obvious variations or modifications derived therefrom remain within the scope of protection of the present invention.

Claims

1. A multi-source heterogeneous knowledge injection prompt learning method for legal crime prediction, characterized by: The following steps are involved: Step 1: Fact element acquisition and encoding 1.1 Constructing a joint semantic space of case description and legal provisions: Use the RoBERTa model for comparative training and optimize the model through the comparative loss function to make the semantic vectors of relevant case descriptions and legal provisions closer; The CAIL-2018 dataset is used to automatically construct positive and negative sample pairs, where the positive sample is a case description and its related legal provisions, and the negative sample is a case description and unrelated legal provisions; 1.2 Search for the N legal provisions most relevant to the case description: Encode the case description and candidate legal provisions into semantic vectors, calculate the correlation through vector inner product, and select the N legal provisions with the highest correlation; 1.3 Extracting Factual Elements Using a Conversational Large Language Model: Build question templates and use a conversational large language model to extract key factual elements from case descriptions; 1.4 Encoding fact elements into semantic vectors: The extracted fact elements are concatenated into sequences, fed into the BiGRU encoder, and semantic vectors are generated and injected into the hint learning framework; Step 2: Knowledge matching and prompt injection 2.1 Matching case description with legal knowledge base: Use the THUOCL_Law knowledge base to match the content in the case description with the legal terms and provisions in the knowledge base; 2.2 Extract relevant knowledge fragments: Extract the most relevant knowledge fragments to the case description from the knowledge base as prompt information; 2.3 Injecting knowledge fragments into the model: Splice the matching knowledge fragments into the original input as prompt information to enhance the model's reasoning ability; Step 3: Legal Language Model Inference 3.1 Constructing legal language model input: The input includes soft prompt tokens, manually constructed template text, masked vocabulary, case descriptions, and knowledge snippets; 3.2 Perform masked vocabulary prediction: Use the legal language model to jointly reason over the input content and predict the words at the masked positions; 3.3 Generate inference results: Generate legal guilt reasoning results related to the case description through language modeling tasks; Step 4: Legal Crime Category Mapping 4.1 Calculate the similarity between the predicted words and the crime label: Use Jaccard similarity to calculate the similarity between the predicted words and the legal crime label text; 4.2 Determine the final crime category: The crime label with the highest similarity is selected as the final prediction result.

2. The multi-source heterogeneous knowledge injection prompt learning method for legal crime prediction according to claim 1 is characterized in that: In step 1.1, the construction of the joint semantic space is achieved through contrastive training, which specifically includes: Use the RoBERTa model to encode case descriptions and legal provisions; The model is optimized by contrasting the loss function, so that the semantic vectors of positive sample pairs are closer and the semantic vectors of negative sample pairs are farther apart.

3. The multi-source heterogeneous knowledge injection prompt learning method for legal crime prediction according to claim 1 is characterized in that: In step 1.3, the use of the conversational large language model includes: Build question templates and query the conversational language model for key factual elements in the case description; Use ChatGLM or ChatGPT as a conversational large language model to extract a list of factual elements.

4. The multi-source heterogeneous knowledge injection prompt learning method for legal crime prediction according to claim 1 is characterized in that: In step 1.4, the specific process of encoding fact elements includes: Concatenate the extracted fact elements into a sequence and input it into the BiGRU encoder; The sequences are processed separately by forward GRU and backward GRU to generate semantic vectors and inject them into the hint learning framework.

5. The multi-source heterogeneous knowledge injection prompt learning method for legal crime prediction according to claim 1 is characterized in that: In step 2.1, the specific process of knowledge matching includes: Use the THUOCL_Law knowledge base to match and extract legal terms and provisions related to the case description; Inject the matching knowledge fragments into the model as prompt information.

6. The multi-source heterogeneous knowledge injection prompt learning method for legal crime prediction according to claim 1 is characterized in that: In step 2.2, the process of extracting knowledge fragments includes: Through semantic matching algorithm, the most relevant knowledge fragments to the case description are extracted from the THUOCL_Law knowledge base; The extracted knowledge fragments include legal terms, article interpretations and relevant cases.

7. The multi-source heterogeneous knowledge injection prompt learning method for legal crime prediction according to claim 1 is characterized in that: In step 2.3, the specific process of knowledge fragment injection includes: Splice the matching knowledge fragments into the original input as prompt information; Through the prompt learning framework, knowledge fragments are combined with case descriptions to enhance the model's reasoning ability.

8. The multi-source heterogeneous knowledge injection prompt learning method for legal crime prediction according to claim 1 is characterized in that: In step 3.2, the specific process of mask vocabulary prediction includes: Jointly reason over the input using a legal language model; Generate legal guilt reasoning results related to the case description through the masked word prediction task.

9. The multi-source heterogeneous knowledge injection prompt learning method for legal crime prediction according to claim 1 is characterized in that: In step 4.1, the calculation of Jaccard similarity includes: Calculate the intersection and union ratio of the predicted vocabulary and the legal crime label text; The crime label with the highest similarity is selected as the final prediction result.

10. The multi-source heterogeneous knowledge injection prompt learning method for legal crime prediction according to claim 1 is characterized in that: In step 4.2, the specific process of determining the final crime category includes: Traverse all crime labels and calculate the Jaccard similarity between them and the predicted words; The crime label with the highest similarity is selected as the final prediction result.

Citation Information

Cited By

  • Crime determination abnormity early warning method based on legal knowledge framework and star graph neural network

    CN121458493A

  • Criminal conviction anomaly early warning method based on legal knowledge framework and star graph neural network

    CN121458493B