Multi-element fusion class case text matching method

By fusing information about facts and degree elements in case text matching and using deep learning models for encoding and weighting fusion, the problem of neglecting the severity of the case in the prior art is solved, and the accuracy of the match is improved.

CN120234645APending Publication Date: 2025-07-01BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510213700.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing method of matching texts in a class only depends on the structured characteristics of the case text, ignoring the differences in the severity of the case, resulting in a low matching accuracy rate when facing cases with similar case facts but different circumstances severity.

Method used

A multi-factor fusion text matching method is used to extract facts and degree elements through BERT-BiLSTM-CRF, and a graph neural network is used to encode fact elements, adjust the weight of degree elements in combination with attention mechanism, and perform semantic coding weighted fusion. Finally, the matching similarity is calculated using cosine similarity.

Benefits of technology

The case text's plot severity information is effectively considered, the accuracy of case matching is improved, and the lack of relying solely on semantic information is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234645A_ABST
    Figure CN120234645A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-element fusion class case text matching method, and belongs to the field of artificial intelligence. According to the method, firstly, a BERT-BiLSTM-CRF is used for extracting fact elements composed of entities, relations and the like and degree elements composed of quantifiers and degree adverbs in a target sample; secondly, generating a representation vector by using GNN coding fact elements; secondly, adjusting the weight of a degree element in a target sample by particularly utilizing an attention mechanism according to a relationship category in the fact element, and generating a semantic code of the target sample by combining a BERT model; carrying out weighted fusion on the representation vector of the fact element and the semantic code of the sample to obtain a final representation vector; and finally, vector similarity between the target sample and other samples is calculated by using cosine similarity, and a high-similarity sample is selected as a class case sample. Aiming at the problem that the matching accuracy of class case texts is influenced by mainly utilizing structural features of the case texts in an existing method, the case plot difference is distinguished by fusing case facts and degree element fine grit, and the matching accuracy is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for matching case texts with multi-factor integration, and belongs to the field of artificial intelligence. Background Art

[0002] With the rapid development of artificial intelligence technology, text matching has become an important research direction in the field of natural language processing. Text matching technology is mainly applied to various tasks such as information retrieval, text classification, and recommendation systems, and plays a crucial role in industries such as law, medicine, and finance. In the legal field, similar case text matching can help legal workers such as judges and lawyers quickly find similar precedents for the current case in a large number of cases, providing reference for judicial decisions. However, the method for matching case texts still faces many challenges in practical applications.

[0003] Existing methods for matching case texts usually focus on analyzing factual elements such as entities and relationships in the text from a structural perspective, using the semantic information of samples to calculate the similarity between different texts, and outputting similar cases of the target case text. However, these methods usually ignore the fine-grained differences in case circumstances, that is, the differences in the severity and complexity of case circumstances. Although two cases may have a similar factual background on the surface, involving the same legal provisions or judgment principles, the severity of the case circumstances may be significantly different. For example, both cases involve theft, but one is sentenced to a light sentence due to minor circumstances and first-time offense, while the other is sentenced to a heavy sentence due to serious circumstances and repeated offenses. Similarly, the amount of theft is also an important part reflecting the severity of the case circumstances. For example, both cases involve theft, but the theft amount in one case is tens of thousands of yuan, while the theft amount in the other case is as high as tens of millions. Although both are theft acts, due to the difference in amount, the severity of the case circumstances is obviously different. This kind of difference in the severity of circumstances often determines the legal consequences of the case and is also an important factor that must be considered in judicial judgment.

[0004] Therefore, the existing method for matching similar case texts only relies on factual elements for similarity matching, and easily ignores the information contained in the severity of case circumstances. Although factual elements can reflect part of the information of the case, for the fine-grained matching of complex cases, especially in the context of similar criminal processes, how to accurately judge the severity of case circumstances is the key to determining the accuracy of the matching.

[0005] In summary, the existing method for matching similar case texts only focuses on the factual elements of the case text and ignores the information on the severity of case circumstances reflected in the case text. Therefore, on the basis of the existing method, further explore and utilize the information on the degree of circumstances in the case text, so that the model not only focuses on the factual element information of the case text, but also focuses on the information on the severity of case circumstances, and improves the accuracy of case matching. Summary of the Invention

[0006] The object of the present invention is to address the problem that the existing similar case text matching only relies on the structured features of the case text and ignores the differences in the severity of the case circumstances, resulting in a low matching accuracy rate of the model when facing cases with similar case facts but different severities. A method for matching similar case texts with multi-factor fusion is proposed.

[0007] The design principle of the present invention is as follows: First, use BERT-BiLSTM-CRF to extract the factual elements composed of entities, relationships, etc., and the degree elements composed of quantifiers and adverbs of degree in the target sample; secondly, use the graph neural network GNN to encode the factual elements to generate representation vectors; then, according to the relationship categories in the factual elements, use the attention mechanism to adjust the weights of the degree elements in the target sample, and then combine with the BERT model to generate the semantic encoding of the target sample; after that, weight and fuse the representation vectors of the factual elements and the semantic encoding of the sample to obtain the final representation vector; finally, use the cosine similarity to calculate the vector similarity between the target sample and other samples, and select the samples with high similarity as the similar case samples. The entire framework can finally obtain the case sample most similar to the input target sample.

[0008] The technical solution of the present invention is realized through the following steps:

[0009] Step 1, train the BERT-BiLSTM-CRF model to extract the factual elements composed of entities and relationships, etc., and the degree elements composed of quantifiers and adverbs of degree in the sample

[0010] Step 1.1, annotate each sample to mark the factual elements (such as entities and relationships) and degree elements (such as quantifiers and adverbs of degree) therein.

[0011] Step 1.2, use the pre-trained BERT model to perform word vector encoding on the sample, input the word vectors of BERT into the BiLSTM network, output the serialized representation, and pass the output of BiLSTM to the CRF layer. The CRF layer is used for structured annotation to accurately identify the factual elements and degree elements.

[0012] Step 1.3, use the cross-entropy loss function for factual element extraction and the binary cross-entropy loss function for degree element extraction, and perform weighted summation on the cross-loss function and the binary cross-loss function to obtain the joint loss function.

[0013] Step 1.4, compare the output of the CRF layer with the actual annotation results, and use the joint loss function to train the BERT-BiLSTM-CRF model.

[0014] Step 1.5: Use the trained BERT-BiLSTM-CRF model to extract the factual elements and relational elements from the target samples.

[0015] Step 2: Count the relationships in the dataset and classify each relationship into a crime category.

[0016] Step 2.1: Count all relationships and classify them into the following nine major crime categories according to the nature of the crime: violent crime, property crime, drug crime, economic crime, traffic crime, sexual crime, cyber crime, environmental crime, and other crimes.

[0017] Step 3: Use a graph neural network to encode the factual elements extracted in Step 1 to obtain the encoded vectors of the factual elements.

[0018] Step 4: According to the crime category corresponding to the relationship in the extracted factual elements, use the attention mechanism to adjust the weights of the degree elements, and combine with the BERT model to generate the semantic encoding of the target sample.

[0019] Step 4.1: Use the relationships in the factual elements extracted in Step 1 and the crime categories divided in Step 2 to determine the crime category to which the target sample belongs.

[0020] Step 4.2: Define different weight adjustment factors for the quantifiers and degree adverbs in each crime category.

[0021] Step 4.3: Use the attention mechanism and the weight adjustment factors to dynamically adjust the weights of the degree elements in the sample, and at the same time use the BERT model to generate the semantic encoding of the target sample.

[0022] Step 5: Weightedly fuse the encoded vectors of the factual elements of the target sample and the semantic encoding of the target sample to obtain the final representation vector.

[0023] Step 5.1: Weightedly fuse the encoded vectors of the factual elements of the target sample and the semantic encoding of the target sample to obtain the final representation vector.

[0024] Step 5.2: For other samples that are case-matched with the target sample, perform the same operations as the target sample to obtain their respective final representation vectors.

[0025] Step 6: Use the cosine similarity to calculate the similarity between the final representation vectors of the target sample and other samples, and select the samples with high similarity as the case samples.

[0026] Step 6.1: Use the cosine similarity to calculate the similarity between the final representation vectors of the target sample and other samples. Sort the similarities of all samples, select the sample with the highest similarity as the case of the target sample, and output the sample with the highest similarity to the target sample.

[0027] Beneficial effects

[0028] Compared with traditional case text matching methods, the present invention fully considers the differences in the severity of case texts. According to the relationships between entities in the samples, it changes the weights corresponding to the degree elements in the sample semantic encoding process, and also performs weighted fusion on the encoding vectors of fact elements and the sample semantic encoding to obtain a better final vector. This avoids relying solely on the semantic information and degree element information of case texts, makes full use of the severity information of cases in the samples, and improves the accuracy of matching. Description of the drawings

[0029] Figure 1 It is a schematic diagram of the case text matching method for multi-element fusion. Detailed implementation manners

[0030] To better illustrate the purpose and advantages of the present invention, the following further elaborates on the implementation manners of the method of the present invention with reference to examples.

[0031] The specific process is as follows:

[0032] The experimental data source is the public dataset CAIL2019 - SCM. The length of a single text in the CAIL2019 - SCM dataset can reach 1062 words, with an average length of 679 words. This task aims to judge the two more similar texts in the text triples. The construction method of the validation set and the test set is to split the original triple pairs into two binary pair pairs, and after testing, the prediction results of the two binary pair pairs are combined. The detailed statistical information of the split dataset is shown in Table 1.

[0033] Table 1. Detailed information of the training set, validation set and test set

[0034]

[0035] The judgment principle for the text matching results is as follows: If the similarity value of the text {a, b} output by the model is greater than the similarity value of {a, c} output, it means that the text {a, b} is more similar; otherwise, the model indicates that the text {a, c} is more similar. The accuracy rate is used to represent the ratio of the number of correctly predicted triples to the total number of triples in the test set. The accuracy formula is shown in (1)

[0036]

[0037] Where TP represents the number of triples correctly predicted by the model, FP represents the number of triples wrongly predicted by the model, TP + FP is equal to the number of triples, and Precision is the accuracy rate.

[0038] This experiment was conducted on a computer with the following specific configuration: Intel i7-9750H processor, CPU main frequency of 2.60 GHz, dual GTX 1080Ti GPUs, 24GB of memory, 24GB of video memory, Windows 11 64-bit operating system, Python 3.6.0 software environment, pre-trained language library Transformers 4.2.1, and open-source Python machine learning library PyTorch 1.7.

[0039] Step 1: Train the BERT-BiLSTM-CRF model to extract fact elements such as entities and relationships, and degree elements such as quantifiers and adverbs of degree from the samples.

[0040] Step 1.1: Annotate each sample in the dataset CAIL2019-SCM to mark the fact elements (such as entities and relationships) and degree elements (such as quantifiers and adverbs of degree) in it.

[0041] Step 1.2: Use the pre-trained BERT model to perform word vector encoding on the samples, input the word vectors of BERT into the BiLSTM network, output serialized representations, and pass the output of BiLSTM to the CRF layer, which is used for structured annotation to identify fact elements and degree elements.

[0042] Step 1.3: For fact element extraction, use the cross-entropy loss function L fact , and for degree element extraction, use the binary cross-entropy loss function L degree . Perform weighted summation on the cross-loss function and the binary cross-loss function to obtain the joint loss function, as shown in formula (2).

[0043] L joint = 0.5L fact + 0.5L degree (2)

[0044] where L joint is the joint loss function, L fact is the cross-entropy loss function, and L degree is the binary cross-entropy loss function.

[0045] Step 1.4: Compare the output of the CRF layer with the actual annotation results, and use the joint loss function to train the BERT-BiLSTM-CRF model.

[0046] Step 1.5: Use the trained BERT-BiLSTM-CRF model to extract fact elements and relationship elements from the target samples.

[0047] Step 2: Count the relationships in the dataset and classify each relationship into a crime category.

[0048] Step 2.1: Count all the relationships in the dataset and classify them into the following nine major crime categories according to the nature of the crime: violent crime, property crime, drug crime, economic crime, traffic crime, sexual crime, cyber crime, environmental crime, and other crimes.

[0049] Step 3: Use the graph neural network GNN to encode the fact elements extracted in Step 1 to obtain the encoded vectors of the fact elements.

[0050] Step 4: According to the crime category corresponding to the relationship in the extracted fact elements, use the attention mechanism to adjust the weights of the degree elements, and combine with the BERT model to generate the semantic encoding of the target sample.

[0051] Step 4.1: Use the relationships in the fact elements extracted in Step 1 and the crime categories divided in Step 2 to determine the crime category to which the target sample belongs.

[0052] Step 4.2: Define weight adjustment factors w1 and w2 for quantifiers and adverbs of degree. The numerical values of the weight adjustment factors for each crime category are different.

[0053] Step 4.3: Use the attention mechanism and the weight adjustment factors to dynamically adjust the weights of the degree elements in the sample, and at the same time use BERT to generate the semantic encoding of the target sample. The weight calculation is shown in formula (3).

[0054]

[0055] where α i is the weight corresponding to the degree element, q T is the transpose of the query vector, h i is the word vector corresponding to the degree element, w p ∈ {w1, w2} is the weight adjustment factor, which can be obtained through training, and N is the total number of words of the degree elements in the target sample.

[0056] Step 5: Weightedly fuse the encoded vectors of the fact elements and the semantic encoding of the sample to obtain the final representation vector

[0057] Step 5.1: Weightedly fuse the encoded vector of the fact elements of the target sample and the semantic encoding of the sample to obtain the final representation vector The weighting method is shown in formula (4).

[0058]

[0059] Among them is the representation vector of factual elements, is the semantic encoding of the sample, and β is the weight, which can be obtained through training, is the final representation vector.

[0060] Step 5.2: Operate the remaining samples according to the operation process of the target sample to obtain various final representation vectors

[0061] Step 6: Calculate the similarity between the final representation vectors of the target sample and other samples using cosine similarity, and select the samples with high similarity as the similar case samples.

[0062] Step 6.1: Calculate the similarity between the final representation vectors of the target sample and other samples using cosine similarity, sort all the similarities, select the sample with the highest similarity as the similar case of the target sample, and output the matching result. The cosine similarity calculation formula is as shown in (5).

[0063]

[0064] Among them is the final representation vector of the remaining vectors, is the final representation vector of the target vector, is the vector 's modulus, is 's modulus, and cosθ is the cosine similarity.

[0065] The above specific description further details the purpose, technical solution and beneficial effects of the invention. It should be understood that the above is only a specific embodiment of the present invention and is not used to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A similar case text matching method based on multi-factor fusion, characterized by The method comprises the following steps: Step 1: First, annotate each sample in the dataset, annotate the factual elements composed of entities and relations, and the degree elements composed of quantifiers and degree adverbs, and then use the pre-trained BERT model to encode the case text into word vectors. Then, input the encoded word vectors into the BiLSTM network to generate a sequence representation, and then use the CRF layer for structured annotation to accurately extract the factual elements and degree elements in the case. The factual elements use a multi-class cross entropy loss function, and the degree elements use a binary cross entropy loss function. The joint loss function is obtained by weighted summation to train the BERT-BiLSTM-CRF model; Step 2: Count all the relationships marked in step 1 and classify them into nine crime categories according to the nature of the crime: violent crime, property crime, drug crime, economic crime, traffic crime, sexual crime, cyber crime, environmental crime, and other crimes; Step 3: Use the successfully trained BERT-BiLSTM-CRF model to extract the factual elements and degree elements in the target sample, use the graph neural network GNN to obtain the encoding vector of the factual elements, and use the attention mechanism to dynamically adjust the weights corresponding to the degree elements in the sample BERT encoding according to the category of the relationship in the factual elements, so as to generate a refined sample semantic encoding, and perform weighted fusion of the sample semantic encoding and the degree encoding of the factual elements to obtain the final representation vector; Step 4: Perform the same operation as step 3 on other samples to obtain their respective final representation vectors. Use cosine similarity to calculate the similarity between the target sample and the final vectors of other samples, sort all similarities in descending order, and select the sample with the highest similarity as the target sample's similar case.

2. The multi-factor fusion similar case text matching method according to claim 1 is characterized by: In step 3, according to the crime category corresponding to the relationship in the factual elements of the sample, the attention mechanism is used to dynamically adjust the weight corresponding to the degree element in the semantic encoding of the sample, so that the weight of the degree element is different in samples of different categories, and a refined semantic encoding is generated. Then, the semantic encoding and the vector encoding of the factual elements are weightedly fused to obtain the final representation vector.