Intelligent litigation case identification method and system based on RAG and knowledge graph

By optimizing the pre-trained text embedding model and comparison learning method, combined with the multi-factor scoring mechanism of the knowledge graph, the problem of low accuracy of entity recognition and disconnection in the litigation case text is solved, and efficient structured processing and credible output of litigation case information is achieved.

CN120542404AActive Publication Date: 2025-08-26STATE GRID GANSU ELECTRIC POWER CO LANZHOU POWER SUPPLY CO

Patent Information

Application Number
CN202511036802.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-08-26
Estimated Expiration
2045-07-28

AI Technical Summary

Technical Problem

In the text processing of litigation cases, the existing technology has problems such as low accuracy of entity recognition and disconnection of the context of knowledge retrieval results, especially in the field of legal professional term modeling and logical reasoning chain construction, resulting in limited credibility and interpretability of the output results.

Method used

By optimizing the parameters of pretrained text embedding models, combining the comparison learning method for entity matching and semantic aggregation, a multi-factor scoring mechanism is constructed, and a knowledge graph is used to generate context corpus and structured information recognition.

Benefits of technology

It improves the accuracy of entity recognition in litigation cases and the completeness of context information, enhances the professional credibility and interpretability of output results, and solves the problems of low accuracy of entity recognition and disconnection of context in the knowledge retrieval results in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120542404A_ABST
    Figure CN120542404A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent litigation case identification method and system based on RAG and a knowledge graph, and relates to the technical field of litigation case identification, and the method comprises the following steps: carrying out semantic vectorization processing on a litigation case text based on a pre-training text embedding model, and outputting a semantic vector sequence; performing similarity matching on the semantic vector sequence and the obtained knowledge graph entity vector to obtain a related graph entity, and constructing a context corpus block corresponding to the litigation case text; and inputting the context corpus block into the large language model, identifying elements of the litigation case through the structured cue word, and outputting structured case information. The method is used for solving the problems that in an existing method, the entity recognition accuracy is low, and the context of a knowledge retrieval result is disjointed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of litigation case identification, and more specifically, to an intelligent litigation case identification method based on RAG and knowledge graph. Background Art

[0002] The structural processing and element identification of litigation case texts are currently of great practical significance in fields such as intelligent justice and legal technology. In particular, in scenarios such as case prediction, compliance analysis, and precedent recommendation, accurately and systematically extracting structured information from original case materials has become a key technical challenge.

[0003] Despite the excellent performance of large language models (LLMs) in natural language processing, they still face challenges in processing legal texts, such as insufficient understanding of specialized terminology and difficulty modeling logical structures. This limits the credibility and interpretability of their output. In recent years, semantic enhancement and retrieval-augmented generation (RAG) methods based on knowledge graphs have emerged as areas for improvement, enhancing the recognition of structured information by introducing external knowledge.

[0004] For example, the invention patent announcement with the publication number CN118885627A discloses a multi-round dialogue processing method and system based on RAG and knowledge graph. The method can effectively align user intentions with answer texts while reducing interactions. Specifically, the method includes: searching and identifying entities associated with user questions based on the proprietary knowledge graph to obtain related entities to build reasoning links and convert them into reasoning text sequences; retrieving text, pictures, and audio data related to user questions through a multimodal large model to obtain relevant multimodal data; integrating the relevant multimodal data, the current reasoning text sequence, the user question, and the generated answer text as model input, inputting the multimodal large model to obtain a model answer; screening out relevant multimodal data, using the screened multimodal data as context and the question to be processed as model inputs of the multimodal large model, and obtaining an answer text corresponding to the question to be processed.

[0005] The above disclosed technical solutions have at least the following technical problems: In the actual implementation process in the legal field, on the one hand, legal scenarios have extremely high requirements for accuracy and controllability, but the current solutions still have shortcomings in professional terminology modeling, logical reasoning chain construction, etc., resulting in low accuracy in entity recognition in litigation cases; on the other hand, the strict logical correlation of legal texts makes the problem of context disconnection of knowledge retrieval results more prominent, making it difficult to effectively integrate retrieval fragments, affecting the accuracy of the output.

[0006] In view of the above problems, the present invention proposes a solution. Summary of the Invention

[0007] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides an intelligent litigation case identification method based on RAG and knowledge graph, which solves the problems of low entity recognition accuracy and disconnection of knowledge retrieval results from the context in the existing methods by optimizing the parameters of the pre-trained text embedding model, constructing a multi-factor scoring mechanism for entity matching, and generating semantically aggregated context corpus blocks.

[0008] To achieve the above object, the present invention provides the following technical solutions: An intelligent litigation case identification method based on RAG and knowledge graph includes the following steps: S1, semantic vectorization processing of litigation case text based on a pre-trained text embedding model, and outputting a semantic vector sequence; S2, similarity matching of the semantic vector sequence with the acquired knowledge graph entity vector to obtain relevant graph entities, and constructing a context corpus corresponding to the litigation case text; S3, inputting the context corpus into a large language model, identifying the elements of the litigation case through structured prompt words, and outputting structured case information.

[0009] In a preferred embodiment, before the semantic vectorization processing of the litigation case text is performed, the litigation case text is also subjected to text cleaning and semantic segmentation.

[0010] In a preferred embodiment, the specific steps for obtaining the semantic vector sequence are as follows: S11, using the upper and lower legal entity labels predefined in the knowledge graph, the input case text is subjected to entity recognition and subclass labeling to obtain the case annotated corpus; S12, based on the annotated corpus, semantic contrast triplets containing positive and negative samples are constructed to obtain a training data set for contrastive learning; S13, using the contrastive learning method to optimize the parameters of the pre-trained text embedding model, so that the semantic vectors of similar entity sentences are close and the semantic vectors of heterogeneous sentences are separated, to obtain a case semantic vector sequence.

[0011] In a preferred embodiment, the contrastive learning method is used to optimize the parameters of the pre-trained text embedding model, and the specific steps are as follows: use the pre-trained Chinese text embedding model to vector encode the semantic contrast triples to obtain the triple sentence vector; construct a loss function based on the triple sentence vector; based on the loss function, use the AdamW optimizer to iteratively update the model parameters to obtain the optimized pre-trained text embedding model.

[0012] In a preferred embodiment, the semantic vector sequence is matched with the acquired knowledge graph entity vector by similarity to obtain related graph entities, specifically: similarity calculation is performed based on the semantic vector sequence and the semantic vector of the entity in the knowledge graph to obtain candidate graph entities; the candidate graph entities are scored by a preset multi-factor combination scoring system, and sorted and screened based on the scoring results to obtain semantically related graph entities; the multi-factor combination scoring system includes a semantic similarity score, an entity hierarchical path score, and an entity connectivity.

[0013] In a preferred embodiment, the specific steps for obtaining the entity hierarchical path score are as follows: determining the path where the candidate entity is located based on the entity hierarchical structure predefined in the knowledge graph; and calculating the entity hierarchical path score based on the number of hierarchical nodes contained in the path, starting from the root node.

[0014] In a preferred embodiment, the specific steps for obtaining the entity connectivity are as follows: based on the construction principles of the knowledge graph, define the connection relationship types between entities in the graph, and assign a preset weight value to each relationship type; extract the core entities of the input case from the case annotation corpus as the target node for measuring the connection relationship; traverse the edge information of the candidate entity in the knowledge graph, and determine whether the candidate entity has a defined connection relationship type with the target node. If so, record the corresponding weight of the connection relationship type; for each candidate entity, add up the weights between it and all target nodes to obtain the entity connectivity of each candidate entity.

[0015] In a preferred embodiment, before constructing the context corpus block corresponding to the litigation case text, it is also necessary to generate a context corpus unit. The specific steps are as follows: the original graph corpus text is vectorized using the optimized pre-trained text embedding model to obtain the semantic representation vector of the entity corpus; the similarity between entities is calculated based on the semantic representation vector, and the entity correlation graph is constructed in combination with the structural relationship in the knowledge graph; the entity correlation graph is divided into communities through a graph clustering algorithm to obtain entity context corpus units; the original graph corpus text is obtained based on the knowledge graph entity mapping method.

[0016] In a preferred embodiment, before outputting structured case information, the process also includes: constructing a multi-hop path index for the case target entity in the knowledge graph; verifying whether the identified field can be derived from the path structure; if the identified field is broken from the path, marking the identified field as an unsupported item for removal.

[0017] An intelligent litigation case identification system based on RAG and knowledge graph includes: a semantic representation generation module, a context corpus construction module and a structured information recognition module; the semantic representation generation module performs semantic vectorization on the litigation case text based on a pre-trained text embedding model and outputs a semantic vector sequence; the context corpus construction module performs similarity matching between the semantic vector sequence and the acquired knowledge graph entity vector to obtain relevant graph entities and construct a context corpus block corresponding to the litigation case text; the structured information recognition module inputs the context corpus block into a large language model, identifies the elements of the litigation case through structured prompt words, and outputs structured case information.

[0018] The technical effects and advantages of the intelligent litigation case identification method and system based on RAG and knowledge graph of the present invention are as follows: 1. This invention improves the accuracy of legal entity recognition and semantic modeling by combining a pre-trained embedding model with a contrastive learning optimization mechanism. This enables the model to distinguish between semantically similar expressions with different legal meanings, thereby enhancing the accuracy of mapping case texts to structured knowledge.

[0019] 2. This invention introduces a knowledge graph-driven entity aggregation and context construction mechanism, integrating hierarchical paths, connection relationships, and semantic similarities in the graph during the entity matching process. It achieves automatic combination and abstract extraction of contextual information, effectively reduces the hallucination phenomenon in the results generated by large language models, and improves the professional credibility of structured recognition results. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 Schematic diagram of the process of the intelligent litigation case identification method based on RAG and knowledge graph of the present invention; Figure 2 Schematic diagram of the structure of the intelligent litigation case identification system based on RAG and knowledge graph of the present invention; Figure 3 The training and validation loss curves for the parameter optimization process of the pre-trained text embedding model of the present invention are shown; Figure 4 This is the cosine similarity distribution diagram of the semantic vectors of positive and negative samples of the present invention; Figure 5 This is the clustering diagram of the t-distributed random neighborhood embedding sentence vector of the present invention. DETAILED DESCRIPTION

[0021] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments derived by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0022] Example 1, Figure 1 The present invention provides an intelligent litigation case identification method based on RAG and knowledge graph, which includes the following steps.

[0023] S1, performs semantic vectorization on the litigation case text based on the pre-trained text embedding model and outputs a semantic vector sequence.

[0024] In this embodiment, before the semantic vectorization processing of the litigation case text is performed, the litigation case text is further cleaned and semantically segmented, specifically: For the input original litigation case text, the text cleaning rules based on regular expressions are used to automatically delete the following content: (1) Line break ( ), tab (\t), and consecutive extra spaces and other typesetting characters; (2) Fixed-format information that is irrelevant to the semantics of the case, such as page number identifiers (e.g., 'Page X of Y pages'), the full name of the court (e.g., 'XX People's Court'), and date stamps (e.g., 'January 1, 2024'); (3) Character encoding errors or unrecognizable garbled characters (such as non-printing characters with ASCII code values ​​less than 32 or greater than 126); During the cleaning process, the factual description of the litigation subject, the judgment opinion and the relevant evidence text are retained to ensure the integrity of the core semantic content; The cleaned text is segmented at the sentence level using Chinese segmentation tools, dividing it into multiple semantic units based on semantic boundaries. For long and complex sentence structures commonly found in the legal field, a dependency parser is further used to identify subject-predicate structures and subordinate relationships, enabling clause segmentation. Call the Chinese word segmentation tool, load a pre-built dictionary of legal terminology, segment the text after segmentation, and annotate part-of-speech information; this dictionary contains high-frequency terms in power grid enterprise litigation, such as "power supply agreement," "property rights dispute," and "line relocation," to improve word segmentation accuracy; By combining rule templates with dictionary-based entity mapping methods, entities such as organization names, personal names, and place names that appear in the text are uniformly standardized. For example, "XX Power Company" and "XX Power Supply Bureau" are uniformly mapped to "a certain power enterprise" to reduce ambiguity.

[0025] In this example, the HanLP tool was used for Chinese sentence segmentation, and the Jieba tool was used for Chinese word segmentation, to achieve semantic unit division and annotation of the input litigation case text. These steps enable structured preprocessing and semantic hierarchical representation of the input case text, providing high-quality, dense semantic input units for subsequent precise association with the knowledge graph and contextual construction.

[0026] Furthermore, the specific steps for obtaining the semantic vector sequence in S1 are as follows: S11 uses predefined upper and lower level legal entity labels in the knowledge graph to perform entity recognition and subcategory annotation on the input case text, generating a case annotated corpus. First, a legal semantically layered entity label system is extracted from the knowledge graph, including top-level labels such as "contractual disputes," "tort liability," and "property ownership," as well as subordinate second-level labels such as "equipment purchase and sales contract" and "line migration responsibility." An entity recognition dictionary is constructed based on this label system, and a BiLSTM-CRF model is used to perform named entity recognition on the input case text, combining word and position features. Subsequently, the identified entities are annotated with subcategories using rule templates and the logical relationships between upper and lower levels of the labels, ultimately generating a fully annotated case corpus. For example, for "migration disputes arising from equipment ownership issues," "property ownership" can be identified as corresponding to the label "property disputes," which can then be further refined to "line migration responsibility."

[0027] S12: Construct semantic contrast triplets containing positive and negative samples based on the annotated corpus to obtain a training dataset for contrastive learning. The specific method is as follows: Each labeled sentence serves as an anchor. A semantically similar positive example (positive) is selected from the same-labeled corpus, and a semantically unrelated negative example (negative) is selected from the different-labeled corpus to form a training triple (anchor, positive, negative). For example, sentences A and B labeled "property rights disputes" are used as anchors, and sentence C, semantically unrelated to "construction nuisance disputes," is selected as a negative example to form a training triple (A, B, C).

[0028] S13 uses a contrastive learning method to optimize the parameters of the pre-trained text embedding model, bringing the semantic vectors of similar entity sentences closer together and separating the semantic vectors of heterogeneous sentences, thereby obtaining a case semantic vector sequence. Using the contrastive learning method based on the SimCSE architecture, the pre-trained Chinese language model RoBERTa-wwm-ext is selected as the basic embedding network. After inputting the aforementioned triples, vector calculations are performed. The model parameters are optimized using a contrastive loss function, ensuring that the semantic vector distances of similar sentences converge and the distances between heterogeneous sentence vectors increase. After training is complete, any case text related to a power grid enterprise is input to output a structured semantic vector sequence, which serves as the basis for subsequent knowledge graph matching and contextual reasoning.

[0029] The contrastive learning method is used to optimize the parameters of the pre-trained text embedding model. The specific optimization steps are as follows: S131, use the pre-trained Chinese text embedding model RoBERTa-wwm-ext to vector encode the semantic contrast triples and obtain the triple sentence vector. , positive sample sentences and negative sample sentences , respectively obtain their sentence vector representations:

[0030] in, is the anchor sentence vector, is the positive sample sentence vector, is the negative sample sentence vector, [CLS] bit vector representing the output of the pre-trained model.

[0031] S132, construct the InfoNCE loss function based on the triple sentence vector. Use the positive sample as the anchor point's enhanced copy, and the negative sample as other instances in the training batch to calculate the softmax similarity distribution. The specific calculation formula is as follows:

[0032] in, is the InfoNCE comparison loss value, represents the cosine similarity, is the sentence vector of any comparison sample in the batch, that is, the sentence vectors of all samples in the current training batch, including positive samples and multiple negative samples, is the temperature coefficient, set to 0.05.

[0033] S133, based on the loss function, uses the AdamW optimizer to iteratively update the model parameters.

[0034] In this embodiment, the training cycle is set to 10 rounds, 1000 sets of triple samples are extracted in each round, the training batch size is set to 64, and the learning rate is 2e-5. Figure 3 As shown in the training and validation loss curves, the vertical axis is the InfoNCE comparison loss value, and the loss values ​​of the training set and the validation set show a gradual downward trend; Figure 4 As shown in the cosine similarity distribution of semantic vectors, after training, the cosine similarity of positive samples is concentrated above 0.8, and the cosine similarity of negative samples is concentrated below 0.3; Figure 5 The t-distributed stochastic neighbor embedding (t-SNE) sentence vector clustering results show that sample sentences with different legal labels are clustered on a two-dimensional plane, indicating that the pre-trained Chinese text embedding model RoBERTa-wwm-ext has enhanced its semantic extraction ability for legal texts after the above optimization.

[0035] Using contrastive learning methods to optimize the pre-trained text embedding model can significantly improve the model's ability to distinguish fine-grained semantic differences in litigation cases, making the case description vectors of the same legal entity category more aggregated and the description vectors of different categories more separated.

[0036] By introducing a training mechanism that compares positive and negative samples, the above steps enable the model to not only rely on context for embedding learning but also proactively identify and determine semantic boundaries. This effectively addresses the semantic confusion and vector overlap issues inherent in traditional embedding models for legal entities. This makes the matching between semantic vectors and knowledge graph entities more accurate and the contextual recall results more targeted in subsequent steps, significantly improving the accuracy and stability of overall case element recognition and ensuring good scalability and portability.

[0037] S2, performs similarity matching between the semantic vector sequence and the acquired knowledge graph entity vector to obtain relevant graph entities and construct the context corpus corresponding to the litigation case text.

[0038] In this embodiment, the semantic vector sequence is matched with the acquired knowledge graph entity vector by similarity to obtain the relevant graph entity, specifically: Calculate the similarity between the semantic vector sequence and the semantic vector of the entity in the knowledge graph to obtain the candidate graph entity; The candidate graph entities are scored through a preset multi-factor combination scoring system, and are sorted and screened based on the scoring results to obtain semantically relevant graph entities; The multi-factor combination scoring system includes semantic similarity score, entity level path score and entity connectivity.

[0039] The specific steps for obtaining the semantic similarity score are as follows: Based on the case semantic vector obtained in S1 , calculate its semantic vector with each entity in the knowledge graph The cosine similarity of , the scoring formula is as follows:

[0040] in, Represents the modulus of the vector. A higher score indicates a closer match between the case semantics and the entity semantics. Using this metric and a top-10 strategy, we select the 10 entities in the knowledge graph with the highest similarity to the case semantic vector as candidate graph entities.

[0041] The specific steps for obtaining the entity-level path score are as follows: Determine the path where the candidate entity is located based on the predefined entity hierarchy in the knowledge graph; Based on the number of hierarchical nodes d contained in the path, starting from the root node, which is the 0th layer, the entity hierarchical path score Score is calculated L , the specific calculation formula is as follows:

[0042] The hierarchical path score is used to reflect the proximity of the candidate entity to the root node in the graph structure, so as to give priority to retaining candidate entities with more reasonable structural positions. For example, in the legal compliance knowledge graph, there is a structure shown in Table 1, and the corresponding path depth is: Candidate entity A (construction contract dispute): The path is "legal dispute -> civil dispute -> contract dispute -> construction contract dispute", the hierarchy depth is d=3, and the path score is 0.25; Candidate entity B (line property rights dispute): The path is "legal dispute -> civil dispute -> property rights dispute -> line property rights dispute", the hierarchy depth is d=3, and the path score is 0.25; Candidate entity C (civil dispute): The path is "legal dispute -> civil dispute", the hierarchical depth is d=1, and the path score is 0.5.

[0043] Table 1

[0044] The specific steps for obtaining the entity connectivity are as follows: Based on the construction principles of the knowledge graph, the connection relationship types between entities in the graph are defined, and a preset weight value is assigned to each relationship type, as shown in Table 2; Table 2

[0045] Extract the core entities of the input case from the case annotation corpus as the target nodes for measuring the connection relationship; Traverse the edge information of the candidate entity in the knowledge graph to determine whether the candidate entity has a defined connection relationship type with the target node. If so, record the corresponding weight of the connection relationship type; For each candidate entity, add up the weights between it and all target nodes to get the entity connectivity score of each candidate entity. C , the specific formula is as follows:

[0046] in, For candidate entity e and target node c k The connection weight between them, K is the number of target nodes.

[0047] The traversal of the edge information of the candidate entity in the knowledge graph determines whether the candidate entity has a defined connection relationship type with the target node, specifically: For each candidate entity, find its outgoing and incoming edges in the graph, that is, all connected adjacent nodes and edge types; Determine whether any of these adjacent nodes contains any target node; If a connection exists, its weight is recorded according to the type of edge.

[0048] By introducing an entity connectivity scoring mechanism based on knowledge graph relationship types, we can effectively enhance the structural relevance judgment between candidate graph entities and case context, avoiding the semantic drift problem caused by relying solely on semantic similarity. This scoring strategy leverages the pre-set legal domain-specific connectivity relationships in the graph, such as regulatory citations, precedent attribution, and business processes, to ensure that the selected graph entities have a significant structural legal connection with the core content of the case, thereby improving the accuracy and semantic integrity of the contextual corpus construction, significantly enhancing the pertinence and discriminative accuracy of subsequent case element identification.

[0049] Finally, for each candidate graph entity, the comprehensive score is calculated by combining the above three scoring indicators. The specific weighted calculation formula is as follows:

[0050] Among them, α, β, and γ are weight coefficients of each scoring item, and in this embodiment, they are 0.4, 0.2, and 0.2 respectively. Taking the case of "line property rights dispute" as an example, the scores of some candidate entities are shown in Table 3: Table 3

[0051] Furthermore, the specific steps for constructing the context corpus corresponding to the litigation case text in S2 are as follows: S21, using the knowledge graph entity mapping method to obtain the attributes, definitions and instance information of the relevant graph entities, and obtain the original graph corpus text, specifically: Extract the entity's "definition", "scope of application", "related regulations", and "practical cases" attributes from the knowledge graph; Use entity uniform resource identifier mapping to match the semantic vector obtained in S13 to the corresponding entity node; Integrate the description texts of all entities and their adjacent entities into the original graph corpus text.

[0052] For example, for the candidate entity “line migration responsibility,” its original graph corpus includes the definition: “refers to the responsibility division standard caused by the migration of property lines due to engineering construction in power grid construction”; an example: “The 110kV line in a certain place was relocated due to high-speed construction, resulting in multi-party ownership disputes”; as well as the superordinate entity “property rights disputes” and its linked regulations.

[0053] S22, based on the semantic relationship and structural connectivity between entities in the original graph corpus text, aggregate and generate context corpus units, specifically: The optimized pre-trained text embedding model is used to vectorize the original graph corpus text to obtain the semantic representation vector of the entity corpus; The similarity between entities is calculated based on the semantic representation vector, and the entity relationship graph is constructed by combining the structural relationships in the knowledge graph. The similarity calculation formula is the same as the semantic similarity score formula mentioned above. For example, suppose the input entity set is: E1 is "line migration responsibility", E2 is "equipment ownership dispute", E3 is "property preservation request", E4 is "property registration system", and E5 is "execution avoidance objection"; Construct the semantic similarity matrix S{s ij}∈R 5×5 as follows:

[0054] Among them, each element s ij Represents entity E i With E j The semantic similarity between entities is the node in the graph. If the similarity s ij If the preset threshold θ=0.8 is exceeded, then the entity E i With E j Create an edge between them with an edge weight of s ij , thereby constructing an entity-related graph; The entity-related graph is divided into communities using the Louvain community discovery graph clustering algorithm. Each clustering result is an "entity context corpus unit", and the corpus will be provided as a whole to the subsequent large language model for processing.

[0055] S23, using a text summarization algorithm to refine the context corpus unit and perform format standardization processing to obtain a context corpus block, specifically: The BERT-SUM model is used to perform semantic compression on contextual corpus units; The paragraph structure of the semantically compressed context corpus unit is unified, redundant punctuation and noise characters are cleaned, and the order of entity description, cited regulations, and case fields is kept consistent to form the final context corpus block.

[0056] It should be noted that in this embodiment, all knowledge graph entity vectors, entity semantic vectors and related graph entity vectors used for calculation are obtained through the optimized pre-trained text embedding model in S13.

[0057] S3, inputs the context corpus block into the large language model, identifies the elements of the litigation case through structured prompt words, and outputs structured case information.

[0058] In this embodiment, the specific steps of S3 are as follows: Construct a prompt word template with legal semantic prompt function to guide the language model to perform inductive reasoning on context information. Example prompt words are as follows: "Based on the following case text and background materials, please identify and complete the legal structure of the case: plaintiff, defendant, focus of dispute, legal basis, request, and resolution." The case text, context corpus, and prompt word template are concatenated to form the model input in the following format: [Case text]: ...

Background

Task

[0059] Based on the above input, the model outputs a structured result in the form of a JSON string. The following is an example output: { "Plaintiff": "A power supply company", "Defendant": "An equipment company", "Dispute Focus": "Equipment Ownership Issue", "Legal Basis": "Article XX of the XX Contract Law", "Request": "Compensation for equipment loss and reinstallation costs", "Disposition Conclusion": "The defendant is ordered to bear the responsibility for relocation"}.

[0060] To ensure the legal rationality of the identification elements, the S3 also includes the following before outputting the structured case information: By constructing a multi-hop path index for the case target entity in the knowledge graph; Verify whether the identified field can be derived from the path structure. If the identified field is disconnected from the path, mark the identified field as an unsupported item and remove it.

[0061] It should be noted that "multi-hop path index" refers to a set of paths in the knowledge graph that starts from the case target entity and connects to other related entities through multiple intermediate entity nodes. The path length is greater than 1 hop and is used to capture indirectly related semantic relationships.

[0062] For example, taking a certain electricity-related litigation case as an example, the original text recognition results are as follows: { "Plaintiff": "A power supply company", "Defendant": "A property management company", "Dispute Focus": "Responsibility for Line Relocation", "Legal basis": "Article XX of the XX Electricity Law", "Request": "Refund of relocation costs and compensation for lost work time" }.

[0063] In the knowledge graph path index, "Line Relocation Responsibility" -> "Line Property Rights" -> "Relocation Obligations" -> "Expense Bearing Rules" can support the "Refund of Relocation Expenses" field, but there is no direct path to the "Compensation for Lost Work Losses" node. Therefore, the system marks the latter as "Unsupported Item" and excludes it from the final output: { "Plaintiff": "A power supply company", "Defendant": "A property management company", "Dispute Focus": "Responsibility for Line Relocation", "Legal basis": "Article XX of the XX Electricity Law", "Request": "Refund of relocation fees" }.

[0064] Example 2, Figure 2 The present invention provides an intelligent litigation case identification system based on RAG and knowledge graph, which includes: a semantic representation generation module, a context corpus construction module and a structured information recognition module; The semantic representation generation module performs semantic vectorization on the litigation case text based on the pre-trained text embedding model and outputs a sequence of semantic vectors; The context corpus construction module performs similarity matching between the semantic vector sequence and the acquired knowledge graph entity vectors to obtain relevant graph entities and construct the context corpus block corresponding to the litigation case text; The structured information recognition module inputs the context corpus into the large language model, identifies the elements of the litigation case through structured prompt words, and outputs structured case information.

[0065] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.

[0066] The above embodiments may be implemented in whole or in part through software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product.

[0067] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0068] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0069] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0070] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. An intelligent litigation case identification method based on RAG and knowledge graph, characterized by: The following steps are involved: Perform semantic vectorization on the litigation case text based on the pre-trained text embedding model and output a semantic vector sequence; Perform similarity matching between the semantic vector sequence and the acquired knowledge graph entity vector to obtain relevant graph entities and construct the context corpus corresponding to the litigation case text; Input the context corpus into the large language model, identify the elements of the litigation case through structured prompt words, and output structured case information; Before constructing the context corpus block corresponding to the litigation case text, it is necessary to generate context corpus units. The specific steps are as follows: The optimized pre-trained text embedding model is used to vectorize the original graph corpus text to obtain the semantic representation vector of the entity corpus; Calculate the similarity between entities based on semantic representation vectors, and build entity relationship graphs based on structural relationships in the knowledge graph; The entity-related graph is divided into communities through graph clustering algorithms to obtain entity context corpus units; The original graph corpus text is obtained based on the knowledge graph entity mapping method.

2. The intelligent litigation case identification method based on RAG and knowledge graph according to claim 1 is characterized in that: Before the semantic vectorization processing of the litigation case text is performed, the litigation case text is also cleaned and semantically segmented.

3. The intelligent litigation case identification method based on RAG and knowledge graph according to claim 2 is characterized in that: The specific steps for obtaining the semantic vector sequence are as follows: Using the predefined upper and lower legal entity labels in the knowledge graph, we perform entity recognition and subclass annotation on the input case text to obtain the case annotation corpus. Construct semantic contrast triplets containing positive and negative samples based on the annotated corpus to obtain a training dataset for contrastive learning; The contrastive learning method is used to optimize the parameters of the pre-trained text embedding model so that the semantic vectors of similar entity sentences are close and the semantic vectors of heterogeneous sentences are separated, and a case semantic vector sequence is obtained.

4. The intelligent litigation case identification method based on RAG and knowledge graph according to claim 3 is characterized in that: The contrastive learning method is used to optimize the parameters of the pre-trained text embedding model. The specific steps are as follows: Use the pre-trained Chinese text embedding model to vectorize the semantic contrast triples and obtain the triple sentence vectors; Construct loss function based on triple sentence vector; Based on the loss function, the AdamW optimizer is used to iteratively update the model parameters to obtain the optimized pre-trained text embedding model.

5. The intelligent litigation case identification method based on RAG and knowledge graph according to claim 4 is characterized in that: The semantic vector sequence is matched with the acquired knowledge graph entity vector by similarity to obtain the relevant graph entity, specifically: Calculate the similarity between the semantic vector sequence and the semantic vector of the entity in the knowledge graph to obtain the candidate graph entity; The candidate graph entities are scored through a preset multi-factor combination scoring system, and are sorted and screened based on the scoring results to obtain semantically related graph entities. The multi-factor combination scoring system includes semantic similarity score, entity hierarchical path score and entity connectivity.

6. The intelligent litigation case identification method based on RAG and knowledge graph according to claim 5 is characterized in that: The specific steps for obtaining the entity-level path score are as follows: Determine the path where the candidate entity is located based on the predefined entity hierarchy in the knowledge graph; The entity-level path score is calculated based on the number of level nodes included in the path, starting from the root node.

7. The intelligent litigation case identification method based on RAG and knowledge graph according to claim 6 is characterized in that: The specific steps for obtaining the entity connectivity are as follows: Based on the construction principles of the knowledge graph, define the connection relationship types between entities in the graph and assign a preset weight value to each relationship type; Extract the core entities of the input case from the case annotation corpus as the target nodes for measuring the connection relationship; Traverse the edge information of the candidate entity in the knowledge graph to determine whether the candidate entity has a defined connection relationship type with the target node. If so, record the corresponding weight of the connection relationship type; For each candidate entity, the weights between it and all target nodes are summed to obtain the entity connectivity of each candidate entity.

8. The intelligent litigation case identification method based on RAG and knowledge graph according to claim 7 is characterized in that: Before outputting the structured case information, the method further includes: By constructing a multi-hop path index for the case target entity in the knowledge graph; Verify whether the identified field can be derived from the path structure. If the identified field is disconnected from the path, mark the identified field as an unsupported item and remove it.

9. A system for implementing the intelligent litigation case identification method based on RAG and knowledge graph according to claims 1-8, characterized in that: include: Semantic representation generation module, context corpus construction module, and structured information recognition module; The semantic representation generation module performs semantic vectorization on the litigation case text based on the pre-trained text embedding model and outputs a sequence of semantic vectors; The context corpus construction module performs similarity matching between the semantic vector sequence and the acquired knowledge graph entity vectors to obtain relevant graph entities and construct the context corpus block corresponding to the litigation case text; The structured information recognition module inputs the context corpus into the large language model, identifies the elements of the litigation case through structured prompt words, and outputs structured case information.

Citation Information

Patent Citations

  • Information retrieval query method for legal data service platform

    CN118981512A

  • Legal element analysis method and system based on large language model and knowledge graph

    CN119046476A

  • Method for constructing knowledge graph based on large language model and vector library

    CN119129722A

  • Method and system for handling small-amount unhealthy asset litigation cases

    CN120256528A

  • Method for training natural language processing model, and method for generating subsequent text of dialogue

    WO2025118396A1

Cited By

  • Legal element map auxiliary class case recommendation method and device and program product

    CN120780850A