Automatic construction method based on legal big language model and legal knowledge graph and intelligent question answering system thereof
Through the method of automatic construction of legal large language models and legal knowledge graphs, the problem that knowledge graph construction in the existing technology depends on predefined entities and relationships is solved, and an unsupervised knowledge graph construction and intelligent question-and-answer system is realized, which improves the accuracy and generalization capabilities of question-and-answer.
Patent Information
- Application Number
- CN202510211104.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art usually relies on predefined entities and relationships when building knowledge graphs, and most of them use supervised learning methods require a large amount of manual annotations, which limits its application scope and efficiency.
A method based on the legal large language model and legal knowledge graph is proposed. Through the legal large language model fine-tuning module, the knowledge graph automatic construction module and the intelligent question and answer module, an unsupervised knowledge graph construction and intelligent question and answer system are realized.
By introducing knowledge graph technology, the accuracy and generalization capabilities of the question-and-answer system are improved, and the model needs to be retrained due to changes in the diversity of target answers is avoided, which significantly enhances the performance and scalability of the model.
Smart Images

Figure CN120123480A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large language model applications, and particularly relates to a method for automatically constructing a legal large language model and a legal knowledge graph and an intelligent question answering system thereof. Background Art
[0002] Automatic knowledge graph construction aims to create a structured representation of knowledge from different data sources without manual intervention. The knowledge graph construction (KGC) process typically uses named entity recognition, relation extraction, and entity resolution techniques to convert unstructured text into a structured representation that can capture entities, entity relationships, and related attributes. Luan et al. developed a multi-task model for identifying entities, relationships, and co-reference clusters in scientific articles to support the creation of scientific knowledge graphs. Mehta et al. proposed an end-to-end knowledge graph construction system and a new deep learning-based predicate mapping model. Their system identifies and extracts entities and relationships from text and maps them to the DBpedia namespace. Although these processes can achieve satisfactory results and generate high-quality knowledge graphs, their methods are usually limited to a predefined set of entities and relationships or rely on a specific ontology. Our proposal addresses these limitations because we do not rely on a predefined set or external ontology. In addition, almost all methods adopt supervised methods that require a large amount of manual annotation. To solve this problem, since the emergence of the first pre-trained language model, the question has been: Can we use the knowledge stored in the pre-trained language model to construct a knowledge graph? Wang et al. were among the first to propose and solve this problem. They designed an unsupervised method called MAMA that constructs a knowledge graph by performing a single forward pass of the pre-trained language model on a corpus without any fine-tuning.
[0003] Information retrieval knowledge graph question answering (IR-KGQA) represents the integration of two powerful technologies in artificial intelligence: information retrieval (IR) and knowledge graph question answering (KGQA). While traditional question answering (QA) systems are good at understanding and responding to user queries in natural language, IR-KGQA aims to enhance this process by combining structured knowledge in the knowledge graph with unstructured data retrieved from text corpora, databases, or other large information sources. This hybrid system aims to provide more accurate, context-rich answers, especially for complex or open-ended questions where a single data source may not be sufficient.
[0004] In summary, the rapid development of the Chinese legal large language model provides strong support for legal AI applications, but continuous optimization is still needed in terms of question answering accuracy and model generalization ability. By continuously combining the knowledge graph and the legal large language, the legal large language model is expected to play a more important role in future legal practice. Summary of the Invention
[0005] In view of the technical defects in the background art, the present invention proposes a method for automatically constructing based on a legal large language model and a legal knowledge graph and its intelligent question-answering system, which solves the above technical problems and meets the actual needs. The specific technical solutions are as follows:
[0006] An intelligent question-answering system automatically constructed based on a legal large language model and a legal knowledge graph includes a legal large language model fine-tuning module, a knowledge graph automatic construction module, and an intelligent question-answering module;
[0007] The legal large language model fine-tuning module includes an embedding module, a pre-trained large model, a prompt module, and a LoRA fine-tuning module; the knowledge graph automatic construction module includes a candidate triple extraction module, an entity / predicate parsing module, and a pattern reasoning module; the intelligent question-answering module includes a dynamic few-shot retrieval module, a multi-query generation module, and an answer selection module.
[0008] As an improvement of the above solution, the embedding module includes a phrase embedding layer, a phrase type embedding layer, and a position embedding layer; the Chinese base language model of the legal large language model fine-tuning module is the ChatGLM-6B second-generation model.
[0009] As an improvement of the above solution, the candidate triple extraction module includes a text splitting module, an entity extraction module, and a relationship extraction module; the entity / predicate parsing module includes a semantic aggregation module, a clustering disambiguation module, and a concept contraction module, and the pattern reasoning module includes a hypernym generation module and a hierarchical aggregation module.
[0010] A method for automatically constructing based on a legal large language model and a legal knowledge graph, applied to an intelligent question-answering system automatically constructed based on a legal large language model and a legal knowledge graph, includes the following steps;
[0011] Step 1: Fine-tuning of the legal large language model:
[0012] Input a legal question-answer pair dataset, and convert the questions and answers in the legal question-answer pair dataset into feature vectors through the embedding module;
[0013] Use the prompt module to optimize the feature vectors and generate prompt vectors containing legal semantics;
[0014] Input the prompt vectors into the Chinese large language model, and calculate the cross-entropy loss between the predicted answer and the target answer;
[0015] Freeze the parameters of the Chinese base language model, and update the bypass parameters of the low-rank matrix in the Chinese large language model through the LoRA fine-tuning module until the model converges.
[0016] After the model converges, the parameters of the low-rank matrix and the parameters of the Chinese large language model are merged to obtain the fine-tuned legal large language model
[0017] Step 2; Automatic construction of legal knowledge graph:
[0018] Preprocess the legal text data, and generate candidate triples through text splitting, entity extraction, and relationship extraction;
[0019] Input the candidate triples into the entity / predicate parsing module, perform semantic aggregation, clustering disambiguation, and concept contraction on the candidate triples to generate parsed entity clusters;
[0020] Input the parsed entity clusters into the pattern reasoning module, and the pattern reasoning module finally outputs a knowledge graph containing legal knowledge;
[0021] Step 3; Intelligent question answering processing;
[0022] Receive the legal questions input by the user, and extract the entities and relationships in the questions;
[0023] Through the dynamic few-shot retrieval module, calculate the cosine similarity between the question and the stored examples in the knowledge graph, and retrieve the top k most relevant examples;
[0024] Generate multiple SPARQL query candidates through the multi-query generation module in combination with the retrieved examples. The generation process adopts the beam search strategy to retain the top-n query sequences with the highest probability;
[0025] Execute the maximum set (LS) heuristic method through the answer selection module, and select the answer with the largest number of returned results as the final output.
[0026] As an improvement to the above solution, in step 1, the LawGPT_zh large model is used for pre-training to obtain the legal question-answer pair dataset. The specific steps for the embedding module to convert the legal question-answer pair dataset into feature vectors are as follows;
[0027] Step 1.1, the phrase embedding layer converts each phrase into a corresponding word vector and adds special tokens;
[0028] Step 1.2, the phrase type embedding layer embeds the word vectors of specific types into the overall word vectors;
[0029] Step 1.3, the position embedding layer integrates the position information of each word vector into the overall word vector, and through the processing of the normalization layer and the linear layer, the alignment of the word vectors is achieved;
[0030] Step 1.4, the word vectors are input into the pre-trained large model Roberta to align the text and the corresponding answers to form special vectors.
[0031] As an improvement to the above solution, in step 1, the feature vector is optimized using a prompter module, specifically including:
[0032] Step 1.5, input the feature vector into the prompter. The input and output of the prompter are repeatedly looped three times. Input the feature vector into the prompter. The prompter is based on the Transformer architecture and generates a legal semantic prompt vector through the self-attention mechanism to align the semantic association between the question and the answer.
[0033] Step 1.6, input the prompt vector into the Chinese large language model, calculate the difference between the target answer and the predicted answer using the cross-entropy loss function, and use this loss function to update the parameters of the low-rank matrix in the Chinese large language model by means of LoRA fine-tuning.
[0034] Step 1.7, through continuous training, the cross-entropy loss gradually decreases until the model converges. After the model converges, merge the parameters of the low-rank matrix and the parameters of the Chinese base language model to obtain the fine-tuned legal large language model.
[0035] As an improvement to the above solution, step 1.6 specifically includes adding a bypass structure in the Chinese large language model. This bypass structure is obtained by multiplying two matrices. During training, freeze the parameters of the Chinese large language model and only update the parameters of the low-rank matrix that multiplies in the bypass structure. Through continuous training, the cross-entropy loss gradually decreases, and the overall performance of the model gradually improves, finally converging to the optimal state.
[0036] As an improvement to the above solution, in step 2, the steps for generating candidate triples specifically include:
[0037] Step 2.1, preprocess the legal text data, and split the long text into short text blocks not exceeding the token limit of the language model through the text splitting module.
[0038] Step 2.2, perform two-stage extraction on each text block, respectively:
[0039] Entity extraction: Use the entity extraction module to identify entities related to the target entity, and use the legal large language model to identify the entities and their types to generate a list of candidate entities.
[0040] Relationship extraction: Use the relationship extraction module to identify the relationships between the context entities generated in the previous step, generate RDF triples based on the semantic association between entities, and perform a normalized description of the predicates.
[0041] Step 2.3, input the candidate triples into the entity / predicate parsing module, perform semantic aggregation and clustering disambiguation on the candidate triples, and unify the entity / predicate identifiers through concept contraction to generate a parsed entity cluster.
[0042] Step 2.4, input the parsed entity clusters into the pattern reasoning module, and construct a hierarchical legal knowledge classification system through the pattern reasoning module, specifically including:
[0043] Hypernym generation: Generate general hypernyms for entity type clustering;
[0044] Hierarchical aggregation: Merge hypernyms to form a multi-level taxonomy until a single top-level classification is generated.
[0045] According to the method automatically constructed based on the legal large language model and the legal knowledge graph as described in claim 4, characterized in that, in step 3, the dynamic few-shot retrieval includes;
[0046] Step 3.1, use a sentence encoder to map the question, entity, and relationship into vector representations, specifically, use a sentence encoder to map the input question q and the entities and relationships extracted from it to vector representations ;
[0047] Step 3.2, calculate the matching degree with the knowledge graph examples through formula G1;
[0048] Step 3.3, take the k examples with the highest matching degree as context hints and input them into the legal large language model to generate a SPARQL query;
[0049] As an improvement of the above solution, the formula G1 is: , where is a similarity function, and k most similar examples S are retrieved based on this score; use a sentence encoder to map the input question q and the entities and relationships extracted from it to vector representations , by concatenating the question, entity, and relationship into a single input sequence , which is processed by a legal language model containing only an encoder, and each example is encoded as a vector , then, the module calculates the similarity between the input question vector and the example vector to retrieve the most relevant examples from the storage.
[0050] The beneficial effects of the present invention are as follows:
[0051] 1. Method of introducing knowledge graph and reasonably connecting questions with answers: The present invention provides a more reasonable and effective solution by introducing a method for generating a knowledge graph to organically connect legal questions with target answers. In addition, through the candidate triple module, legal questions and answers are converted into triples and used as the input of the entity / predicate parsing module, thereby achieving excellent training effects and improving the overall performance of the model.
[0052] 2. Improving generalization ability and calculation accuracy: Based on knowledge graph technology, the present invention combines the knowledge graph with a large language model to achieve efficient integration of questions and answers. By avoiding the need to retrain the model due to the diverse changes in target answers, the present invention significantly enhances the generalization ability of the model and effectively improves the calculation accuracy.
[0053] 3. Clear logic, strong flexibility, and scalability: The model logic of the present invention is clear at each stage, the calculation method is flexible and has strong scalability. By flexibly setting the network structure, different calculation requirements can be met, and it can be quickly extended to a distributed and parallel development environment. In particular, the combination of the knowledge graph and the large language model demonstrates unique innovation and practicality in distributed computing. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 It is a technical framework diagram of the present invention.
[0055] Figure 2 It is a logical structure diagram for implementing fine-tuning of a Chinese large language model based on a prompter of the present invention.
[0056] Figure 3 It is a logical structure diagram for implementing a legal intelligent question-answering system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] The following describes the embodiments of the present invention in conjunction with the accompanying drawings and related embodiments. The embodiments of the present invention are not limited to the following embodiments, and the related necessary components involved in the present invention should be regarded as well-known technologies in the technical field, which can be known and mastered by those skilled in the technical field.
[0058] An intelligent question-answering system automatically constructed based on a legal large language model and a legal knowledge graph includes a legal large language model fine-tuning module, a knowledge graph automatic construction module, and an intelligent question-answering module;
[0059] The legal large language model fine-tuning module includes an embedding module, a pre-trained large model, a prompt module, and a LoRA fine-tuning module; the knowledge graph automatic construction module includes a candidate triple extraction module, an entity / predicate parsing module, and a pattern reasoning module; the intelligent question-answering module includes a dynamic few-shot retrieval module, a multi-query generation module, and an answer selection module.
[0060] Further, in the above solution, the embedding module includes a phrase embedding layer, a phrase type embedding layer, and a position embedding layer; the Chinese base language model of the legal large language model fine-tuning module is the ChatGLM-6B second-generation model.
[0061] Further, in the above solution, the candidate triple extraction module includes a text splitting module, an entity extraction module, and a relationship extraction module; the entity / predicate parsing module includes a semantic aggregation module, a clustering disambiguation module, and a concept contraction module, and the pattern reasoning module includes a hypernym generation module and a hierarchical aggregation module.
[0062] A method based on the automatic construction of a legal large language model and a legal knowledge graph, applied to an intelligent question-answering system based on the automatic construction of a legal large language model and a legal knowledge graph, includes the following steps;
[0063] Step 1; Legal large language model fine-tuning:
[0064] Input the legal question-answer pair dataset, and convert the questions and answers in the legal question-answer pair dataset into feature vectors through the embedding module;
[0065] Use the prompt module to optimize the feature vectors and generate prompt vectors containing legal semantics;
[0066] Input the prompt vectors into the Chinese large language model, and calculate the cross-entropy loss between the predicted answer and the target answer;
[0067] Freeze the parameters of the Chinese base language model, and update the bypass parameters of the low-rank matrix in the Chinese large language model through the LoRA fine-tuning module until the model converges.
[0068] After the model converges, merge the parameters of the low-rank matrix and the parameters of the Chinese large language model to obtain the fine-tuned legal large language model
[0069] Step 2; Automatic construction of legal knowledge graph:
[0070] Preprocess the legal text data, and generate candidate triples through text splitting, entity extraction, and relationship extraction;
[0071] Input the candidate triples into the entity / predicate parsing module, perform semantic aggregation, clustering disambiguation, and concept contraction on the candidate triples, and generate parsed entity clusters;
[0072] The parsed entity clusters are input into the pattern reasoning module, and the pattern reasoning module finally outputs a knowledge graph containing legal knowledge;
[0073] Step 3; Intelligent question answering processing;
[0074] Receive the legal questions input by the user, and extract the entities and relationships in the questions;
[0075] Through the dynamic few-shot retrieval module, calculate the cosine similarity between the question and the stored examples in the knowledge graph, and retrieve the top-k most relevant examples;
[0076] Through the multi-query generation module, combine the retrieved examples to generate multiple SPARQL query candidates. The generation process adopts the beam search strategy to retain the top-n query sequences with the highest probability;
[0077] Through the answer selection module, execute the largest set (LS) heuristic method, and select the answer with the largest number of returned results as the final output.
[0078] Furthermore, in the above solution, in step 1, the large model LawGPT_zh is used for pre-training to obtain the legal question-answer pair dataset. The specific steps for the embedding module to convert the legal question-answer pair dataset into feature vectors are as follows;
[0079] Step 1.1, the phrase embedding layer converts each phrase into a corresponding word vector and adds special tokens;
[0080] Step 1.2, the phrase type embedding layer embeds the word vectors of specific types into the overall word vectors;
[0081] Step 1.3, the position embedding layer integrates the position information of each word vector into the overall word vector. After the processing of the normalization layer and the linear layer, the alignment of the word vectors is achieved;
[0082] Step 1.4, the word vectors are input into the pre-trained large model Roberta to align the text and the corresponding answers to form special vectors.
[0083] Furthermore, in the above solution, in step 1, the prompt module is used to optimize the feature vectors, specifically including:
[0084] Step 1.5, input the feature vectors into the prompt. The input and output of the prompt are repeated three times. Input the feature vectors into the prompt. The prompt is based on the Transformer architecture and generates legal semantic prompt vectors through the self-attention mechanism to align the semantic associations between the questions and the answers;
[0085] Step 1.6, input the prompt vector into the Chinese large language model, calculate the difference between the target answer and the predicted answer using the cross-entropy loss function, and use this loss function to update the parameters of the low-rank matrix in the Chinese large language model by means of LoRA fine-tuning;
[0086] Step 1.7, through continuous training, the cross-entropy loss gradually decreases until the model converges. After the model converges, merge the parameters of the low-rank matrix and the parameters of the Chinese base language model to obtain the fine-tuned legal large language model.
[0087] Further, in the above solution, Step 1.6 specifically includes adding a bypass structure in the Chinese large language model. This bypass structure is obtained by multiplying two matrices. During training, freeze the parameters of the Chinese large language model and only update the parameters of the low-rank matrix that multiplies in the bypass structure. Through continuous training, the cross-entropy loss gradually decreases, and the overall performance of the model gradually improves, and finally converges to the optimal state.
[0088] Further, in the above solution, in Step 2, the steps of generating candidate triples specifically include:
[0089] Step 2.1, preprocess the legal text data, and split the long text into short text blocks not exceeding the token limit of the language model through the text splitting module;
[0090] Step 2.2, perform two-stage extraction on each text block, which are respectively:
[0091] Entity extraction: Use the entity extraction module to identify entities related to the target entity, use the legal large language model to identify entities and their types, and generate a list of candidate entities;
[0092] Relationship extraction: Use the relationship extraction module to identify the relationships between the context entities generated in the previous step, generate RDF triples based on the semantic associations between entities, and perform a normalized description of the predicates;
[0093] Step 2.3, input the candidate triples into the entity / predicate parsing module, perform semantic aggregation and clustering disambiguation on the candidate triples, and unify the entity / predicate identifiers through concept contraction to generate a parsed entity cluster;
[0094] Step 2.4, input the parsed entity cluster into the pattern reasoning module, and construct a hierarchical legal knowledge classification system through the pattern reasoning module, which specifically includes:
[0095] Hypernym generation: Generate common hypernyms for entity type clustering;
[0096] Hierarchical aggregation: Merge hypernyms to form a multi-level taxonomy until a single top-level classification is generated.
[0097] The method for automatically constructing based on a legal large language model and a legal knowledge graph according to claim 4, characterized in that, in step 3, the dynamic few-shot retrieval includes;
[0098] Step 3.1, using a sentence encoder to map the question, entity, and relationship into vector representations. Specifically, using the sentence encoder to map the input question q and the entities and relationships into vector representations ;
[0099] Step 3.2, calculating the matching degree with the knowledge graph examples through formula G1;
[0100] Step 3.3, taking the k examples with the highest matching degree as context prompts and inputting them into the legal large language model to generate a SPARQL query;
[0101] Further, in the above solution, the formula G1 is: , where is a similarity function. Based on this score, k most similar examples S are retrieved; using the sentence encoder to map the input question q and the entities and relationships into vector representations , by connecting the question, entity, and relationship into a single input sequence , and processing it by a legal language model containing only an encoder. Each example is encoded as a vector , then, the module calculates the similarity between the input question vector and the example vector to retrieve the most relevant examples from the storage.
[0102] The intelligent question-answering system of the present invention mainly consists of three core modules: a legal large language model fine-tuning module, a knowledge graph automatic construction module, and an intelligent question-answering module;
[0103] First, the system fine-tunes the Chinese language model using a legal dataset through the legal large language model fine-tuning module to make the Chinese language model a legal language model, and then generates a legal knowledge graph through the legal language model in the knowledge graph automatic construction module; in the intelligent question-answering module, the system processes the current legal question through the legal knowledge graph and the legal language model, takes the legal answers trusted by the user as a reward, and inputs it into the dynamic few-shot retrieval. Then, through operations such as the legal language model, context prompting, multi-question generation, and question-answer selection, a legal prediction answer is generated. In the application module of the intelligent question-answering system, users can input legal questions to be consulted, and the system will reason and calculate the questions based on the trained model and the generated legal knowledge graph to generate a prediction answer.
[0104] The fine-tuning logic structure of the legal large language model is as Figure 2 shown, and the main components include an embedding module, a pre-trained large model, a prompter module, a LoRA fine-tuning module, and a Chinese large language model. The fine-tuning process of the Chinese large language model is divided into the following three steps:
[0105] Step 1: First, the legal question and the target answer are converted into word vectors through the embedding module. The embedding module consists of three parts: a phrase embedding layer, a phrase type embedding layer, and a position embedding layer. The phrase embedding layer converts each phrase into a corresponding word vector and adds special tokens such as [CLS] and [SEP]. The encoding method used here is Word2Vec. Word embedding not only retains the information of the original text vocabulary but also preserves the semantic relevance between the vocabulary, which can enhance the embedding vector before inputting it into the pre-trained large model; the phrase type embedding layer embeds specific types of word vectors (such as open-ended questions, closed-ended questions, or questions of specific organ types) into the overall word vector; the position embedding layer integrates the position information of each word vector into the overall word vector. Subsequently, through the processing of the normalization layer and the linear layer, the alignment of the word vectors is achieved. Finally, these word vectors are input into the pre-trained large model Roberta. The pre-trained large model consists of several encoders, which are mainly responsible for aligning the text and the corresponding answers, and then the feature vectors are input into the prompter. The input and output of the prompter are repeated three times. The design refers to the Transformer architecture and applies the self-attention module, cross-attention module, residual connection, and forward propagation layer. The prompter further aligns the information between the question and the answer and optimizes the ability of the model to generate reliable answers.
[0106] Step 2: Input the prompt vector generated by the prompter into the Chinese large language model. The Chinese base large language model uses the ChatGLM large language model, which is the second generation version based on the open-source Chinese-English bilingual dialogue model ChatGLM-6B. Input the prompt vector containing question information into the ChatGLM large language model to further align the features of images and texts.
[0107] Step 3: Use the cross-entropy loss function to calculate the difference between the target answer and the predicted answer, and use this loss function to update the parameters of the low-rank matrix in the Chinese large language model by means of LoRA fine-tuning. The specific approach is to add a bypass structure in the Chinese large language model. This bypass structure is obtained by multiplying two matrices. During training, freeze the parameters of the Chinese large language model and only update the parameters of the low-rank matrix that multiplies in the bypass structure. Through continuous training, the cross-entropy loss gradually decreases, and the overall performance of the model gradually improves, eventually converging to the optimal state. After reaching the optimal state, directly merge the parameters of the low-rank matrix and the parameters of the Chinese large language model to obtain the fine-tuned legal large language model.
[0108] This invention relates to the automatic construction of a legal knowledge graph. As Figure 3 shown, first use legal data as the input of the system to let the system automatically generate a legal knowledge graph, and then combine the knowledge graph and the legal language model. Input legal questions into the system, and the system automatically generates a predicted answer.
[0109] The detailed process of the legal knowledge graph generation paradigm is as Figure 3 shown, mainly including candidate triple extraction, entity / predicate parsing, and pattern reasoning; specifically:
[0110] Extract candidate triples. Exploratory stage: Iteratively perform task decomposition and prompt definition through trial and error until each subtask can be executed reliably and accurately. Tuple extraction process design: Preprocessing stage: The text is first split into smaller chunks so that the legal language model can process these inputs more effectively. Two-stage extraction process: 1. Entity extraction: First, identify the entities in the text chunks to ensure that each entity can be clearly extracted. 2. Triple extraction: Then, check the relationships between these entities through an iterative method to generate the final triples.
[0111] Furthermore, the text splitting module: When extracting entities and triples, due to token limitations in the legal large language model, the text splitting module is used to process long texts. For example, the token limit of GPT-3.5 is 4096, including input and output tokens. To process long texts without exceeding this limit, the text splitting module divides the input into smaller chunks so that the model can process the text in several steps. Its function is to efficiently extract entities and triples from long documents while preventing token overflow during the processing.
[0112] Preprocessing legal text data through the entity extraction module: The entity extraction module is responsible for identifying and extracting entities from text chunks. This module processes features by using detailed system prompts to guide the legal large language model to complete the extraction process. It consists of three specific parts: 1) a clear explanation of entity composition, 2) direct instructions for retrieving entity mentions and a list of descriptions and types for each entity, 3) detailed formatting information for the output. 2) and 3) define the task and format, and 1) provides a more targeted entity definition. The function of the entity extraction module is to improve the accuracy and relevance of entity extraction. When including 1), it narrows down the scope to specific nouns and named entities, resulting in fewer but more detailed extractions while excluding abstract nouns.
[0113] Phrase selection: The phrase selection module is used to extract key text excerpts from the full text T , and this module integrates information related to specific entities . This module processes features by isolating smaller relevant parts of the text, thereby reducing complexity and improving the accuracy of triple extraction using the legal large language model. It generates based on the design prompt of a simple declarative sentence centered on by performing classic query-centered text summarization or open-ended question-answering tasks. This module can integrate through directly using the descriptions generated by the entity extraction module for explicit extraction.
[0114] Mention recognition: The mention recognition module is used to identify entities related to the target entity . This module processes features by obtaining the generated summary (focusing on ) and detecting mentions of other entities from the predefined list E. The function of this module is to create a subset of entities that only contains entities mentioned in . This task is performed by issuing a system prompt to the legal language model to identify Entity mentions in it. The working principle of this module is to provide a numbered entity list from a predefined list E and the text to a legal large language model and ask it to append "yes" or "no" after each entity in the list to indicate whether it is mentioned in the text. Using the numbered list helps to preserve the correct association between entity IDs and labels, thus improving the reliability of the results. The final subset is formed by selecting the entities marked as "yes".
[0115] Relationship extraction module: The relationship extraction module is used to identify the relationships between the context entities generated in the previous steps. This module processes features by obtaining the summary text and the entity list and using a system prompt to query the legal language model to extract relationships in the form of RDF triples. The triples consist of a subject and an object (with name and ID) from and an expressive predicate describing the relationship between them. The function of this module is to generate a set of triples that accurately represent the relationships between entities in . The system prompt includes a detailed explanation of the meaning of the expressive predicate, guiding the legal language model to select predicates that are diverse enough to be reused in multiple triples, thus improving the quality and reusability of the extracted relationships. This normalization of predicates ensures that relationships are well represented without being overly specific.
[0116] Predicate description: The predicate description module is used to generate an explanation for each unique predicate. Since this process involves a certain degree of complexity, it is implemented as a separate final step to simplify the relationship extraction process. The process starts with a system prompt that asks the legal language model to generate a description for each unique predicate. This module references the text provided by the user and the list of RDF triples (R i ), and these triples have clear divisions and markings. Functionally, this module is a key component for understanding and describing the relationships between different elements in the data, providing clarity and making the data easier to understand.
[0117] Input candidate triples into the entity / predicate parsing module: The entity / predicate parsing module combines semantic aggregation and prompts from the legal language model. The purpose of semantic aggregation is to group related entities or relationships that represent similar semantic meanings or correspond to higher-level concepts. For example, the concepts of "car" and "motorcycle" will be grouped into the higher-level concept of "vehicle". This initial aggregation process allows the information to be divided and passed in the form of a series of prompts rather than a single prompt. The parsing process within the module is as follows: The entity / predicate parsing is detailed as follows:
[0118] Semantic Aggregation Module: The first stage of semantic aggregation involves collating entities and relationships with semantic similarity. This process designs a function to calculate two specific similarity scores - one for entities (S e ), and another for relationships (S r ). These scores are determined by evaluating the output lengths of knowledge graph components, including entity / relationship, description, and type. First, entity and relationship pairs are considered. For each pair (i, j), the module calculates the similarity score for each entity or relationship pair by determining the Levenshtein distance, which represents the similarity between two entity labels. For entities, this score is denoted as e i,j , and for relationships, this score is denoted as r i,j . These scores are then normalized to the range of 0 to 1, with a score of 1 indicating the same entity. The module also uses the same strategy to calculate the similarity between each pair of types, which only applies to entities. Considering the complexity of descriptions, the module uses a more sophisticated method to evaluate descriptions. Entities and relationships are projected into an embedding space, and the similarity between entities and relationships is determined by the cosine similarity metric. A legal language model is used as the embedding model. The final similarity scores for relationships and entities are obtained through weighted fusion, as shown in the given formula:
[0119] (1)
[0120] (2)
[0121] The fixed coefficient values are = 0.35,[[]] = 0.65,[[]] = 0.25, and = 0.75. The module outputs a set of aggregations of entities and relationships. Each aggregation is a set of semantically similar entities or relationships, which can also be said that each group represents an entity / relationship cluster. Each cluster is integrated into the output sent to the legal language model.
[0122] Clustering Disambiguation Module: This module parses a group consisting of semantically related elements, which may not all correspond to a single concept, consistent with the results of semantic aggregation. For example, an entity cluster containing various vehicles, these entities have similarities, but they do not represent the same entity. The role of this module is to determine which subsets in the cluster correspond to the same entity. The legal language model is given a specific prompt to require it to identify subsets of semantically equivalent entities or relationships. This step is repeated in all clusters, and the final output result is a set of semantically identical entities or relationships. Therefore, the function of this module is to process features in a way that enables the legal language model to distinguish and group the same entities or relationships within semantically linked clusters.
[0123] Concept Contraction Module: The concept contraction process aims to identify and unify mentions of the same entity, providing a single, unique identifier for that entity. During the process, relationships that may have different textual representations but express the same relationship are identified and these relationships are marked with unique identifiers. The Cluster Disambiguation Module helps achieve this by using each set of equivalent entities or relationships to create a new prompt. The purpose of this new prompt is to request a unique label that can be used to represent the group of related entities or relationships. This label ultimately serves as the final unique representation of the specified entity or relationship.
[0124] Input the parsed entity clusters into the Pattern Inference Module: This module processes the clusters of parsed entities and aims to infer their corresponding types. Pattern inference is carried out through an iterative process, with each iteration involving two core steps. The function of this module is to gradually refine and discover appropriate pattern structures, helping to automate the pattern generation process and reduce manual effort. In summary, this module uses the clusters of parsed entities to infer and organize their types, and automatically iteratively creates patterns using a method based on a legal language model.
[0125] Hypernym Generation: This module processes clusters of entity types to identify appropriate hypernyms. It first removes duplicates within each cluster, ensuring that entities belonging to the same type do not repeat. Then the remaining types are embedded into a prompt, which is transmitted to the legal language model. The task of this module is to find a common hypernym applicable to the entire cluster and determine the relationship connecting the hypernym to the entity types. Based on the cluster size and the semantic similarity between types, the module generates multiple hypernyms, each associated with a specific subset of the cluster. This module processes the entity clusters by embedding them into hypernym prompts generated by the legal language model, thereby identifying the hierarchical relationships of the refined architecture.
[0126] Hierarchical Agglomeration: Hierarchical agglomeration involves merging the hypernyms and relationships generated in all clusters to eliminate redundancy. The system applies semantic aggregation techniques to group these higher-level hypernyms into new clusters. Next, the hypernym generation and hierarchical agglomeration processes are iteratively applied to construct a higher-level taxonomy. This module processes the patterns by merging hypernyms and relationships, organizing them into a hierarchical structure. It uses semantic aggregation, hypernym generation, and hierarchical agglomeration to iteratively refine the upper layer of the taxonomy until a single top-level hypernym is determined.
[0127] Dynamic Few-Shot Retrieval: Use a sentence encoder to map the input question q and its extracted entities and relationships to vector representations . By concatenating the question, entities, and relationships into a single input sequence , it is processed by a legal language model consisting only of an encoder. Each example is encoded as a vector Then, the module calculates the input question vector and the example vector to retrieve the most relevant examples from the store. This module uses a sentence encoder to process the input features (question, entity, and relation), mapping them into a vector representation. Its function is to dynamically retrieve the most relevant examples from the stored set based on similarity, which helps improve the accuracy of SPARQL query generation:
[0128] : (4)
[0129] where is the similarity function. Based on this score, the k most similar examples S are retrieved.
[0130] Context hint: The legal language model processes the entire prompt, generates a SPARQL query as a response, and encloses the result in ` <sparql>< / sparql> ` tags. Then, the system parses the generated query and the SPARQL engine executes the query on the knowledge graph to retrieve the answer to the input question. This module processes the features by constructing a single prompt from the task description, retrieved demonstrations, and the input question along with their respective entities and relations. Its function is to guide the legal language model to generate an effective SPARQL query to answer the input question.
[0131] Multi-query generation: This module processes the features by generating multiple SPARQL queries (instead of a single query) through beam search to capture different possible subject-object permutations. Its function is to reduce triple flip errors and improve query accuracy by considering multiple hypotheses in the final output. This module generates multiple SPARQL queries instead of a single query by retaining all the hypotheses generated during beam search, where multiple possible sequences (or query hypotheses) are explored. The model does not return only the most likely SPARQL query but outputs a set of queries , where is the number of final hypotheses generated during beam search.
[0132] Answer selection: This module processes the results of multiple SPARQL queries and applies the largest set (LS) or first set (FS) heuristic to select the most appropriate answer. Its function is to ensure that the most reliable result is returned from the set of possible answers by selecting the largest answer set or the first valid answer set. The largest set executes all queries and for each query obtains a (possibly empty) answer set . Then the largest set Select the largest one among them, i.e.:
[0133] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An intelligent question-answering system automatically constructed based on a legal language model and legal knowledge graph, characterized by: It includes the legal language model fine-tuning module, the knowledge graph automatic construction module and the intelligent question-answering module; The legal large language model fine-tuning module includes an embedding module, a pre-trained large model, a prompter module and a Lora fine-tuning module; the knowledge graph automatic construction module includes a candidate triple extraction module, an entity / predicate parsing module and a pattern reasoning module; the intelligent question and answer module includes a dynamic few-sample retrieval module, a multi-query generation module and an answer selection module.
2. The intelligent question-answering system automatically constructed based on the legal big language model and legal knowledge graph according to claim 1 is characterized in that: The embedding module includes a phrase embedding layer, a phrase type embedding layer and a position embedding layer; the Chinese base language model of the legal large language model fine-tuning module is the ChatGLM-6B second-generation model.
3. The intelligent question-answering system automatically constructed based on the legal big language model and legal knowledge graph according to claim 1 is characterized in that: The candidate triple extraction module includes a text splitting module, an entity extraction module and a relationship extraction module; the entity / predicate parsing module includes a semantic aggregation module, a clustering disambiguation module and a concept contraction module; the pattern reasoning module includes a hypernym generation module and a hierarchical clustering module.
4. A method for automatically constructing a legal large language model and a legal knowledge graph, applied to the intelligent question-answering system of claims 1 to 3, characterized in that: The steps include: Step 1: Fine-tune the legal language model: Input the legal question-answer dataset, and convert the questions and answers in the legal question-answer dataset into feature vectors through the embedding module; The feature vector is optimized using the prompter module to generate a prompt vector containing legal semantics; Input the prompt vector into the Chinese large language model and calculate the cross entropy loss between the predicted answer and the target answer; Freeze the parameters of the Chinese base language model and update the bypass parameters of the low-rank matrix in the Chinese large language model through the Lora fine-tuning module until the model converges; After the model converges, the parameters of the low-rank matrix are merged with the parameters of the Chinese large language model to obtain the fine-tuned legal large language model. Step 2: Automatic construction of legal knowledge graph: Preprocess legal text data and generate candidate triples through text segmentation, entity extraction, and relationship extraction; The candidate triples are input into the entity / predicate parsing module, and semantic aggregation, cluster disambiguation and concept contraction are performed on the candidate triples to generate parsed entity clusters; The parsed entity clusters are input into the pattern reasoning module, and the pattern reasoning module finally outputs a knowledge graph containing legal knowledge; Step 3; Intelligent question-answering processing; Receive legal questions input by users and extract entities and relations in the questions; Through the dynamic few-shot retrieval module, the cosine similarity between the question and the examples stored in the knowledge graph is calculated, and the top k most relevant examples are retrieved; Generate multiple SPARQL query candidates by combining the retrieval examples with the multi-query generation module, wherein the generation process adopts a beam search strategy to retain the query sequence with the highest top-n probability; The maximum set (LS) heuristic method is executed by the answer picking module to select the answer that returns the largest number of results as the final output.
5. The method for automatically constructing a legal large language model and a legal knowledge graph according to claim 4 is characterized in that: In step 1, the legal question-answer pair dataset is obtained by pre-training with the LawGPT_zh large model, and the embedding module converts the legal question-answer pair dataset into a feature vector. The specific steps are as follows: Step 1.1, the phrase embedding layer converts each phrase into a corresponding word vector and adds special tags; Step 1.2, the phrase type embedding layer embeds the word vector of a specific type into the overall word vector; Step 1.3: The position embedding layer integrates the position information of each word vector into the overall word vector, and then aligns the word vectors after being processed by the normalization layer and the linear layer. In step 1.4, the word vector is input into the pre-trained large model Roberta to align the text and the corresponding answer to form a special vector.
6. The method for automatically constructing a legal large language model and a legal knowledge graph according to claim 4 is characterized in that: In step 1, the feature vector is optimized using the prompter module, which includes: Step 1.5, input the feature vector into the prompter, the input and output of the prompter are repeated three times, and the feature vector is input into the prompter. The prompter is based on the Transformer architecture and generates a legal semantic prompt vector through a self-attention mechanism to align the semantic association between the question and the answer; Step 1.6, input the prompt vector into the Chinese language model, use the cross entropy loss function to calculate the difference between the target answer and the predicted answer, and use the loss function to update the parameters of the low-rank matrix in the Chinese language model through Lora fine-tuning; Step 1.7, through continuous training, the cross entropy loss is gradually reduced until the model converges. After the model converges, the parameters of the low-rank matrix and the parameters of the Chinese base language model are merged to obtain a fine-tuned legal large language model.
7. The method for automatically constructing a legal large language model and a legal knowledge graph according to claim 6 is characterized in that: The step 1.6 specifically includes adding a bypass structure in the Chinese language model, which is obtained by multiplying two matrices. During training, the parameters of the Chinese language model are frozen and only the parameters of the low-rank matrix multiplied in the bypass structure are updated. Through continuous training, the cross entropy loss is gradually reduced, the overall performance of the model is gradually improved, and finally converges to the optimal state.
8. The method for automatically constructing a legal large language model and a legal knowledge graph according to claim 4 is characterized in that: In step 2, the steps of generating candidate triples specifically include: Step 2.1, pre-process the legal text data and split the long text into short text blocks that do not exceed the language model token limit through the text splitting module; Step 2.2, perform two-stage extraction on each text block, namely: Entity extraction: Use the entity extraction module to identify entities related to the target entity, use the legal large language model to identify entities and their types, and generate a list of candidate entities; Relation extraction: Use the relationship extraction module to identify the relationship between the context entities generated in the previous step, generate RDF triples based on the semantic associations between entities, and normalize the predicates; Step 2.3, input the candidate triples into the entity / predicate resolution module, perform semantic aggregation and clustering disambiguation on the candidate triples, unify the entity / predicate identifiers through concept contraction, and generate resolved entity clusters; Step 2.4, input the parsed entity clusters into the pattern reasoning module, and construct a hierarchical legal knowledge classification system through the pattern reasoning module, which specifically includes: Hypernym generation: Generate common hypernyms by clustering entity types; Hierarchical aggregation: Hypernyms are combined to form a multi-level taxonomy until a single top-level taxonomy is generated.
9. The method for automatically constructing a legal large language model and a legal knowledge graph according to claim 4 is characterized in that: In step 3, the dynamic few-sample retrieval includes: Step 3.1: Use sentence encoder to map questions, entities and relations into vector representations. Specifically, use sentence encoder to map the input question q and its extracted entities and relationship Mapping to vector representation ; Step 3.2, calculate the matching degree with the knowledge graph example through formula G1; In step 3.3, the k examples with the highest matching degree are used as contextual prompts and input into the legal language model to generate SPARQL queries.
10. The method for automatically constructing a legal large language model and a legal knowledge graph according to claim 9 is characterized in that: The formula G1 is: , among which is a similarity function, based on this score, the k most similar examples S are retrieved; the sentence encoder is used to convert the input question q and its extracted entities and relationship Mapping to vector representation , by concatenating questions, entities, and relations into a single input sequence , processed by a legal language model containing only an encoder, each example Encoded as a vector , then the module calculates the input problem vector and the example vector to retrieve the most relevant examples from storage.
Citation Information
Cited By
Intelligent question answering method and device based on domain knowledge lightweight fine tuning and cross-domain dynamic knowledge base and readable storage medium thereof
CN120296110A
Intelligent legal instrument automatic generation system
CN120850987A
Intelligent auditing method and system based on multi-modal large model
CN120876132A