Case Q&A Method, Medium and Device Based on Large Language Model
By building a multimodal corpus and fine-tuning of the model based on transfer learning, combining the thinking chain of CLIP module and multi-dimensional logical association, the problem of inaccurate and insufficient answers in case intelligent Q&A is solved, and more efficient and reliable Q&A performance is achieved.
Patent Information
- Application Number
- CN202510198989.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-24
AI Technical Summary
The existing intelligent question-and-answer technology in case does not perform well in generating accurate and reliable answers, and due to the particularity of the case question-and-answer task, the logic and accuracy of the intelligent question-and-answer are not sufficient.
By building a multimodal corpus, training pre-trained models based on transfer learning, combining LoRA and P-Tuning fine-tuning large language models, introducing CLIP modules for multimodal feature fusion, and introducing a multi-dimensional logical association thinking chain in the generative model. Finally, training and evaluation models based on comparative learning to improve the accuracy and reliability of question-and-answer.
It significantly improves the performance of question-and-answer technology in generating accurate and reliable answers, reduces the generation of misinformation or hallucinations, enhances the model's ability to understand the specific terms and situations of the case, and improves reasoning and interpretability.
Smart Images

Figure CN119692484B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent question - answering methods, and particularly to a case question - answering method, medium and device based on a large language model. Background Art
[0002] With the rise of artificial intelligence (LegalAI), all industries are undergoing significant transformations. It benefits various groups by automating tasks and answering questions. It improves people's work efficiency by reducing the heavy burden of paper work and simplifies the access process of the general public to daily life services, consultation services and remote case consultations.
[0003] However, the application of these technologies also faces challenges such as answer hallucination and insufficient reasoning ability of the model. Therefore, solving the problem of hallucination in intelligent judgment and improving the reasoning ability of the model have become the current research hotspots.
[0004] Combined with the research and collation of current question - answering technologies and large model technologies in the specified application fields, it can be seen that there is less labeled corpus for existing cases, and due to the particularity of the case question - answering task, some methods of intelligent question - answering cannot be directly transferred to this field. Although large language models have achieved excellent results in various tasks, using large language models to solve case question - answering problems is currently a research hotspot, but there are still the following problems:
[0005] 1) Although large language models have shown excellent results in many NLP tasks, the commonly used fine - tuning methods still have unsatisfactory effects in the case intelligent question - answering task.
[0006] 2) Although large language models have good intelligent question - answering effects in the general field, due to the particularity of the case question - answering task, the logic and accuracy of intelligent question - answering are still insufficient. Summary of the Invention
[0007] The purpose of the present invention is to provide a case question - answering method, medium and device based on a large language model that can improve the performance of the question - answering technology in generating accurate and reliable answers and reduce the generation of false information or hallucinations. The specific technical solutions are as follows:
[0008] The present invention provides a case question - answering method based on a large language model, including the following steps:
[0009] Obtain the graphic and text information of the case and the related questions of the case;
[0010] Input the obtained graphic and text information of the case and the related questions of the case into an intelligent question - answering model to obtain the optimal answer. The intelligent question - answering model is obtained through the following steps:
[0011] S1: Construct a multi-modal corpus containing case-related knowledge, and train a pre-trained model through the multi-modal corpus based on transfer learning to obtain a base model suitable for cases, where: the data format of the multi-modal corpus includes text and pictures;
[0012] S2: Fine-tune the parameter information of the base model based on LoRA and P-Tuning to obtain a fine-tuned large language model;
[0013] S3: Introduce a CLIP module on the basis of the fine-tuned large language model to perform multi-modal feature fusion on the text and picture information of the input case, and obtain a preliminary question-answering model that can understand and fuse text and picture information;
[0014] S4: Introduce a multi-dimensional logical association thinking chain into the preliminary question-answering model to obtain a generation model; the generation model outputs multiple candidate answers with an inference chain according to the relevant questions of the input case;
[0015] S5: Train the existing scoring model based on the contrast learning method and the candidate answers generated by the generation model to obtain an evaluation model that can output the optimal answer among multiple candidate answers.
[0016] Optionally, the S1 includes:
[0017] S1.1: Collect multi-source data in the case-related field, and preprocess the collected multi-source data in the case-related field to obtain a multi-modal corpus containing case-related knowledge, where: the text data of case-related knowledge includes case-related regulations and case judgment information; the preprocessing includes data cleaning of the multi-source data in the case-related field, and annotation of the unannotated data in the multi-source data in the case-related field;
[0018] S1.2: Select a language model pre-trained on a large-scale general corpus as the pre-trained model, and initialize the pre-trained model;
[0019] S1.3: Extract the annotated data set in the case-related field from the multi-modal corpus, and continue to train the pre-trained model through the annotated data set in the case-related field to obtain a fine-tuned model. The specific formula is as follows:
[0020] ;
[0021] Where, is the loss function on the target task; is the pre-trained model; is the annotated data set in the case-related field, is the fine-tuned model;
[0022] S1.4. Adapt the fine-tuned model to a specific case Q&A task to obtain an adapted model, where the case Q&A task includes question understanding, answer generation, and information extraction;
[0023] S1.5. Optimize the adapted model to obtain a base model applicable to cases.
[0024] Optionally, S1.5 includes:
[0025] Perform multi-task learning on the adapted model to obtain an optimized model capable of performing multiple tasks simultaneously;
[0026] Introduce a regularization term and a learning rate scheduling strategy into the optimized model to obtain a regularized model;
[0027] Introduce a knowledge graph or a multi-modal corpus into the regularized model to obtain a base model applicable to cases.
[0028] Optionally, S2 includes:
[0029] Optimize the weight matrix of specific layers in the base model by introducing a low-rank matrix based on the LoRA fine-tuning method to obtain the base model after the first fine-tuning;
[0030] Optimize the prompt for the case Q&A task in the input end of the base model after the first fine-tuning based on the P-Tuning fine-tuning method to obtain a fine-tuned large language model.
[0031] Optionally, S3 includes:
[0032] Use the CLIP model to encode the text and image information of the input case respectively to generate an embedding vector of the image and an embedding vector of the text;
[0033] Embed the image and text into a shared semantic space for contrast learning and fusion to obtain a preliminary Q&A model capable of understanding and fusing text and image information.
[0034] Optionally, S4 includes:
[0035] S4.1. Construct a thinking chain with multi-dimensional logical associations based on syllogism and analogical reasoning; specifically:
[0036] Take the relevant regulations as the major premise and the text and image information of the case as the minor premise and input them into the preliminary Q&A model;
[0037] Judge whether the text and image information of the input case conforms to the relevant regulations under the major premise. If it conforms, the preliminary Q&A model outputs a preliminary conclusion. If not, search for historical cases and judgments similar to the current case and analogically derive a similar conclusion, where the major premise, the minor premise, and the reasoning conclusion based on the major premise and the minor premise are the three elements of syllogism;
[0038] Combine the inference conclusions of syllogism and analogical reasoning to generate a complete inference chain and the answer to the case consultation;
[0039] S4.2. Introduce the thinking chain of multi-dimensional logical association into the preliminary question-answering model to obtain a generation model.
[0040] Optionally, the above S5 includes:
[0041] S5.1. Construct a contrastive learning loss function and a ranking loss function;
[0042] S5.2. Based on the contrastive learning loss function and the ranking loss function, train the existing scoring model to learn to distinguish high-quality answers from low-quality answers among multiple candidate answers and rank them, so as to obtain an evaluation model that can output the optimal answer among multiple candidate answers, where: high-quality answers are positive samples, and low-quality or irrelevant answers are negative samples;
[0043] Specifically:
[0044] S5.2.1. Input the relevant questions and candidate answers of the case: Input the relevant questions and multiple candidate answers into the existing scoring model;
[0045] S5.2.2. Calculate sentence embeddings: Use a pre-trained language model to generate sentence embeddings of the questions and candidate answers;
[0046] S5.2.3. Construct a total loss function: Construct a total loss function composed of a contrastive learning loss function and a ranking loss function;
[0047] S5.2.4. Calculate the loss: Evaluate the quality of each candidate answer by calculating the total loss function and rank them according to the relevance between the candidate answers and the questions;
[0048] S5.2.5. Gradient update: Use the gradient descent algorithm to optimize the total loss function through backpropagation to update the weights of the existing scoring model;
[0049] S5.2.6. Iterative training: Repeat S5.2.4 to S5.2.5 until the total loss function of the existing scoring model converges to obtain an evaluation model that can output the optimal answer among multiple candidate answers.
[0050] Optionally, the contrastive learning loss function is a classification objective function or a triplet objective function, and the classification objective function The specific formula is as follows:
[0051] ;
[0052] Where: is the time step for training the existing scoring model; is a trainable weight matrix, ; is the dimension of the sentence embedding, is the number of label categories in the classification objective function; 、 are two different sentence embeddings, is 、 the distance between;
[0053] The specific formula of the triple objective function is as follows:
[0054] ;
[0055] Among them: is the sentence embedding of the given anchor sentence , is the sentence embedding of the positive sentence , is the sentence embedding of the negative sentence ; is the margin;
[0056] The specific formula of the ranking loss function is as follows:
[0057] ;
[0058] Among them: is the candidate set composed of candidate answers, and are the th candidate answer and the th candidate answer respectively, 、 are the labels of the candidate answers respectively, ; represents the similarity between the candidate answer and the question ; is the similarity between the candidate answer and the question , 、 are the labels of the question respectively, .
[0059] The present invention also provides a readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above-mentioned case question-answering method based on a large language model is implemented.
[0060] The present invention also provides an electronic device, including: at least one processor, at least one memory, and computer program instructions stored in the memory. When the computer program instructions are executed by the processor, the case Q&A method based on the large language model is as described above.
[0061] The base model constructed by the present invention based on transfer learning allows the model to learn from a wide range of data sources and apply it to the professional field of cases, which can effectively enhance the understanding ability of the base model for case-specific terms and scenarios;
[0062] Through the fine-tuning method combining LoRA and P-Tuning, the knowledge of the large language model is maximally transferred to the target task, reducing the dependence on a large amount of domain-specific labeled data, and realizing the rapid adaptation and deployment of the model; and using the parameter information of the base model as the initialization parameters of the new model to be trained in the target domain, and then updating some parameters of the model through the labeled corpus in the target domain to obtain the fine-tuned large language model, which can accelerate and optimize the learning efficiency of the fine-tuned large language model, without making the model learn from scratch, greatly reducing the training cost;
[0063] Through training, the CLIP model can map images and texts in the multimodal corpus to a shared semantic space for multimodal feature fusion, obtaining a preliminary Q&A model. Thus, when the model conducts intelligent Q&A, it can understand and combine image and text information to provide more accurate and comprehensive answers;
[0064] Moreover, the present invention also introduces a thinking chain with multi-dimensional logical associations on the basis of the preliminary Q&A model to obtain a generation model, so that the generation model not only depends on the matching of surface facts, but also can make more reasonable judgments at multiple levels through reasoning ability;
[0065] Finally, by training the existing scoring model, an evaluation model capable of outputting the optimal candidate answer is obtained to reduce the hallucination during model Q&A.
[0066] In addition to the purposes, features and advantages described above, the present invention has other purposes, features and advantages. The present invention will be further described in detail below with reference to the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] The drawings constituting a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0068] Figure 1 It is a schematic flowchart of S2 to S3 in an embodiment of the present invention;
[0069] Figure 2This is the specific technical roadmap of S3 in the embodiments of the present invention;
[0070] Figure 3 This is the specific technical roadmap of S4 in the embodiments of the present invention. Specific embodiments
[0071] The following will describe the embodiments of the present invention in detail with reference to the accompanying drawings. However, the present invention can be implemented in many different ways as defined and covered by the claims.
[0072] In one embodiment, referring to Figures 1 to 2 , a case Q&A method based on a large language model includes the following steps:
[0073] 1. Obtain the graphic and text information of the case and the related questions of the case;
[0074] 2. Input the obtained graphic and text information of the case and the related questions of the case into the intelligent Q&A model to obtain the optimal answer;
[0075] The intelligent Q&A model is obtained through the following steps:
[0076] S1: Construct a multi-modal corpus containing case-related knowledge, and based on transfer learning, train a pre-trained model through the multi-modal corpus to obtain a base model applicable to cases, where: the data format of the multi-modal corpus includes text and pictures;
[0077] In some specific emerging fields, if one wants to build a model to handle Q&A tasks, it is usually very difficult to obtain enough labeled data to support model training. Therefore, the top priority is to build a reliable model that only requires a small amount of labeled training data. Transfer learning is an ideal solution to solve this problem. It uses existing knowledge to solve problems in the target domain that are different but related to the source domain. For the case Q&A task in the case-related field, continue to pre-train a large language model that has been trained on a large-scale corpus, transfer some language features of the source domain to the target domain (such as the case-related field) to improve the performance of the model, and then adapt the obtained pre-trained model to the case Q&A model in the target domain, maximizing the reuse of existing knowledge, without having to invest a large amount of cost to re-collect the labeled corpus in the target domain, and realizing the rapid transfer and application of the model.
[0078] S1 includes:
[0079] S1.1. Collect multi-source data in the case-related field, and preprocess the collected multi-source data in the case-related field to obtain a multi-modal corpus containing case-related knowledge.
[0080] Among them: the text data of case-related knowledge includes case-related regulations and case judgment information;
[0081] If the labeled data is scarce, unlabeled corpus can be preferentially used, which can help the model learn the language features in the case-related fields. Through self-supervised learning, these unlabeled data can be used to optimize the pre-training stage.
[0082] Preprocessing includes data cleaning of multi-source data in the case-related fields and labeling the unlabeled data in the multi-source data of the case-related fields. Data cleaning is a key step to ensure the quality of model training, removing noise and irrelevant information to guarantee data quality; the labeling process can be achieved through expert review or semi-automated tools, especially for accurate labeling of specific terms and case scenarios in the case-related fields. To alleviate the problem of scarce labeled data, data augmentation techniques can be used, such as expanding the diversity of the dataset through text generation models.
[0083] S1.2. Select a language model pre-trained on a large-scale general corpus as the pre-training model and initialize the pre-training model;
[0084] S1.3. Extract the labeled dataset in the case-related fields from the multi-modal corpus, and continue to train the pre-training model with the labeled dataset in the case-related fields to obtain a fine-tuned model. The specific formula is as follows:
[0085] ;
[0086] where is the loss function on the target task; is the pre-training model; is the labeled dataset in the case-related fields, is the fine-tuned model;
[0087] Transfer learning extracts language knowledge from a large-scale multi-modal corpus and then further trains the model with the labeled data unique to the case-related fields, enabling the model to understand the case-specific terms and scenarios. Specifically, by fine-tuning on the data in the case-related fields, the model can master the terms, relevant regulations, crime patterns, and case reasoning processes in this field. The fine-tuning stage is the core step of transfer learning. In this stage, the pre-trained language model is further trained on the specific corpus in the target domain (such as cases), enabling the model to adjust the language features of the source domain and learn the knowledge of specific professional terms, case types, behavior patterns, etc. in the case-related fields. The main task of the fine-tuning stage is to adjust the parameters of the model to make it adapt to the specific needs of the target domain. Due to the high professionalism of the knowledge in the case-related fields, the model needs to optimize the parameters through gradient descent on the limited labeled data to maximize its performance in this field.
[0088] S1.4. Adapt the fine-tuned model to a specific case Q&A task to obtain an adapted model, where: the case Q&A task includes question understanding, answer generation, and information extraction;
[0089] The target task usually includes various task forms such as classification and sequence labeling, and the choice of loss function depends on the task type. In this embodiment, the goal of case analysis is a classification task, and the loss function uses cross-entropy loss to define:
[0090] ;
[0091] where, is the label of the sample in the labeled dataset in the case-related field of is the class probability of the model output sample of is the number of samples in the labeled dataset in the case-related field. This loss function is optimized during training so that the class probability output by the model is as close as possible to the true label.
[0092] The label refers to the class associated with each sentence or sentence pair, which may be a question class or an answer classification in a Q&A system; in this embodiment, when doing a "question-answer matching" task, the label may be "match" or "not match". In other types of tasks, the label may represent different classes, such as case classes, question types, etc.
[0093] The label is used to classify or distinguish different types of questions or answers by the model through a contrastive learning task. In another example of this embodiment, the task of the model is to classify according to the similarity between the description of a criminal case and a question pair, then the label may correspond to different types of cases or case classes.
[0094] In summary, the label is related to the semantic matching of questions and answers in this context, and it determines whether a sentence pair belongs to the same class (such as similar case descriptions and question types).
[0095] In a classification task, the true label refers to the class or label that the sample actually belongs to in the training data. In the context of case analysis, the true label refers to the class that each training sample (i.e., the description or related data of the case) actually belongs to. In this embodiment, cases may be classified according to different types. For each training sample, the true label represents the actual type of the case.
[0096] The questions and candidate answers involve the analysis of cases. The relationship between the true label and the questions and candidate answers may be reflected in the following aspects:
[0097] Problem: The problem can be a form of case analysis task, such as "What type of case does this case belong to?" or "What is the nature of this case?".
[0098] Candidate answers: The candidate answers are some category labels.
[0099] True label: For each training sample (i.e., each case), the true label is the actual type of this case (i.e., one of the categories in the corresponding candidate answers). During training, the model will output the probability of each candidate answer (such as the probability of each type) according to the case description and the problem. The goal of the cross-entropy loss function is to make the probability output by the model as close as possible to the true label, that is, to make the model correctly classify the case.
[0100] For sequence labeling tasks (such as case summary generation or entity recognition), a sequence labeling loss function can be adopted. In this embodiment, the conditional random field (CRF) loss function is adopted , and the specific formula is as follows:
[0101] ;
[0102] where: is the feature vector at time step ; is the label output at time step ; is the parameter of label ; is the total number of time steps; is all possible labels, is the transpose.
[0103] After the model fine-tuning is completed, verification and evaluation become key links. Traditional evaluation metrics, such as accuracy, recall, etc., although having certain reference values, cannot fully reflect the actual performance of the model when dealing with cases. Therefore, during the evaluation process, it is necessary to design more diverse evaluation criteria in combination with the requirements of domain-specific tasks, covering aspects such as the reasoning ability, interpretability, and generation ability of the model. In addition, the evaluation of domain experts is particularly important in this process. Through the feedback loop with experts, the model can be continuously optimized to ensure that it can provide solutions that meet the actual needs.
[0104] S1.5. Optimize the adapted model to obtain a base model applicable to cases;
[0105] In this embodiment, before performing multi-task learning on the adapted model, it also includes: model optimization;
[0106] The model optimization process is carried out by minimizing the loss function, and usually uses Gradient Descent and its variants (such as the Adam optimizer) to optimize the model parameters. The model parameters in this embodiment are set to , then the optimization process can be expressed by the following formula:
[0107] ;
[0108] Where: is the iteration time step for optimizing the adapted model; is the model parameter at time step , is the model parameter at time step ; is the learning rate; is the loss function with respect to the model parameter gradient; Through iterative optimization, the model gradually adjusts the parameters to reduce the task loss. During the optimization process, the update of the parameters is based on the gradient of the entire dataset, and the time step in the CRF loss function is the specific annotation position when processing the sequence.
[0109] S1.5 includes:
[0110] Perform multi-task learning on the adapted model to obtain an optimized model that can perform multiple tasks simultaneously;
[0111] Introduce a regularization term and a learning rate scheduling strategy into the optimized model to obtain a regularized model;
[0112] The specific formula of the regularization term is as follows:
[0113] ;
[0114] The loss function after adding the regularization term is:
[0115] ;
[0116] Where, and are different regularization hyperparameters; is the parameter of the model, is the number of the model parameter. This can constrain the model to avoid overfitting during training.
[0117] Introduce a knowledge graph or a multi-modal corpus into the regularized model to obtain a base model suitable for the case.
[0118] In the model optimization phase, the model can be enhanced through a knowledge graph or a multi-modal corpus. For example, by introducing a multi-modal corpus related to the case or relevant regulations , the model can search for relevant information during inference, thereby improving the inference accuracy of the model. Assuming there is a certain relationship between the multi-modal corpus and the probability output by the model, the loss function of the model can be adjusted to , and the specific formula is as follows:
[0119] ;
[0120] where is the data number related to the label in the multi-modal corpus; represents the weight related to the label in the multi-modal corpus. By introducing the information of the multi-modal corpus, the model can more accurately identify the specific features of the case during the inference process.
[0121] The focus of the model optimization phase is to introduce domain knowledge into the model to further enhance its inference and decision-making capabilities. The multi-modal corpus can be introduced to enrich the background knowledge of the model, thereby improving its understanding and inference capabilities for complex cases. At the same time, as more labeled data is collected, the model should be updated and iteratively trained periodically to adapt to new threats and challenges emerging in the case-related fields.
[0122] Finally, in the model deployment and application phase, customized interfaces and APIs should be designed according to specific application scenarios (such as case trial support, network security monitoring, criminal case analysis, etc.) to embed the model into the actual work process. At the same time, considering the continuous changes in the case form, the monitoring and update mechanism of the model should also be continuously carried out to ensure its long-term effectiveness and adaptability. In this way, it can be ensured that the constructed base model can continuously provide effective support in the dynamically complex case-related fields.
[0123] S2: Fine-tune the parameter information of the base model based on LoRA and P-Tuning to obtain the fine-tuned large language model;
[0124] In intelligent question answering for cases, the application of LoRA fine-tuning can significantly improve the performance of the model in handling complex questions. For questions that require multi-step logical reasoning or in-depth context understanding, LoRA fine-tuning can help the model more accurately parse the key elements of the question and effectively extract relevant information from a large amount of data. This is crucial for improving the accuracy and adaptability of intelligent question answering. In addition, the role of P-Tuning fine-tuning in intelligent question answering is reflected in generating more accurate and relevant question answers. By continuously updating and optimizing the prompts, the model can better understand the specific requirements of the question and accurately quote and integrate relevant information in the answer. This method is particularly suitable for question-answering tasks that require understanding complex contexts and performing in-depth information processing. Through fine-tuning, the knowledge of the large language model is maximally transferred to the target task, reducing the dependence on a large amount of domain-specific labeled data and enabling the rapid adaptation and deployment of the model. The parameter information of the base model is used as the initialization parameters of the new model to be trained in the target domain, and then some parameters of the model are updated through the labeled corpus in the target domain. This method can accelerate and optimize the learning efficiency of the model without making the model learn from scratch, greatly reducing the training cost.
[0125] In the intelligent question-answering task in the field related to cases, the combination of two fine-tuning methods, LoRA (Low-Rank Adaptation) and P-Tuning, can make up for the deficiencies of a single fine-tuning method and jointly improve the processing ability of the model. They optimize the model from different perspectives, and after combination, they can significantly improve the performance of the model in complex reasoning tasks, especially in case reasoning and the handling of technical details.
[0126] S2 includes:
[0127] Based on the LoRA fine-tuning method, the weight matrix of a specific layer in the base model is optimized by introducing a low-rank matrix to obtain the base model after the first fine-tuning.
[0128] In this embodiment, the weight matrix of a certain layer of the model is , after LoRA fine-tuning, the weight matrix of this layer is decomposed into the product of low-rank matrices. The mathematical formula of LoRA fine-tuning is as follows:
[0129] ;
[0130] Where: is the original weight matrix. is the increment of the low-rank matrix, and are low-rank matrices respectively to reduce the calculation and storage overhead.
[0131] Through LoRA fine-tuning, the model can effectively adjust the weights of these layers to enhance the model's performance in case tasks, such as enhancing the ability to understand technical details and the accuracy of case reasoning.
[0132] Based on the P-Tuning fine-tuning method, the prompt for the case question-answering task in the input end of the base model after the first fine-tuning is optimized to obtain the fine-tuned large language model.
[0133] In this embodiment, the core of P-Tuning fine-tuning is to enable the model to better understand the context of a specific task by optimizing the task prompt. Suppose the input prompt is , P-Tuning adjusts the output of the model by optimizing the prompt . Suppose the original prompt of the input is , after P-Tuning fine-tuning, the optimized prompt is obtained:
[0134] ;
[0135] where: is the original prompt. is the prompt adjustment obtained through gradient descent or other optimization methods during the P-Tuning process.
[0136] By optimizing the prompt, P-Tuning can guide the model to be more accurate when processing specific tasks (such as case reasoning, case fact matching), thereby improving the performance of the question-answering system.
[0137] In summary, the reasoning process after combining LoRA and P-Tuning fine-tuning can be expressed as:
[0138] ;
[0139] where: represents the model weight matrix optimized based on LoRA fine-tuning for reasoning about relevant regulations and case facts process. represents the reasoning process based on the optimized task prompt .
[0140] During the fine-tuning process, a loss function is usually defined to measure the gap between the model prediction and the true label, and the goal is to minimize this loss function. Under the combination of LoRA and P-Tuning fine-tuning, the loss function can be expressed as:
[0141] ;
[0142] Wherein: is the loss function optimized by LoRA fine-tuning, which measures the reasoning error based on relevant regulations and case facts. is the loss function optimized by P-Tuning fine-tuning, which measures the Q&A error based on prompt optimization.
[0143] Finally, the minimization of the loss function updates the low-rank matrix increment in LoRA fine-tuning through backpropagation and , as well as the prompts in P-Tuning fine-tuning .
[0144] After the model combines LoRA and P-Tuning fine-tuning, the final reasoning formula can be expressed as:
[0145] ;
[0146] Here, the softmax function is used to normalize the output of the model to generate the final answer. This output represents the reasoning result of the model on the matching of relevant regulations and case facts in the case.
[0147] The key advantages of the combination of LoRA and P-Tuning fine-tuning are as follows:
[0148] (1) Complementary: LoRA optimizes the inside of the model, and P-Tuning optimizes the input and output.
[0149] (2) LoRA improves the reasoning ability of the model in dealing with specific tasks (such as the technical issues and case reasoning of cases) by optimizing the internal structure and weights of the model, especially performing outstandingly in tasks with deep reasoning levels and high requirements for background knowledge. It helps the model better understand the relationship between case facts and relevant regulations.
[0150] (3) P-Tuning focuses on optimizing the prompts at the input end, enabling the model to better understand and parse the background information of complex problems and generate accurate answers. By optimizing the prompts, P-Tuning can guide the model to focus on specific case details and case information, thereby improving the accuracy of the Q&A system.
[0151] (4) Enhanced reasoning ability: LoRA optimizes multi-level reasoning, and P-Tuning optimizes task context.
[0152] (5) In cases, the reasoning task may involve multiple levels of information processing, including extracting information from case facts, matching with relevant regulations, and performing logical reasoning. LoRA fine-tuning optimizes the reasoning ability of the model at these levels, enabling the model to efficiently capture the key information of the case and perform effective reasoning.
[0153] (6) At the same time, P-Tuning enables the model to accurately grasp the core of the problem at each level by optimizing the prompt. For example, when dealing with the problem of "whether it constitutes a cyber attack", P-Tuning helps the model extract the most relevant technical details from the case facts and guides the model to use the correct relevant regulations for reasoning.
[0154] (7) Reduce the number of parameters: LoRA remains lightweight, and P-Tuning improves efficiency.
[0155] (8) The low-rank matrix structure of LoRA does not significantly increase the number of parameters of the model, making the fine-tuning process more efficient and ensuring the deployability of the model in resource-constrained environments.
[0156] (9) P-Tuning reduces the computational overhead by optimizing the task prompt rather than the model parameters, making the fine-tuning process more efficient and more scalable.
[0157] S3: Introduce the CLIP module on the basis of the fine-tuned large language model to perform multi-modal feature fusion on the text and image information of the input case, and obtain a preliminary question-answering model that can understand and fuse text and image information;
[0158] S3 includes:
[0159] Use the CLIP model to encode the text and image information of the input case respectively to generate the embedding vector of the image and the embedding vector of the text;
[0160] Embed the image and text into a shared semantic space for contrastive learning and fusion to obtain a preliminary question-answering model that can understand and fuse text and image information.
[0161] Specifically: The analysis of cases usually involves a large amount of multi-modal data, such as text, pictures, log files, etc. Traditional text analysis methods cannot effectively process image information, while the multi-modal feature fusion method based on the CLIP (Contrastive Language-Image Pre-training) model can combine image and text information, so as to achieve more accurate case analysis and intelligent question answering. This solution aims to develop an intelligent question-answering system that can understand and fuse text and image information by combining the CLIP model and the large language model, so as to improve the automated processing ability of cases.
[0162] This study will adopt a method that combines the CLIP model and large language models for feature fusion and multimodal data processing. The CLIP model encodes images, maps the image content to a semantic space, enabling contrastive learning and fusion with text features. The large language model is responsible for processing and generating text-based answers or inferences.
[0163] The CLIP model was proposed by OpenAI. Using the method of contrastive learning, it embeds images and text into a shared semantic space, capable of processing and understanding the relationship between images and text. The CLIP model consists of two parts:
[0164] Image encoder: Encodes images through a convolutional neural network (CNN) or Vision Transformer (ViT) to generate embedding vectors of the images ;
[0165] Text encoder: Encodes text through a pre-trained Transformer model to generate embedding vectors of the text ;
[0166] The core of the CLIP model is to optimize the similarity between images and text through contrastive learning, that is, by maximizing the similarity between image and text vectors with the same semantics and minimizing the similarity between image and text vectors with different semantics.
[0167] The core idea of the CLIP model is to optimize the common representation of images and text through contrastive learning, making images and text with the same semantics as close as possible in the shared semantic space, while images and text with different semantics should be as far apart as possible. This is achieved by designing a specific loss function.
[0168] The goal of contrastive learning is to maximize the similarity between positive samples (i.e., semantic consistency between image and text pairs) while minimizing the similarity between negative samples (i.e., semantic inconsistency between image and text pairs). To achieve this, the CLIP model optimizes a contrastive loss function to measure the similarity between image and text vectors.
[0169] The loss function used by CLIP is based on contrastive loss and is usually called the information maximization loss. This loss function calculates the similarity between image and text vectors in the shared semantic space and encourages the model to make image and text vectors with similar semantics close to each other, while making image and text vectors with different semantics as far apart as possible.
[0170] For each pair of images and text, the loss function of the CLIP model can be expressed in the following form:
[0171] ;
[0172] wherein: is the number of negative samples; is the embedding vector of the image, representing the semantic features of the image. is the embedding vector of the text, representing the semantic features of the text. and are the embedding vectors of other images and texts respectively. is usually a hyperparameter less than 1, which controls the compactness of the embedding space and affects the scale of similarity calculation. The numerator part calculates the similarity between the image and text vectors. The denominator part contains the sum of the similarities between all images and all texts and the current image or text, ensuring that the influence of negative samples is effectively minimized.
[0173] In case analysis, CLIP can effectively perform contrastive learning on case-related images (such as screenshots, on-site photos, network traffic diagrams, etc.) and descriptive texts (such as case summaries, human descriptions, event logs, etc.). Through training, the CLIP model can map images and texts into a shared semantic space, so that when the preliminary question-answering model conducts intelligent question answering, it can understand and combine image and text information to provide more accurate and comprehensive answers.
[0174] S4: Introduce a multi-dimensional logical association thinking chain into the preliminary question-answering model to obtain a generation model; the generation model outputs multiple candidate answers with reasoning chains according to the relevant questions of the input case;
[0175] See Figure 3 , S4 includes:
[0176] S4.1. Construct a multi-dimensional logical association thinking chain based on syllogism and analogical reasoning; specifically:
[0177] ①. Input the relevant regulations as the major premise and the graphic and text information of the case as the minor premise into the preliminary question-answering model;
[0178] Traditional case reasoning relies on deductive reasoning. By combining relevant regulations (major premise) with case facts (minor premise), the conclusion of a specific case is obtained. However, in cases, relevant regulations are often rather general, and case facts are complex and changeable, and the relationship between the two is not always linear or direct. Therefore, simply relying on traditional syllogistic reasoning may lead to insufficient reasoning and even incorrect reasoning results.
[0179] ②. Determine whether the graphic and text information of the input case complies with the relevant regulations under the main premise. If it complies, the preliminary Q&A model outputs a preliminary conclusion. If it does not comply, similar historical cases and judgments similar to the current case are searched, and a similar conclusion is derived by analogical reasoning. Among them: the main premise, the secondary premise, and the reasoning conclusion based on the main premise and the secondary premise are the three elements of syllogism. Specifically:
[0180] Combine the reasoning conclusions of syllogism and analogical reasoning to generate a complete reasoning chain and the answer to the case consultation;
[0181] In this embodiment, according to the reasoning structure of syllogism, if the case facts meet the applicable relevant regulations, we can draw a conclusion , that is, the judgment of the case or the relevant regulations apply:
[0182] If and are true, then the conclusion is drawn, and respectively represent the relevant regulations and case facts used in the current step;
[0183] That is: ;
[0184] Among them, is the conclusion of the case. If and are both true, then the reasoning result is correct, and it is determined that is true.
[0185] Analogical reasoning helps the model to reason by identifying precedent cases similar to the current case. Suppose we have a set of precedent cases , where each case has its factual description and judgment result. The core idea of analogical reasoning is that if a case is similar to the current case in some key features, then we can speculate on the judgment result of the current case. Define a similarity metric to measure the similarity between the case and the precedent case :
[0186] ;
[0187] Among them are the factual features of the current case , are the factual features of the precedent case , is a similarity metric function, which can usually be calculated by methods such as cosine similarity and Euclidean distance.
[0188] Once the similarity measure of the case is calculated, the process of analogical reasoning can be expressed as:
[0189] ;
[0190] Where, is the predicted conclusion of the current case, is the judgment result of the precedent case . Through the weighted similarity measure, the model will predict the most likely judgment result.
[0191] Two results are obtained according to syllogism and analogical reasoning: and , then the final reasoning result can be expressed as:
[0192] ;
[0193] Where, is a weight coefficient that controls the influence of the two reasoning methods. Through this weighting mechanism, the model can flexibly combine deductive reasoning and analogical reasoning, thereby improving the accuracy and adaptability of reasoning. For different cases, the weight can be adjusted according to the characteristics of the case to better adapt to the relationship between different types of case facts and relevant regulations.
[0194] In some complex cases, the reasoning may require multiple iterations or recursions. In this case, we can regard the reasoning process as a recursive model. Assume that the reasoning chain contains multiple steps, and the reasoning result of each step depends on the conclusion of the previous step. We can express it as a recursive formula:
[0195] ;
[0196] Where, is the result of the th step of reasoning, is the result of the previous round of reasoning.
[0197] Through recursive reasoning, the model can gradually optimize and improve the reasoning process, and gradually enhance the reasoning ability for complex cases.
[0198] S4.2. Introduce the thinking chain of multi-dimensional logical association into the preliminary question-answering model to obtain a generation model.
[0199] In this embodiment, introducing the thinking chain has the following advantages:
[0200] (1). Enhancement of reasoning ability
[0201] ZK-CoT not only relies on the strict rules of syllogism but also introduces the flexibility of analogical reasoning. By enhancing the model's ability to handle potential mismatches between relevant regulations and case facts, the model can provide more reasonable and accurate reasoning results when faced with complex or ambiguous cases.
[0202] (2), Multi-level reasoning structure
[0203] ZK-CoT combines deductive reasoning and analogical reasoning through a multi-level reasoning structure to ensure that the model has flexible adaptability when dealing with different types of case problems. The syllogism provides a rigorous logical framework, while analogical reasoning provides the model with greater flexibility and adaptability, enabling it to make more practical inferences in complex and ambiguous situations.
[0204] (3), Enhanced interpretability
[0205] Analogical reasoning can help the model understand and apply the underlying logic of relevant regulations. By drawing on the application of relevant regulations and judgment logic in similar cases, analogical reasoning helps improve the model's interpretability.
[0206] (4), Support for reasoning in complex cases
[0207] In cases, the relationship between case facts and relevant regulations is often complex and not straightforward. ZK-CoT can effectively handle case reasoning tasks in complex cases. Especially when case facts involve new types of criminal acts or technical issues, analogical reasoning can provide more appropriate answers.
[0208] S5: Train the existing scoring model based on the candidate answers generated by the contrastive learning method and the generative model to obtain an evaluation model that can output the optimal answer among multiple candidate answers.
[0209] S5 includes:
[0210] S5.1, Construct a contrastive learning loss function and a ranking loss function;
[0211] S5.2, Train the existing scoring model based on the contrastive learning loss function and the ranking loss function to learn to distinguish high-quality answers and low-quality answers among multiple candidate answers and rank them, obtaining an evaluation model that can output the optimal answer among multiple candidate answers, where: high-quality answers are positive samples, and low-quality or irrelevant answers are negative samples;
[0212] Specifically:
[0213] S5.2.1, Input the relevant questions and candidate answers of the case: Input the relevant questions and multiple candidate answers into the existing scoring model;
[0214] S5.2.2. Calculate sentence embeddings: Use a pre-trained language model to generate sentence embeddings for questions and candidate answers;
[0215] S5.2.3. Construct the total loss function: Construct a total loss function composed of a contrastive learning loss function and a ranking loss function;
[0216] The contrastive learning loss function is a classification objective function or a triplet objective function. The classification objective function The specific formula is as follows:
[0217] ;
[0218] Where: is the time step for training the existing scoring model; is a trainable weight matrix, , is the set of real numbers, is the dimension of the sentence embedding, is the number of label categories in the classification objective function; , are two different sentence embeddings;
[0219] The classification objective function can help the model distinguish different candidate answers, improve the classification accuracy of the model in different types of questions, and ensure that the generated answers have high semantic relevance.
[0220] The triplet objective function The specific formula is as follows:
[0221] ;
[0222] Where: is the sentence embedding of the given anchor sentence , is the sentence embedding of the positive sentence , is the sentence embedding of the negative sentence ; is the margin;
[0223] Given an anchor sentence , a positive sentence and a negative sentence , the triplet loss function adjusts the network so that and the distance between is less than and the distance between.
[0224] By comparing positive and negative samples, it helps the model distinguish correct and incorrect answers in complex reasoning tasks, especially performing excellently in case reasoning and case analysis.
[0225] Ranking loss function The specific formula is as follows:
[0226] ;
[0227] is the candidate set composed of candidate answers, that is, the candidate set filtered from all possible labels in the current task; and are the th candidate answer and the th candidate answer respectively, , are the labels of the candidate answers respectively, ; represents the similarity between the candidate answer and the question ; is the similarity between the candidate answer and the question ; , are the labels of the question respectively, .
[0228] In this embodiment, is obtained by calculating the cosine similarity, and the specific formula is as follows:
[0229] ;
[0230] Among them: and are the sentence embedding vectors of the candidate answer and the question respectively, and these vectors are generated by a pre-trained model (such as RoBERTa). represents the dot product of vectors, that is, calculating the inner product between two vectors. and are the L2 norms of the vectors and respectively, that is, their respective magnitudes.
[0231] The core idea of the ranking loss function is to make the model correctly rank among multiple candidate answers according to their relevance to the question by minimizing the objective loss. Specifically, the ranking learned by the model should be: for the question , the more relevant answers should be ranked first; it improves the model's ability to calculate the similarity between questions and answers and is applicable to question matching and selection of similar answers.
[0232] Evaluate the total loss function of the model is the weighted sum of the contrastive learning loss and the ranking loss:
[0233] ;
[0234] Where: and are hyperparameters used to adjust the contrastive learning loss and the ranking loss in the total loss function. By jointly optimizing these two loss functions, the evaluation model can not only distinguish good and bad answers but also effectively rank multiple candidate answers, thereby improving the model's ability to generate high-quality answers.
[0235] S5.2.4. Calculate the loss: Evaluate the quality of each candidate answer by calculating the total loss function and rank them according to the relevance between the candidate answer and the question;
[0236] S5.2.5. Gradient update: Use the gradient descent algorithm through backpropagation to optimize the total loss function and update the weights of the existing scoring model;
[0237] S5.2.6. Iterative training: Repeat S5.2.4 to S5.2.5 until the total loss function of the existing scoring model converges, and obtain an evaluation model that can output the optimal answer among multiple candidate answers. In this embodiment, when the loss reduction between two rounds of iterative training is less than 0.001, it is considered that the training converges.
[0238] Under this training mechanism, the evaluation model will imitate the real evaluation metrics to rank the candidate answers. This means that even in the absence of a standard answer, the model can effectively rank the answers based on the context of the original document and the question. A key advantage of this method is that it can improve the model's ability to generate high-quality answers when there is no direct reference answer, especially when dealing with complex questions that require comprehensive understanding and reasoning.
[0239] This embodiment also provides a readable storage medium on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above-mentioned case question-answering method based on a large language model is implemented.
[0240] Exemplarily, the computer program may be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device.
[0241] This embodiment further includes an electronic device, comprising: at least one processor, at least one memory, and computer program instructions stored in the memory, and when the computer program instructions are executed by the processor, the case question-answering method based on the large language model is as described above.
[0242] The electronic device may be a computing device such as a mobile phone, a desktop computer, a notebook, a palm computer, and a cloud server. The electronic device may include, but is not limited to, a processor and a memory. For example, the electronic device may further include input / output devices, network access devices, a bus, etc.
[0243] The foregoing are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A case question answering method based on a large language model, characterized in that: The steps include: Obtain graphic and text information of the case and relevant issues of the case; The obtained case graphic information and case-related questions are input into the intelligent question-answering model to obtain the optimal answer. The intelligent question-answering model is obtained through the following steps: S1: Construct a multimodal corpus containing case-related knowledge, and train a pre-trained model through the multimodal corpus based on transfer learning to obtain a base model suitable for the case, wherein: the data format of the multimodal corpus includes text and pictures; S2: Fine-tune the parameter information of the base model based on LoRA and P-Tuning to obtain a fine-tuned large language model; S3: Based on the fine-tuned large language model, the CLIP module is introduced to perform multimodal feature fusion on the image and text information of the input case, and a preliminary question-answering model that can understand and fuse image and text information is obtained; S4: Introduce a multi-dimensional logically associated thinking chain into the preliminary question-answering model to obtain a generative model; the generative model outputs multiple candidate answers with reasoning chains based on the relevant questions of the input case; The S4 includes: S4.
1. Construct a multi-dimensional logically related thinking chain based on syllogism and analogical reasoning; specifically: Input the relevant regulations as the main premise and the case's graphic and text information as the secondary premise into the preliminary question-answering model; Determine whether the input case graphic information complies with the relevant provisions under the main premise. If so, the preliminary question-answering model outputs a preliminary conclusion. If not, search for historical cases and judgments similar to the current case and deduce similar conclusions by analogy. The main premise, the secondary premise, and the inference conclusion based on the main premise and the secondary premise are the three elements of the syllogism. Combine the reasoning conclusions of syllogism and analogy to generate a complete chain of reasoning and answers to case consultations; According to syllogism and analogical reasoning, two results are obtained: C deductive and C analogical , then the final inference result C is expressed as: C=α·C deductive +(1-α)·C analogical ; Among them, α is a weight coefficient that controls the influence of the two inference methods; Introducing recursive reasoning in complex cases, the specific formula is as follows: C (t″) =f(C (t″-1) ,P(L),P(F)); Among them, C (t”) is the result of the t”th step reasoning, C (t”-1) is the result of the previous round of reasoning; P(L) and P(F) represent the relevant provisions and case facts used in the current step respectively; S4.2, introduce the multi-dimensional logically related thinking chain into the preliminary question-answering model to obtain the generative model; S5: Train the existing scoring model based on the contrastive learning method and the candidate answers generated by the generative model to obtain an evaluation model that can output the best answer among multiple candidate answers.
2. The case question answering method based on a large language model according to claim 1, characterized in that: The S1 includes: S1.
1. Collect multi-source data in case-related fields, and pre-process the collected multi-source data in case-related fields to obtain a multimodal corpus containing case-related knowledge, wherein: the text data of case-related knowledge includes case-related regulations and case judgment information; the pre-processing includes data cleaning of the multi-source data in case-related fields, and labeling of unlabeled data in the multi-source data in case-related fields; S1.
2. Select a language model pre-trained on a large-scale general corpus as a pre-trained model, and initialize the pre-trained model; S1.
3. Extract annotated data sets in case-related fields from the multimodal corpus, and continue to train the pre-trained model with the annotated data sets in case-related fields to obtain a fine-tuned model. The specific formula is as follows: in, is the loss function on the target task; is a pre-trained model; D target It is a labeled dataset in case-related fields. is the fine-tuned model; S1.4, adapting the fine-tuned model to a specific case question-answering task to obtain an adapted model, wherein the case question-answering task includes question understanding, answer generation, and information extraction; S1.
5. Optimize the adapted model to obtain a base model suitable for the case.
3. The case question answering method based on a large language model according to claim 2 is characterized in that: The S1.5 includes: Perform multi-task learning on the adapted model to obtain an optimized model that can perform multiple tasks simultaneously; Introduce regularization terms and learning rate scheduling strategies into the optimization model to obtain a regularized model; The knowledge graph or multimodal corpus is introduced into the regularized model to obtain a base model suitable for the case.
4. The case question answering method based on a large language model according to claim 3 is characterized in that: The S2 includes: Based on the LoRA fine-tuning method, a low-rank matrix is introduced to optimize the weight matrix of a specific layer in the base model to obtain the base model after the first fine-tuning. Based on the P-Tuning fine-tuning method, the prompt for the case question answering task at the input of the base model after the first fine-tuning is optimized to obtain the fine-tuned large language model.
5. The case question answering method based on a large language model according to claim 4 is characterized in that: The S3 includes: The CLIP model is used to encode the image and text information of the input case respectively, generating the image embedding vector and the text embedding vector; Images and texts are embedded in a shared semantic space for comparative learning and fusion, resulting in a preliminary question-answering model that can understand and integrate image and text information.
6. The case question answering method based on a large language model according to claim 5 is characterized in that: The S5 includes: S5.
1. Construct contrastive learning loss function and ranking loss function; S5.
2. Based on the contrastive learning loss function and the ranking loss function, the existing scoring model is trained to learn to distinguish high-quality answers from low-quality answers among multiple candidate answers and to rank them, thereby obtaining an evaluation model that can output the best answer among multiple candidate answers, where: high-quality answers are positive samples, and low-quality or irrelevant answers are negative samples; Specifically: S5.2.
1. Input relevant questions and candidate answers of the case: Input relevant questions and multiple candidate answers into the existing scoring model; S5.2.
2. Compute sentence embeddings: Generate sentence embeddings for questions and candidate answers using a pre-trained language model. S5.2.
3. Construct a total loss function: Construct a total loss function consisting of a contrastive learning loss function and a ranking loss function; S5.2.
4. Calculate loss: Evaluate the quality of each candidate answer by calculating the total loss function and sort the candidate answers by their relevance to the question; S5.2.5, Gradient update: Update the weights of the existing scoring model by optimizing the total loss function using the gradient descent algorithm through the back-propagation algorithm; S5.2.
6. Iterative training: Repeat S5.2.4 to S5.2.5 until the total loss function of the existing scoring model converges, and an evaluation model that can output the best answer among multiple candidate answers is obtained.
7. The case question answering method based on a large language model according to claim 6 is characterized in that: The contrastive learning loss function is a classification objective function or a triplet objective function. The specific formula of the classification objective function O is as follows: Where: t0 is the time step for training the existing scoring model; is the trainable weight matrix, n is the dimension of sentence embedding, k is the number of label categories in the classification objective function; u and v are two different sentence embeddings, |uv| is the distance between u and v; The specific formula of the triplet objective function O' is as follows: O'=max(||h a -h p ||-||h a -h z ||+∈,0); Where: h a is the sentence embedding of a given anchor sentence a, h p is the sentence embedding of the positive sentence p, h z is the sentence embedding of the forward sentence z; ∈ is the boundary; Ranking loss function The specific formula is as follows: in: is the candidate set consisting of candidate answers, x i and x j are the i-th candidate answer and the j-th candidate answer, i and j are the labels of the candidate answers, i≠j; S(y b ,x j ) represents the candidate answer x j With the problem b The similarity of S(y c ,x i ) is the candidate answer x i With the problem c similarity, b and c are the labels of the questions respectively, b≠c.
8. A readable storage medium, characterized in that: Computer program instructions are stored thereon, and when the computer program instructions are executed by a processor, a case question answering method based on a large language model as described in any one of claims 1 to 7 is implemented.
9. An electronic device, characterized in that: include: At least one processor, at least one memory, and computer program instructions stored in the memory, when the computer program instructions are executed by the processor, the case question answering method based on a large language model as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Visual knowledge reasoning question and answer method for multi-source heterogeneous knowledge joint enhancement
CN117010500A
Multi-modal large model implementation method and system for organizational knowledge management
CN117709356A