Knowledge retrieval-oriented legal element identification method and system
Through the legal element identification method for knowledge retrieval, the open source large language model, the maximum likelihood method and gradient descent method are used to optimize parameters, and the existing system cannot meet the comprehensive search needs of complex legal issues, improving the accuracy and efficiency of legal information retrieval.
Patent Information
- Application Number
- CN202510276224.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-06
AI Technical Summary
The existing legal information service system cannot meet the comprehensive retrieval needs of complex legal issues, and single-modal data processing affects the accuracy of the model.
The legal element recognition method for knowledge retrieval is adopted. By obtaining the target information to be processed, searching using the knowledge base, extracting features and inputting them into the trained open source large language model, and legal element recognition and knowledge reasoning are performed. During the training process, the maximum likelihood method and gradient descent method are used to optimize the model parameters.
In-depth analysis and comprehensive search of complex legal issues have been achieved, and the accuracy of model parameters and subsequent search recognition have been improved.
Smart Images

Figure CN120104722A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field related to legal retrieval, and in particular, relates to a legal element identification method and system for knowledge retrieval. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] With the development of artificial intelligence technology, the demand for intelligence, informationization and automation in the legal service field is growing. Existing legal information service systems are often limited to single-mode data processing and cannot meet the comprehensive retrieval needs of complex legal issues. In addition, the comprehensive retrieval needs will also affect the accuracy of the model. Summary of the invention
[0004] In order to overcome the deficiencies of the above-mentioned prior art, the present invention provides a legal element identification method and system for knowledge retrieval. The open source large language model has excellent semantic understanding, can realize in-depth analysis and understanding of legal cases and laws and regulations, and realize comprehensive retrieval needs for complex legal issues; the maximum likelihood method and gradient descent are used to improve the accuracy of model parameters and improve the accuracy of subsequent retrieval and identification.
[0005] In order to achieve the above object, the present invention adopts the following technical solution:
[0006] In a first aspect, the present invention provides a legal element identification method for knowledge retrieval, comprising:
[0007] Get the target information to be processed;
[0008] Retrieving the target information to be processed through the knowledge base to obtain relevant legal knowledge;
[0009] Extract features of the target information to be processed and the corresponding relevant legal knowledge respectively, splice the extracted features, input the spliced features into a trained open source large language model, and use the open source large language model to perform legal element recognition and knowledge reasoning on the target information to be processed;
[0010] In the process of training the open source large language model, the maximum likelihood method is used to estimate the model parameters, the model parameters corresponding to maximizing the log-likelihood function are determined, and the gradient descent method is used to update the model parameters.
[0011] In a second aspect, the present invention provides a legal element identification system for knowledge retrieval, comprising:
[0012] An acquisition module is configured to: acquire target information to be processed;
[0013] A retrieval module is configured to: retrieve the target information to be processed through a knowledge base to obtain relevant legal knowledge;
[0014] The recognition module is configured to: extract features from the target information to be processed and the corresponding relevant legal knowledge respectively, splice the extracted features, input the spliced features into a trained open source large language model, and perform legal element recognition and knowledge reasoning on the target information to be processed through the open source large language model;
[0015] In the process of training the open source large language model, the maximum likelihood method is used to estimate the model parameters, the model parameters corresponding to maximizing the log-likelihood function are determined, and the gradient descent method is used to update the model parameters.
[0016] In a third aspect, the present invention provides an electronic device comprising a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method described in the first aspect is performed.
[0017] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions, wherein when the computer instructions are executed by a processor, the method described in the first aspect is performed.
[0018] In a fifth aspect, the present invention provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method described in the first aspect.
[0019] One or more of the above technical solutions have the following beneficial effects:
[0020] In the present invention, the target information to be processed and the retrieved legal knowledge are input into a trained open source large language model, and legal element recognition and knowledge reasoning are performed through the open source large language model; in the training process of the open source large language model, the maximum likelihood method is used to estimate the model parameters, the model parameters corresponding to the maximum log-likelihood function are determined, and the gradient descent method is used to update the model parameters; the open source large language model has excellent semantic understanding, can achieve in-depth analysis and understanding of legal cases and laws and regulations, and realize the comprehensive retrieval needs of complex legal issues; the maximum likelihood method and gradient descent are used to improve the accuracy of model parameters and improve the accuracy of subsequent retrieval and recognition.
[0021] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0023] Figure 1 This is a flow chart of the open source large language model training in Embodiment 1 of the present invention;
[0024] Figure 2 This is a legal element identification framework diagram for knowledge retrieval in the first embodiment of the present invention. DETAILED DESCRIPTION
[0025] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0026] It should be noted that the terms used herein are for describing specific embodiments only and are not intended to be limiting of exemplary embodiments according to the present invention.
[0027] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.
[0028] Embodiment 1
[0029] This embodiment discloses a legal element identification method for knowledge retrieval, including:
[0030] Get the target information to be processed;
[0031] Retrieving the target information to be processed through the knowledge base to obtain relevant legal knowledge;
[0032] Extract features of the target information to be processed and the corresponding relevant legal knowledge respectively, splice the extracted features, input the spliced features into a trained open source large language model, and use the open source large language model to perform legal element recognition and knowledge reasoning on the target information to be processed;
[0033] In the process of training the open source large language model, the maximum likelihood method is used to estimate the model parameters, the model parameters corresponding to maximizing the log-likelihood function are determined, and the gradient descent method is used to update the model parameters.
[0034] In this embodiment, historical legal cases are obtained, and legal elements such as parties involved and legal clauses are marked. The legal elements of the legal cases are marked, and the marked content covers key elements such as parties involved, legal clauses, case facts, etc., and the text is converted into a format that can be processed by an open source large language model, such as word segmentation, encoding, etc.; through historical legal cases, the open source large language model is allowed to perform label learning, cross entropy is used as the loss function, and the Adam optimizer is used for model optimization.
[0035] The classification algorithm formula used in case label learning is:
[0036]
[0037] Among them, x is the case data, f(x) is the output of the open source large language model, and Label is the predicted label.
[0038] Among them, legal cases come from public data such as Peking University Yinghua, CAIL2024, C3RD, and DISC-Law-SFT.
[0039] Introduce multimodal data (such as text, voice, pictures, etc.) for multimodal fusion training to improve the model's ability to process complex legal information:
[0040] z=φ(x 1 , x 2 , …, x M )
[0041] Among them, x 1 , x 2 , …, x M They represent data of different modes respectively, φ is the fusion function, and z is the feature representation after fusion.
[0042] In this embodiment, it also includes cleaning and preprocessing of historical legal cases, including removing noise data, text standardization, word segmentation, removing stop words, etc., and converting the text into a format that can be processed by an open source large language model, such as word segmentation, encoding, etc.
[0043] When processing text data, the long short-term memory model LSTM (Long Short-Term Memory) is used for word segmentation. This is a special RNN (Recurrent Neural Network) that can effectively process long sequence data. This technology is based on a large-scale pre-trained model and can more accurately understand the professional terms and contextual relationships in legal texts. At the same time, the encoding is not a traditional one-layer encoding, but a multi-layer encoding structure. By fusing the encoding features of different levels, the model's ability to understand the semantics of complex legal information is enhanced.
[0044] This example uses the multi-layer attention mechanism of the Transformer encoder to automatically capture the key elements in the legal text and the relationship between the elements. Compared with the traditional rule-based word segmentation method and simple encoding technology, this processing method has significantly improved the accuracy and depth of semantic understanding. The specific steps are as follows:
[0045] First, input embedding.
[0046] The input text sequence x 1 , x 2 , …, x T Convert to an embedded representation, where x t is the input word at time step t:
[0047] e t =W embed x t +b embed
[0048]
[0049] Among them, e t is the word embedding vector at time step t, is the position embedding vector at time step t, is the total embedding at time step t, W embed is the word embedding matrix, W pos is the position embedding matrix, p t is the position index at time step t, b embed is the bias term of word embedding, b pos is the bias term for position embedding.
[0050] Second, participle.
[0051] Use LSTM cells to calculate the hidden state and memory unit at each time step. The LSTM cell contains a forget gate f t , input gate i t , candidate memory cells Memory Unit C t , output gate o t And the hidden state h t The specific calculation formula is as follows:
[0052]
[0053] h t =o t tanh(C t )
[0054] Among them, σ is the sigmoid activation function, tanh is the hyperbolic tangent activation function, Wf , W i , W C , W o etc. correspond to the weight matrices of different units, b f 、b i etc. correspond to the bias items of different units, h t-1 and C t-1 are the hidden state and memory unit of the previous time step respectively.
[0055] The word segmentation probability distribution of each time step is generated through the fully connected layer and the softmax layer. The hidden state h of LSTM t It usually goes through a fully connected layer and a softmax layer to generate the word segmentation probability distribution for each time step:
[0056] z t =W z h t +b z
[0057] y t =softmax(z t )
[0058]
[0059] Among them, z t is the output of the fully connected layer, W t is the input of the fully connected layer, b t is the bias term of the fully connected layer, y t is the probability distribution of word segmentation at time step t, and K is the number of categories of word segmentation labels.
[0060] Third, encoding.
[0061] The multi-head self-attention mechanism generates attention weights by calculating the similarity between Query, Key, and Value. The specific formula is as follows:
[0062] Linear transformation:
[0063] Q=e total W Q , K = e total W K , V = e total W V
[0064] Where Q is the Query matrix, K is the Key matrix, and V is the Value matrix. Q , W k , W V is the weight matrix of the linear transformation.
[0065] Attention score and weight:
[0066]
[0067] α=softmax(Sim(Query,Key))
[0068] Among them, d k It is the dimension of Key.
[0069] Attention value and multi-head attention:
[0070] MultiHead(Q,K,V)=Concat(head 1 , head 2 , …, head h )·W O
[0071] head i =Attention(Q·W i Q , K.W. i K , V.W. i V ).
[0072] Among them, W i Q , W i K , W i V is the weight matrix of the ith head, W O is the output weight matrix.
[0073] The output MultiHead (Q, K, V) at each position is then transformed nonlinearly through a feedforward neural network. Layer normalization is used to stabilize and accelerate the training process, and residual connections are used to alleviate the gradient vanishing problem in deep networks:
[0074] FFN(x)=max(0,x·W 1 +b 1 )·W 2 +b 2
[0075]
[0076] Among them, x is the input vector of the feedforward neural network, W 1 and W 2 is the weight matrix, b 1 and b 2 is the bias term, x norm is the input vector of the normalization layer, μ is the mean, σ 2is the variance, and ∈ is a small constant used to prevent division by zero errors.
[0077] The encoder consists of a multi-head attention mechanism and a feedforward neural network. The specific formula is as follows:
[0078] EncoderLayer(x encoder )
[0079] =LayerNorm(x encoder +MultiHead(x encoder , x encoder , x encoder ))+LayerNorm(x encoder +FFN(x encoder ))
[0080] Among them, x encoder Represents the input vector of each layer of the encoder.
[0081] Multiple encoders are stacked to form a deep Transformer model:
[0082] Encoder(x input )
[0083] =EncoderLayer N (EncoderLayer N-1 (…EncoderLayer 1 (x input )…))
[0084] Among them, x input is the initial input vector to the first encoder and N is the number of encoder layers.
[0085] Fourth, output.
[0086] The final output is obtained through a fully connected layer and a softmax layer:
[0087] z=W z Encoder(x input )+b z
[0088] y=softmax(z)
[0089] Here, z is the output of the fully connected layer.
[0090] Fifth, optimization.
[0091] The cross entropy loss function is used to measure the difference between the prediction and the true label.
[0092]
[0093] Where T represents the number of time steps, K represents the number of categories, is the true label distribution at time step t, is the predicted label distribution at time step t.
[0094] Use the Adam optimizer to minimize the loss function and update the model parameters.
[0095]
[0096] Among them, θ is the model parameter, η is the learning rate, is the gradient of the loss function with respect to the model parameters.
[0097] In this embodiment, the open source large language model is fine-tuned in a targeted manner to learn more than 60,000 relevant laws and regulations such as the Criminal Law, Judicial Interpretation, Constitution, and Civil Code.
[0098] The training of the open source large language model can be expressed as:
[0099]
[0100] in, is the loss function of the ith law, x i is the legal text entered, y i is the legal interpretation of the model’s predictions.
[0101] Optional. A total of 64,426 laws and regulations and 465,003 case data are learned, about 0.7s / persample.
[0102] Perform inference tests on the trained open source large language model, obtain output results through forward propagation, and compare the true value with the predicted value. The comparison indicators include accuracy, precision, recall and F1 score to ensure the efficiency and accuracy of training the open source large language model.
[0103]
[0104] in, is the prediction result of the model, x is the input data, and θ is the model parameter.
[0105] The accuracy of the model parameters of the open source large language model is continuously improved through maximum likelihood estimation. The likelihood function L(θ) is the probability of observing the training data set D when the parameter θ is given, which can be expressed as:
[0106]
[0107] Where N represents the number of samples in the training data set D, that is, the total number of samples contained in the data set, P(y i |xi , θ) is the open source large language model for sample x i The predicted value y given i The conditional probability of , D is the set of legal and regulatory text data in the legal knowledge base.
[0108] Taking the logarithm of both sides of the formula, we get the log-likelihood function:
[0109]
[0110] The goal of maximum likelihood estimation is to find the parameter θ that maximizes the log-likelihood function, update the open source large language model parameters through the gradient descent algorithm, maximize the log-likelihood function, iteratively update θ to gradually approach the maximum value, and converge to Right now:
[0111]
[0112] Where η is the learning rate, It is the optimal parameter model obtained by maximum likelihood estimation.
[0113] This embodiment introduces model pruning and quantization technology.
[0114] Through model pruning, unimportant weight parameters in the model are removed, thereby reducing the complexity of the open source large language model, while maintaining or even improving the performance of the open source large language model.
[0115] Calculate the importance of weights, determine the pruning threshold, and perform pruning operations:
[0116] Importance(w i )=|w i |
[0117] Threshold = α mean(|w|)
[0118]
[0119] Among them, w i is the weight parameter of the open source large language model, α is the pruning ratio, which is used to determine the pruning threshold, and w′ i is the weight parameter after pruning.
[0120] Quantization technology converts the parameters of the open source large language model from floating point numbers to fixed point numbers, reducing memory usage and computational complexity, and improving the operating efficiency of the open source large language model.
[0121] Define the quantization function and perform the dequantization operation:
[0122]
[0123] w′i =q(w i )·Δ+w min
[0124] Among them, w min is the minimum value of the weight, Δ is the quantization step size, q(w i ) is the quantized weight parameter, w′ i is the weight parameter after dequantization.
[0125] In addition, we also tried a fine-tuning strategy based on reinforcement learning. By setting a reward function, we guided the open source large language model to pay more attention to the recognition accuracy of legal elements during the optimization process, making the open source large language model more responsive to the comprehensive retrieval needs of complex legal issues.
[0126] Define the reward function and perform policy gradient update:
[0127]
[0128] Among them, r t is the reward value at time step t.
[0129] In addition, adversarial training is used to enhance the robustness of the open source large language model. By introducing adversarial samples during the training process, the open source large language model's resistance to input perturbations is improved.
[0130] Generate adversarial samples and update the parameters of the open source large language model:
[0131]
[0132] Among them, x′ t is the adversarial sample, ∈ is the perturbation amplitude, indicating the perturbation size of the adversarial sample, is the adversarial loss function.
[0133] Use the open source large language model trained above to determine whether the target accuracy has been achieved, and then proceed to the next step; if the obtained open source large language model does not reach the target accuracy, return to re-learn the laws, regulations and case labels.
[0134] The accuracy rate indicates the proportion of samples predicted correctly by the classifier to the total number of samples. The formula is as follows:
[0135]
[0136] The precision rate indicates the proportion of samples that are actually positive among all samples predicted to be positive. The formula is as follows:
[0137]
[0138] The recall rate indicates the proportion of samples that are correctly predicted as positive among all samples that are actually positive. The formula is as follows:
[0139]
[0140] The F1 score is the harmonic mean of precision and recall, which can comprehensively measure the performance of the classifier. The formula is as follows:
[0141]
[0142] Among them, TP (True Positive) is a true positive example, TN (True Negative) is a true negative example, FP (False Positive) is a false positive example, and FN (False Negative) is a false negative example.
[0143] z=φ(x 1 , x 2 , …, x M ) Compare the accuracy of different models and training methods and select the best model for deployment. The final model has an F1 score of up to 92%. Based on feedback from actual applications, we continue to optimize model parameters and training strategies to ensure the long-term stability and accuracy of the system.
[0144] In actual application, first, obtain multimodal information such as legal text, voice, and pictures input by the user; retrieve the legal knowledge related to the user input information from the knowledge base, and obtain the legal knowledge related to the user input as the input of the open source large language model to further perform legal element recognition and knowledge reasoning.
[0145] Specifically: The multimodal information (text, voice, picture, etc.) input by the user is generated through the encoder into an input vector v input , the legal knowledge in the knowledge base (case library, knowledge graph, expert knowledge, etc.) is generated into document vectors through the same encoder Here, i represents the i-th document.
[0146] Determine the legal knowledge most relevant to the input information by calculating the cosine similarity between the input vector and each document vector:
[0147]
[0148] According to the calculated semantic similarity, the first k documents with the highest similarity are selected as the retrieval results. The retrieved legal knowledge is passed to an open-source large language model for further legal element identification and knowledge reasoning.
[0149] The multimodal information (text, voice, picture, etc.) input by the user is extracted through the encoder to generate the input vector v input , specifically:
[0150] The input I of image feature extraction is a H×W×C tensor, where H is the height, W is the width, and C is the number of channels. The convolution kernel K is used to perform a convolution operation on the image to extract local features from the image features, and an activation function is used to increase the nonlinear expression ability of the model:
[0151] O=I*K+b
[0152] O = max(0, i*K+b)
[0153] Among them, * represents the convolution operation and b is the bias term.
[0154] Use maximum pooling or average pooling to reduce the size of the feature map and retain important features:
[0155] O′=Pooling(O)
[0156] Flatten the pooled feature map and pass it through the fully connected layer to get the feature vector of the image:
[0157] f image =W image ·flatten(O′)+b image
[0158] Text feature extraction: Each word t i Convert to word vector e i :
[0159] e i =W embed t i +b embed
[0160] Use the LSTM layer to process the word vector sequence and extract the context information of the text:
[0161] h i =LSTM(e i ,h i-1 )
[0162] Among them, h i is the hidden state of the ith word.
[0163] The last hidden state h of LSTM N Through the fully connected layer, the feature vector of the text is obtained:
[0164] f text =W text h N +btext
[0165] Speech feature extraction:
[0166] The input speech V is a sequence of M sampling points, where M is the length of the speech. The acoustic features of the speech are extracted using the Mel frequency cepstral coefficients:
[0167] MFCC=MFCC(V)
[0168] Use the LSTM layer to process the MFCC feature sequence and extract the timing information of the speech:
[0169] h i =LSTM(MFCC i ,h i-1 )
[0170] Among them, h t is the hidden state of the i-th sampling point.
[0171] The last hidden state h of LSTM M Through the fully connected layer, the feature vector of the speech is obtained:
[0172] f voice =W voice h M +b voice
[0173] Feature fusion: Introducing weight parameter α image , α text and α voice , normalize the weighted eigenvectors to ensure that the sum of the weights is 1:
[0174]
[0175] Concatenate the feature vectors of image, text, and speech together to form a multimodal feature vector:
[0176] f multi =Concat(α′ image f image , α′ text f text , α′ voice f voice )
[0177] The concatenated feature vectors are processed through the fully connected layer to obtain the final multimodal feature representation:
[0178] v input =W final f multi +b final
[0179] The same multimodal feature extraction is performed on the legal knowledge text retrieved from the knowledge base to obtain the feature vector of the retrieval result. Perform operations such as concatenation, fusion or weighted average of the user's input features and the features of the search results:
[0180]
[0181] The final input of the open source large language model is f combined .
[0182] Use open source large language models to identify legal elements and conduct knowledge reasoning, and output the identification results and retrieved legal knowledge to users.
[0183] In summary, this embodiment realizes the specific application of the legal big model technology through the above steps, and significantly improves the accuracy and efficiency of legal retrieval.
[0184] Through testing and verification on multi-type and multi-source data sets, the accuracy, precision and AUC value of the present invention have been significantly improved compared with traditional models.
[0185] Embodiment 2
[0186] The purpose of this embodiment is to provide a legal element identification system for knowledge retrieval, including:
[0187] An acquisition module is configured to: acquire target information to be processed;
[0188] A retrieval module is configured to: retrieve the target information to be processed through a knowledge base to obtain relevant legal knowledge;
[0189] The recognition module is configured to: extract features of the target information to be processed and the corresponding relevant legal knowledge respectively, splice the extracted features, input the spliced features into a trained open source large language model, and perform legal element recognition and knowledge reasoning on the target information to be processed through the open source large language model;
[0190] In the process of training the open source large language model, the maximum likelihood method is used to estimate the model parameters, the model parameters corresponding to maximizing the log-likelihood function are determined, and the gradient descent method is used to update the model parameters.
[0191] In further embodiments, there is also provided:
[0192] An electronic device includes a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method described in Embodiment 1 is performed. For the sake of brevity, no further description is given here.
[0193] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0194] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0195] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the method described in embodiment 1 is completed.
[0196] The method in the first embodiment can be directly embodied as a hardware processor, or a combination of hardware and software modules in the processor. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware. To avoid repetition, it will not be described in detail here.
[0197] A computer program product includes a computer program, and when the computer program is executed by a processor, the method described in the first embodiment is implemented.
[0198] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer executable instructions, such as instructions included in a program module, which are executed in a device on a real or virtual processor of the target to perform the process / method as described above. Typically, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of program modules can be combined or divided between program modules as needed. Machine executable instructions for program modules can be executed in local or distributed devices. In distributed devices, program modules can be located in local and remote storage media.
[0199] The computer program code for implementing the method of the present invention can be written in one or more programming languages. These computer program codes can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device, so that the program code, when executed by the computer or other programmable data processing device, causes the function / operation specified in the flow chart and / or block diagram to be implemented. The program code can be executed completely on a computer, partially on a computer, as an independent software package, partially on a computer and partially on a remote computer or completely on a remote computer or server.
[0200] In the context of the present invention, computer program codes or related data may be carried by any appropriate carrier to enable a device, apparatus or processor to perform the various processes and operations described above. Examples of carriers include signals, computer readable media, etc. Examples of signals may include electrical, optical, radio, acoustic or other forms of propagation signals, such as carrier waves, infrared signals, etc.
[0201] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with this embodiment can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0202] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.
Claims
1. A legal element identification method for knowledge retrieval, characterized in that: include: Get the target information to be processed; Retrieving the target information to be processed through the knowledge base to obtain relevant legal knowledge; Extract features of the target information to be processed and the corresponding relevant legal knowledge respectively, splice the extracted features, input the spliced features into a trained open source large language model, and use the open source large language model to perform legal element recognition and knowledge reasoning on the target information to be processed; In the process of training the open source large language model, the maximum likelihood method is used to estimate the model parameters, the model parameters corresponding to maximizing the log-likelihood function are determined, and the gradient descent method is used to update the model parameters.
2. A method for identifying legal elements for knowledge retrieval according to claim 1, characterized in that: Training the open source large language model includes: Obtaining a training data set; wherein the training data set includes historical legal cases with legal elements annotated, as well as laws and regulations; Pre-training the open source large language model using the historical legal cases annotated with legal elements; The open source large language model is fine-tuned using the laws and regulations.
3. The method for identifying legal elements for knowledge retrieval according to claim 1, characterized in that: The model parameters are estimated using the maximum likelihood method to determine the model parameters that maximize the log-likelihood function, and the model parameters are updated using the gradient descent method, specifically: Calculate the conditional probability of the predicted value given by the open source large language model for the input; Constructing a likelihood function based on the conditional probability, and taking the logarithm of the likelihood function to obtain a log-likelihood function; The model parameters are updated by using a gradient descent algorithm, and the model parameters are iteratively updated by maximizing the log-likelihood function until convergence.
4. The method for identifying legal elements for knowledge retrieval according to claim 1, characterized in that: Preprocess the historical text legal cases, and train the open source large language model with the preprocessed historical legal cases; the preprocessing of historical legal cases is specifically as follows: Convert historical textual legal case information into embedded representations; Using LSTM to perform word segmentation on the embedded representation; A Transformer encoder based on a multi-layer attention mechanism is used for encoding to capture the key elements in the legal text and the relationships between them.
5. The method for identifying legal elements for knowledge retrieval according to claim 1, characterized in that: Feature extraction is performed on the target information to be processed and the corresponding relevant legal knowledge respectively, and the extracted features are spliced, specifically: Extracting features of the target information to be processed according to different modalities, and concatenating feature vectors extracted from the target information to be processed of different modalities to obtain a first multimodal feature representation; Extract features of the retrieved relevant legal knowledge according to different modalities, and concatenate the feature vectors extracted from the relevant legal knowledge of different modalities to obtain a second multimodal feature representation; The first multimodal feature representation and the second multimodal feature representation are concatenated and fused to obtain the input of the trained open source large language model.
6. A legal element identification method for knowledge retrieval as claimed in claim 1, characterized in that: Also includes: A reward function is set to guide the open source large language model to pay more attention to the recognition accuracy of legal elements during the optimization process; and generative adversarial samples are introduced to update the model parameters of the open source large language model.
7. A legal element identification system for knowledge retrieval, characterized in that: include: An acquisition module is configured to: acquire target information to be processed; A retrieval module is configured to: retrieve the target information to be processed through a knowledge base to obtain relevant legal knowledge; The recognition module is configured to: extract features from the target information to be processed and the corresponding relevant legal knowledge respectively, splice the extracted features, input the spliced features into a trained open source large language model, and perform legal element recognition and knowledge reasoning on the target information to be processed through the open source large language model; In the process of training the open source large language model, the maximum likelihood method is used to estimate the model parameters, the model parameters corresponding to maximizing the log-likelihood function are determined, and the gradient descent method is used to update the model parameters.
8. An electronic device, characterized in that: The method comprises a memory and a processor and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method according to any one of claims 1 to 6 is completed.
9. A computer-readable storage medium, characterized in that: Used to store computer instructions, which, when executed by a processor, complete the method described in any one of claims 1 to 6.
10. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.
Citation Information
Cited By
Model generation method, computer program product, equipment and storage medium
CN120930781A