Learning gene-based large legal model lightweight deployment method, storage medium and program product

By extracting legal learning genes from the Transformer architecture model to reconstruct a lightweight model, the problems of deploying large legal models and losing professional knowledge in resource-constrained environments are solved, enabling efficient deployment and rapid adaptation to legal and regulatory updates in grassroots courts and small law firms.

CN120995251APending Publication Date: 2025-11-21SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510869013.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing large-scale legal models are difficult to deploy in resource-constrained grassroots courts and small law firms due to their massive parameter scale. Furthermore, the compression of these models results in a significant loss of professional knowledge, affecting legal reasoning ability and the accuracy of legal application, and making it difficult to flexibly adapt to different judicial scenarios and updates to laws and regulations.

Method used

By extracting key parameters from a Transformer architecture model pre-trained from a large-scale legal corpus as legal learning genes, a lightweight model is reconstructed, and core professional capabilities are retained by adapting it to the required domain through linear mapping techniques and fine-tuning.

Benefits of technology

It enables deployment in resource-constrained environments, maintains the core knowledge and reasoning capabilities of the legal big data model, adapts to different judicial scenarios, quickly adapts to updates in laws and regulations, and improves the efficiency of judicial practice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120995251A_ABST
    Figure CN120995251A_ABST
Patent Text Reader

Abstract

The invention discloses a large legal model lightweight deployment method based on learning genes, a storage medium and a program product. The method comprises the following steps: (1) selecting a large-scale legal language model which is pre-trained on a large-scale legal corpus and is based on a Transform architecture as an ancestor model; (2) a plurality of key parameters are extracted from the ancestor model to serve as legal learning genes, and the higher the actual contribution degree of the parameters to the legal task performance is, the higher the probability of being selected as the legal learning genes is; (3) reconstructing a lightweight legal progeny model through a linear mapping technology based on the extracted legal learning gene; and (4) performing fine adjustment on the reconstructed progeny model by using the required legal field data to enable the reconstructed progeny model to adapt to the required legal field. The deployed legal model is few in parameter and high in accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, in particular to a legal large model lightweight deployment method based on learning genes, a storage medium and a program product. BACKGROUND

[0002] With the development of artificial intelligence technology, legal large models have shown great potential in legal retrieval, judicial auxiliary decision-making, contract review, and legal document generation. Currently, legal large models are mainly based on large-scale corpus pre-training, such as judicial decisions, laws and regulations, and law journal papers, and master legal concepts, rules of application of law articles, and judicial reasoning logic through deep learning. However, these legal large models usually contain tens of billions or even hundreds of billions of parameters, which are difficult to deploy directly in practical application scenarios such as courts, law firms, and legal aid centers due to limitations on computing resources.

[0003] Existing model compression methods such as knowledge distillation, pruning, and quantization often result in significant performance degradation, especially in legal fields that require high professionalism and accuracy. Currently, legal large models face the following main challenges: large parameter size, difficult to deploy in resource-limited environments such as grassroots courts and small law firms; legal professional knowledge loss after model compression, affecting legal reasoning ability and law article application accuracy; different legal application scenarios (such as criminal case analysis, civil dispute mediation, administrative litigation consultation, etc.) require models of different sizes and professional directions, which are difficult to adapt flexibly; legal regulations update frequently, and models need to efficiently adapt to new legal norms and judicial interpretations. Therefore, there is an urgent need for a lightweight deployment method that can retain the core professional capabilities of legal large models to meet the needs of different application scenarios in judicial practice. SUMMARY

[0004] To solve the problems in the prior art, the present application aims to provide a legal large model lightweight deployment method based on learning genes with fewer parameters and high accuracy, a storage medium and a program product.

[0005] To achieve the above-mentioned application purposes, the present application provides the following technical solutions:

[0006] A legal large model lightweight deployment method based on learning genes, comprising the following steps:

[0007] (1) Select a large-scale legal language model based on the Transformer architecture pre-trained on a large-scale legal corpus as an ancestor model;

[0008] (2) Extract several key parameters from the ancestor model as legal learning genes, wherein the higher the actual contribution of the parameters to the performance of the legal task, the higher the probability of being selected as legal learning genes;

[0009] (3) Based on the extracted legal learning genes, a lightweight legal descendant model is reconstructed through linear mapping technology;

[0010] (4) Fine-tune the reconstructed descendant model with the required legal domain data to adapt it to the required legal domain.

[0011] Further, step (2) specifically includes:

[0012] (2.1) The ancestor model is divided into multiple functional stages;

[0013] (2.2) According to the actual contribution of each parameter in each functional stage of the ancestor model to the performance of the legal task, the comprehensive importance score of each parameter is calculated;

[0014] (2.3) According to the preset parameter proportion of each functional stage, from the parameters of each functional stage, the parameters corresponding to the preset parameter proportion are selected according to the comprehensive importance score from high to low, as the legal learning genes.

[0015] Further, step (2.1) specifically includes:

[0016] (2.1.1) The Embedding layer and the first several layers of the Transformer in the ancestor model are divided into a legal concept understanding functional stage;

[0017] (2.1.2) The middle several layers of the Transformer in the ancestor model are divided into a statute retrieval functional stage;

[0018] (2.1.3) The last several layers of the Transformer in the ancestor model are divided into a legal reasoning functional stage;

[0019] (2.1.4) The output layer of the ancestor model is divided into a conclusion generation functional stage.

[0020] Further, step (2.2) specifically includes:

[0021] (2.2.1) On the legal task evaluation dataset D, the forward propagation and back propagation of the ancestor model are performed, and the basic importance value of each part of the model parameters of each functional stage of the ancestor model is calculated according to the following formula:

[0022]

[0023] θ j,k ∈θ j ,θ={θ j |j=1,...,J}

[0024] Where I(θ j,k ) represents θj,k the importance value of the jth functional stage, θ j,k represents the kth parameter of the jth functional stage in the ancestral model, J represents the number of functional stages, θ j represents the parameter of the jth functional stage in the ancestral model, and θ represents the set of model parameters of the ancestral model. represents the expected calculation of the sample pair (x, y) in the legal task evaluation data set D, y represents the label of the sample x, f(x; θ) represents the output of the ancestral model when the input sample x, and L() represents the loss function.

[0025] (2.2.2) The activation mean and variance of the model parameters of each part of the ancestral model are calculated according to the following formula:

[0026]

[0027] θ j,k ∈θ j , θ = {θ j |j = 1,..., J}

[0028] Where N represents the total number of samples; a i (θ j,k ) represents the activation value controlled by θ j,k in the forward propagation of the ith sample; represents the mean of the activation value controlled by θ j,k ; represents the variance of the activation value controlled by θ j,k ;

[0029] (2.2.3) The activation frequency of each parameter of the ancestral model is calculated according to the following formula:

[0030]

[0031] θ j,k ∈θ j , θ = {θ j |j = 1,..., J}

[0032] Where AF(θ j,k ) represents the frequency ratio of θ j,k being significantly activated in all samples, τ represents the activation threshold, and I represents the indicator function, which is 1 when the condition is true, and 0 otherwise.

[0033] (2.2.4) After normalizing multiple indicators, linearly weighted fusion is performed to obtain the final comprehensive importance score:

[0034]

[0035] Where I final (θ j,k ) represents the importance score of θj,k The comprehensive importance score of each legal learning gene, and a, b, and g represent corresponding weight coefficients.

[0036] Further, step (3) specifically comprises:

[0037] (3.1) obtaining the network layer where each legal learning gene is located from the ancestor model, and connecting them in order to form the offspring model structure;

[0038] (3.2) establishing a linear mapping matrix optimization model, the linear mapping matrix optimization model taking the minimum output difference between the offspring model and the ancestor model corresponding to the mapped parameter matrix of each legal learning gene as the objective function, the mapped parameter matrix of the legal learning gene being the matrix obtained by mapping the parameter matrix corresponding to the legal learning gene through the corresponding linear mapping matrix;

[0039] (3.3) solving the linear mapping matrix optimization model to obtain the linear mapping matrix of each legal learning gene;

[0040] (3.4) calculating the mapped parameter matrix of each legal learning gene according to the linear mapping matrix of each legal learning gene;

[0041] (3.5) assigning the mapped parameter matrix of each legal learning gene to the corresponding network layer in the offspring model structure to obtain the reconstructed offspring model.

[0042] Further, the linear mapping matrix optimization model is specifically:

[0043]

[0044] W d,r = S r · W a,r · P r

[0045] r∈Ω

[0046] In the formula, W a,r represents the parameter matrix of the legal learning gene r in the ancestor model, W d,r represents the mapped parameter matrix of the legal learning gene r in the offspring model, Ω represents the set of legal learning genes, S r and P r respectively represent the linear mapping matrix of the legal learning gene r to be optimized, f a (x; W a,r ) represents the output of the ancestor model at the network layer corresponding to the legal learning gene r when the input sample x is input, and f d (x; W d,r ) represents the output of the offspring model at the network layer corresponding to the legal learning gene r when the input sample x is input.

[0047] Further, step (4) specifically comprises:

[0048] (4.1) constructing a fine-tuning dataset of the required legal field;

[0049] (4.2) using the fine-tuning dataset to fine-tune the reconstructed descendant model with a small number of samples, and the loss function during fine-tuning is:

[0050] L total = L task + λ·L legal_knowledge

[0051] Wherein, L total is the total loss, L task is the loss function of the required legal field, L legal_knowledge is the legal knowledge preservation regularization term, and λ is the balance coefficient.

[0052] Further, the legal knowledge preservation regularization term is specifically:

[0053]

[0054] In the formula, N represents the total number of samples, x i represents the i-th sample, f a (x i ) and f d (x i ) respectively represent the outputs of the ancestor model and the descendant model when the input sample x i .

[0055] A computer storage medium having a computer program / instruction stored thereon, the computer program / instruction realizing the above method when executed by a processor.

[0056] A computer program product comprising a computer program / instruction, the computer program / instruction realizing the above method when executed by a processor.

[0057] Compared with the prior art, the present application has the following beneficial effects:

[0058] 1. Significantly reducing the number of model parameters and the demand for computing resources, realizing the deployment of a large legal model in resource-limited environments such as primary courts and small law firms;

[0059] 2. Retaining the core knowledge and reasoning ability of the large legal model, ensuring the accuracy of legal concept understanding, statute application and judicial reasoning;

[0060] 3. Supporting the construction of legal models of different scales and professional directions to adapt to different judicial practice scenarios;

[0061] 4. Only a small amount of samples are needed to quickly adapt to legal regulations updates and new cases, and improve the efficiency of judicial practice. BRIEF DESCRIPTION OF DRAWINGS

[0062] Figure 1 is a flowchart of a method for lightweight deployment of a large legal model based on learning genes provided by an embodiment of the present application. DETAILED DESCRIPTION

[0063] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application.

[0064] Embodiment one

[0065] The embodiments of the present application provide a method for lightweight deployment of a large legal model based on learning genes, as shown in the following steps. Figure 1

[0066] (1) Select a large-scale legal language model based on the Transformer architecture pre-trained on a large-scale legal corpus as an ancestor model.

[0067] Specifically, a large-scale language model based on the Transformer architecture pre-trained on a large-scale Chinese legal corpus (including laws and regulations, judicial interpretations, guiding cases, court judgments, etc.) is selected as an ancestor model, such as Chinese-Legal-BERT or Chinese Legal GPT model. The model has learned legal concept system, rules of application of legal provisions and judicial reasoning logic through large-scale legal text learning, and has good language representation ability.

[0068] (2) Extract several key parameters from the ancestor model as legal learning genes.

[0069] Among them, the higher the actual contribution of the parameter to the performance of the legal task, the higher the probability of being selected as a legal learning gene. The specific extraction steps are as follows:

[0070] (2.1) Divide the ancestor model into multiple functional stages.

[0071] The Transformer language model has a multi-layer stacked structure, and the model parameter quantity is large. Direct full participation in pruning or migration will result in high computational cost and slow reasoning speed. The purpose of this step is to divide the model into "functional stages" with relatively independent semantic functions, so as to better perform the subsequent parameter screening and model reconstruction steps. Specifically, the ancestor model is divided into a legal concept understanding functional stage, a legal provision retrieval functional stage, a legal reasoning functional stage and a conclusion generation functional stage, and the specific division steps are as follows:

[0072] ​(2.1.1) The first several layers of the Embedding layer and the Transformer in the ancestor model (such as the first layer to the Nth layer) are divided into a legal concept understanding function stage; the shallow layer of the Transformer is mainly good at low-level feature extraction and is suitable as an entry for semantic understanding to capture basic legal concepts in the text such as case description, statute expression, and judgment language, and to learn the context semantics of vocabulary, entity, and basic legal terminology;

[0073] (2.1.2) The middle several layers of the Transformer in the ancestor model (such as the N+1th layer to the Mth layer) are divided into a statute retrieval function stage; the middle layer of the Transformer is usually stronger in context modeling and is suitable for capturing paragraph-level matching relationships. Based on the input case, the statute retrieval function stage focuses on the model learning the matching relationship between "case→statute", simulating the process of quickly identifying relevant legal provisions when a judge or lawyer handles a case;

[0074] (2.1.3) The last several layers of the Transformer in the ancestor model (such as the M+1th layer to the Pth layer) are divided into a legal reasoning function stage; the deep network has stronger abstract reasoning ability and is suitable for complex logic modeling. The legal reasoning function stage, based on the legal facts and statutes understood in the previous stage, performs compliance judgment, causal logic modeling, role analysis, and other judicial reasoning tasks;

[0075] (2.1.4) The output layer of the ancestor model, such as the classification head (one or more fully connected layers) and the task-specific decoder Decoder, is divided into a conclusion generation function stage. The conclusion generation function stage generates the final legal task conclusion, such as judgment result classification, legal suggestion generation, etc., according to the semantic representation and reasoning results output by the previous stage.

[0076] (2.2) According to the actual contribution degree of each parameter in each function stage of the ancestor model to the performance of the legal task, the comprehensive importance score of each parameter is calculated.

[0077] The purpose of this step is to measure the actual contribution degree of each part of the parameter in the ancestor model to the performance of the legal task, and to calculate the comprehensive importance score of each parameter. The specific steps include:

[0078] (2.2.1) On the legal task evaluation dataset D, use the ancestor model to perform forward propagation and back propagation, and calculate the basic importance value of each part of the model parameter of each function stage of the ancestor model according to the following formula:

[0079]

[0080] θ j,k ∈θ j ,θ={θj |j=1,...,J}

[0081] where I(θ j,k ) denotes the base importance value of θ j,k , θ j,k denotes the k-th parameter of the j-th functional stage in the ancestor model (such as Token Embedding matrix, Q, K, V linear transformation matrix in attention layer, MLP matrix of feedforward neural network, etc.), J denotes the number of functional stages, and θ j denotes the parameter of the j-th functional stage of the ancestor model, and θ denotes the expected calculation of the sample pair (x, y) in the legal task evaluation data set D, y denotes the label of the sample x, f(x; θ) denotes the output of the ancestor model when the input sample x, and L() denotes the loss function;

[0082] (2.2.2) The activation mean and variance of the ancestor model are calculated according to the following formula:

[0083]

[0084] θ j,k ∈θ j ,θ={θ j |j=1,...,J}

[0085] where N denotes the total number of samples; a i (θ j,k ) denotes the activation value (such as the result after ReLU, GELU function activation) controlled by θ j,k in the forward propagation of the i-th sample; denotes the mean value of the activation value controlled by θ j,k ; denotes the variance of the activation value controlled by θ j,k , which measures the amplitude of the activation change;

[0086] (2.2.3) The activation frequency of each parameter of the ancestor model is calculated according to the following formula:

[0087]

[0088] θ j,k ∈θ j ,θ={θ j |j=1,...,J}

[0089] where AF(θ j,k ) denotes the activation frequency of θ j,kThe proportion of frequencies that are significantly activated in all samples, tau represents the activation threshold, which is usually set to the average or empirical value of the activation function, I represents the indicator function, which is 1 when the condition is true, and 0 otherwise;

[0090] (2.2.4) Linearly weighted fusion after normalization of multiple indicators to obtain the final comprehensive importance score:

[0091]

[0092] Where I final (θ j,k ) represents the comprehensive importance score of θ j,k , and α, β, γ represent the corresponding weight coefficients.

[0093] (2.3) According to the preset parameter proportion of each functional stage, from the parameters of each functional stage, the parameters corresponding to the preset parameter proportion are selected according to the comprehensive importance score from high to low, as the legal learning genes.

[0094] Where the preset parameter proportion of each functional stage can be determined according to patent experience, and the parameter proportion of the functional stage with high importance is set to be higher, and the parameter proportion of the functional stage with lower importance is set to be lower. For example, the parameter proportions of the four stages can be set to 15%, 25%, 25%, and 35%, respectively. Then, for each functional stage, the parameters are sorted according to the comprehensive importance score from high to low, and the highest parameters are selected as the legal learning genes, the number of which is the product of the total number of parameters of the corresponding functional stage and the preset parameter proportion. All legal learning genes are stored in the legal learning gene set Ω.

[0095] (3) Based on the extracted legal learning genes, a lightweight legal offspring model is reconstructed through linear mapping technology.

[0096] Specifically includes:

[0097] (3.1) Obtain the network layer where each legal learning gene is located from the ancestor model, and connect them in order to form the offspring model structure;

[0098] In the offspring model structure, the output of the previous layer is the input of the next layer, but in the offspring model structure established by (3.1), the output of each layer and the input of the next layer do not match in dimension, so linear mapping matrix is used for dimension matching in the following steps;

[0099] (3.2) establishing a linear mapping matrix optimization model, the linear mapping matrix optimization model taking the output difference between the offspring model and the ancestor model corresponding to the mapping parameter matrix of each legal learning gene as the objective function, the mapping parameter matrix of the legal learning gene being the parameter matrix of the legal learning gene corresponding to the mapping matrix of the corresponding linear mapping matrix; the linear mapping matrix optimization model is specifically:

[0100]

[0101] W d,r =S r ·W a,r ·P r

[0102] r∈Ω

[0103] In the formula, W a,r represents the parameter matrix of the legal learning gene r in the ancestor model, W d,r represents the mapping parameter matrix of the legal learning gene r in the offspring model, Ω represents the set of legal learning genes, S r and P r respectively represent the linear mapping matrix of the legal learning gene r to be optimized, f a (x; W a,r ) represents the output of the ancestor model at the network layer corresponding to the legal learning gene r when the input sample x is input, and f d (x; W d,r ) represents the output of the offspring model at the network layer corresponding to the legal learning gene r when the input sample x is input.

[0104] (3.3) solving the linear mapping matrix optimization model to obtain the linear mapping matrix of each legal learning gene; wherein, when starting the optimization calculation, the linear mapping matrix of general initialization is used as the initial value, and the final linear mapping matrix is finally obtained through the optimization process;

[0105] (3.4) according to the linear mapping matrix of each legal learning gene, the mapping parameter matrix of each legal learning gene is calculated; specifically, the formula W d,r =S r ·W a,r ·P r is used for calculation.

[0106] (3.5) assigning the mapping parameter matrix of each legal learning gene to the corresponding network layer in the offspring model structure to obtain the reconstructed offspring model.

[0107] (4) fine-tuning the reconstructed offspring model using the required legal field data to adapt it to the required legal field.

[0108] Step (4) specifically comprises:

[0109] (4.1) Constructing a fine-tuning dataset of the required legal field; such as criminal case analysis, civil dispute mediation, administrative litigation consultation, etc.

[0110] (4.2) Fine-tuning the reconstructed offspring model with a small number of samples using the fine-tuning dataset, and the loss function during fine-tuning is:

[0111] L total =L task +λ·L legal_knowledge

[0112] Wherein, L total is the total loss, L task is the loss function of the required legal field, commonly used such as cross-entropy loss function Or L legal_knowledge is the legal knowledge preservation regularization term, which ensures that key legal knowledge is not lost during fine-tuning, and can adopt a representation alignment loss strategy, specifically: x i represents the i-th sample, f a (x i ), f d (x i ) respectively represent the output of the ancestor model and the offspring model when the input sample x i , λ is a balance coefficient, which regulates the importance weight of the main task loss and the knowledge preservation loss.

[0113] Embodiment two

[0114] The embodiment of the application provides a storage medium containing a computer executable program, which is used to execute the method of embodiment one when executed by a computer processor.

[0115] The storage medium of the embodiments of the present application can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination thereof. More specific examples (non-exhaustive list) of the computer-readable storage medium include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus or device.

[0116] Of course, the storage medium provided by the embodiments of the present application includes computer executable programs, and the computer executable programs are not limited to the method operations described above, but can also perform related operations in the method provided by any embodiment of the present application.

[0117] Embodiment three

[0118] The embodiments of the present application also provide a computer product, such as an app on a mobile phone, a tablet, an installation program on a computer, etc. The product includes computer programs / instructions, which, when executed by a processor, implement the method described in embodiment one. The code of the computer executable program for executing the operation of the present application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" language or similar programming languages. The program code can be executed entirely on a user computer, partially on a user computer, as an independent software package, partially on a user computer and partially on a remote computer, or entirely on a remote computer or server. In the case involving a remote computer, the remote computer can be connected to the user computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, through the Internet by using an Internet service provider).

[0119] It should be understood that the above embodiments and descriptions described in the specification are only the principles, main features and advantages of the present application, and various changes and improvements can be made to the present application without departing from the spirit and scope of the present application, and these changes and improvements all fall within the scope of protection of the present application.

Claims

1. A lightweight deployment method for a large-scale legal model based on learning genes, characterized in that, Includes the following steps: (1) Select a large-scale legal language model based on the Transformer architecture that has been pre-trained on a large-scale legal corpus as the ancestor model; (2) Extract several key parameters from the ancestor model as legal learning genes. The higher the actual contribution of the parameter to the performance of the legal task, the higher the probability of it being selected as a legal learning gene. (3) Based on the extracted legal learning genes, a lightweight legal offspring model is reconstructed through linear mapping technology; (4) Fine-tune the reconstructed descendant model using the required legal domain data to adapt it to the required legal domain.

2. The lightweight deployment method for a large-scale legal model based on learning genes according to claim 1, characterized in that, Step (2) specifically includes: (2.1) Divide the ancestor model into multiple functional stages; (2.2) Calculate the comprehensive importance score of each parameter based on the actual contribution of each parameter in each functional stage of the ancestor model to the performance of the legal task; (2.3) Based on the preset parameter ratios of each functional stage, select parameters from the parameters of each functional stage according to their comprehensive importance scores from high to low, and use them as legal learning genes.

3. The lightweight deployment method for a large-scale legal model based on learning genes according to claim 2, characterized in that, Step (2.1) specifically includes: (2.1.1) The Embedding layer in the ancestor model and the first few layers of the Transformer are divided into the legal concept understanding function stage; (2.1.2) Divide the intermediate layers of the Transformer in the ancestor model into the legal provision retrieval function stage; (2.1.3) Divide the last few layers of the Transformer in the ancestor model into the legal reasoning function stage; (2.1.4) The output layer in the ancestor model is divided into the conclusion generation function stage.

4. The lightweight deployment method for a large legal model based on learning genes according to claim 2, characterized in that, Step (2.2) specifically includes: (2.2.1) On the legal task evaluation dataset D, perform forward and backward propagation using the ancestor model, and calculate the basic importance values ​​of each part of the model parameters for each functional stage of the ancestor model according to the following formula: i j,k ∈θ j ,θ={θ j |j=1,...,J} Among them, I(θ j,k ) represents θ j,k The basic importance value, θ j,k θ represents the k-th parameter of functional stage j in the ancestor model, J represents the number of functional stages, and θ j θ represents the parameter of the functional stage j of the ancestor model, and θ represents the set of parameters of the ancestor model. Let f(x; θ) represent the expectation calculation for sample pairs (x, y) in the legal task evaluation dataset D, where y represents the label of sample x, f(x; θ) represents the output of the ancestor model when sample x is input, and L() represents the loss function. (2.2.2) Calculate the activation mean and variance of each part of the ancestor model parameters according to the following formula: i j,k ∈θ j ,θ={θ j |j=1,...,J} Where N represents the total number of samples; a i (θ j,k ) represents the i-th sample being propagated by θ during forward propagation. j,k The activation value of the control; Represents θ j,k The average value of the activation value under control; Represents θ j,k Controlling the variance of activation values; (2.2.3) Calculate the activation frequency of each parameter of the ancestor model according to the following formula: i j,k ∈θ j ,θ={θ j |j=1,...,J} Wherein, AF(θ) j,k ) represents θ j,k The proportion of frequencies that are significantly activated in all samples, where τ represents the activation threshold and Ι represents the indicator function, which is 1 when the condition is met and 0 otherwise. (2.2.4) After normalizing multiple indicators, a linear weighted fusion is performed to obtain the final comprehensive importance score: Among them, I final (θ j,k ) represents θ j,k The overall importance score is represented by α, β, and γ, which represent the corresponding weighting coefficients.

5. The lightweight deployment method for a large-scale legal model based on learning genes according to claim 1, characterized in that, Step (3) specifically includes: (3.1) Obtain the network layer where each law learning gene is located from the ancestor model and connect them in chronological order to form the offspring model structure; (3.2) Establish a linear mapping matrix optimization model. The linear mapping matrix optimization model takes the minimum output difference between the descendant model and the ancestor model corresponding to the mapped parameter matrix of each legal learning gene as the objective function. The mapped parameter matrix of the legal learning gene is the matrix after mapping the parameter matrix corresponding to the legal learning gene through the corresponding linear mapping matrix. (3.3) Solve the linear mapping matrix optimization model to obtain the linear mapping matrix of each legal learning gene; (3.4) Based on the linear mapping matrix of each legal learning gene, calculate the mapped parameter matrix of each legal learning gene; (3.5) Assign the mapped parameter matrix of each legal learning gene to the corresponding network layer in the offspring model structure to obtain the reconstructed offspring model.

6. The lightweight deployment method for a large legal model based on learning genes according to claim 5, characterized in that, The specific optimization model for the linear mapping matrix is ​​as follows: W d,r =S r ·W a,r ·P r r∈Ω In the formula, W a,r W represents the parameter matrix of the law learning gene r in the ancestor model. d,r The mapped parameter matrix of the legal learning gene r in the offspring model, where Ω represents the set of legal learning genes, and S r P r Let f represent the linear mapping matrix of the legal learning gene r to be optimized, respectively. a (x;W a,r f represents the output of the ancestor model in the network layer corresponding to the legal learning gene r when inputting sample x. d (x;W d,r ) represents the output of the descendant model in the network layer corresponding to the legal learning gene r when inputting sample x.

7. The lightweight deployment method for a large legal model based on learning genes according to claim 1, characterized in that, Step (4) specifically includes: (4.1) Construct a fine-tuned dataset for the required legal domain; (4.2) The reconstructed descendant model is fine-tuned with a small number of samples using the fine-tuning dataset. The loss function during fine-tuning is: L total =L task +λ·L legal_knowledge Among them, L total For the total loss, L task For the loss function of the required legal domain, L legal_knowledge To maintain the regularization term for legal knowledge, λ is the balance coefficient.

8. The lightweight deployment method for a large legal model based on learning genes according to claim 7, characterized in that, The legal knowledge regularization term specifically refers to: In the formula, N represents the total number of samples, x i Let f represent the i-th sample. a (x i ), f d (x i ) represent the ancestor model and the descendant model respectively in the input sample x i Output at that time.

9. A computer storage medium storing computer programs / instructions thereon, characterized in that, The computer program / instructions, when executed by a processor, implement the method of any one of claims 1-8.

10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the method of any one of claims 1-8.