Judicial affairs processing method and device based on multi-task merging judicial big model
By combining the parameter modules of the judicial task model, a unified judicial model is formed, which solves the problems of resource waste and knowledge sharing difficulties caused by independent training of various judicial task models in the existing technology, and achieves efficient judicial affairs processing and flexible task expansion.
Patent Information
- Application Number
- CN202510191825.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-02-20
AI Technical Summary
In the field of judicial science and technology, various judicial tasks (such as case classification, legal provision recommendation, case similarity matching, etc.) usually require separate training of models, resulting in high computing and storage costs, and difficult to share knowledge between tasks, affecting overall efficiency.
By analyzing parameter modules for multiple task models, the parameter modules are divided into low-correlation parameter modules and high-correlation parameter modules according to task correlation. The parameters of various parameter modules are merged by static merging and dynamic merging to form a unified judicial model for multi-task inference.
It has achieved resource conservation and performance guarantee, improved the efficiency of judicial affairs handling, supported the handling of multiple judicial tasks, and had the ability to expand tasks to adapt to the changing needs of judicial scenarios.
Smart Images

Figure CN119692482B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of artificial intelligence technology, and more particularly to a method and device for handling judicial affairs based on a multi-task merged judicial big model. Background Art
[0002] The field of judicial technology involves a large number of complex and diverse tasks, such as case classification, legal provision recommendation, case similarity matching and judicial decision prediction. However, the above-mentioned businesses currently rely on a large amount of manual operations and professional knowledge. This manual processing mode has certain technical defects in the complex modern judicial system.
[0003] With the development of artificial intelligence technology, the field of judicial technology has gradually introduced intelligent models to assist in case handling. Early methods in the field of judicial technology that used artificial intelligence technology usually trained models separately for each task (i.e. case classification, legal provision recommendation, and case similarity matching, etc.). This not only increased computing and storage costs, but also easily led to difficulties in sharing knowledge between tasks, affecting overall efficiency. Summary of the invention
[0004] In order to solve the problems existing in the prior art, the embodiments of this specification provide a judicial affairs processing method and device based on a judicial big model of multi-task merging. By performing parameter module analysis on multiple task models, the parameter modules are divided into low-correlation parameter modules and high-correlation parameter modules according to task relevance, and the parameters of the low-correlation parameter modules of multiple task models are merged in a static merging manner, and the parameters of the high-correlation parameter modules of multiple task models are merged in a dynamic merging manner. Finally, the final judicial big model is obtained by integration and used for multi-task reasoning.
[0005] The specific technical solutions of the embodiments of this specification are as follows:
[0006] On the one hand, the embodiments of this specification provide a judicial affairs processing method based on a multi-task merged judicial big model, the method comprising:
[0007] Obtaining pending judicial matters;
[0008] The judicial affairs are input into a judicial big model obtained by merging multiple task models for analysis and processing to obtain the processing result of the judicial affairs, wherein the step of obtaining the judicial big model by merging multiple task models includes:
[0009] Analyzing the parameter modules in the model architecture of the plurality of task models to obtain a task relevance score for each type of parameter module, wherein the task models are pre-trained models for processing corresponding types of judicial affairs, and the model architectures of the task models are the same;
[0010] According to the task relevance score of each type of parameter module, determine the low-relevance parameter module and the high-relevance parameter module;
[0011] Merging the parameters of the low-correlation parameter modules of the multiple task models in a static merging manner to obtain first merged model parameters;
[0012] Merging the parameters of the high-correlation parameter modules of the multiple task models in a dynamic merging manner to obtain second merged model parameters;
[0013] The first merged model parameters and the second merged model parameters are integrated to obtain the judicial big model, and the judicial big model is used to handle all types of judicial affairs corresponding to each task model.
[0014] Furthermore, obtaining text data related to the target judicial institution and judicial technology applications further includes:
[0015] Acquire judicial text data related to the target judicial institution and the judicial technology application;
[0016] Distributing a questionnaire related to the judicial technology application to the staff of the target judicial institution, and obtaining text data of the feedback results of the questionnaire;
[0017] The judicial text data and the feedback result text data are used as the text data.
[0018] Furthermore, the parameter modules in the model architecture of the multiple task models include a multi-layer perceptron (MLP) module and a multi-head attention (ATT) module.
[0019] Further, according to the task relevance score of each type of parameter module, a low-relevance parameter module and a high-relevance parameter module are determined;
[0020] The task relevance scores of the MLP module and the ATT module are compared, and the module with a low task relevance score is used as a low-relevance parameter module, and the module with a high task relevance score is used as a high-relevance parameter module.
[0021] Furthermore, analyzing the parameter modules in the model architecture of the multiple task models to obtain the task relevance score of each type of parameter module further includes:
[0022] Calculating the distance between the parameters of each type of parameter module in each layer of the network structure of the task model and the parameters of the corresponding parameter module of the corresponding layer of the network structure of the pre-trained model corresponding to the task model, wherein the pre-trained model is a deep learning model trained based on a large-scale general data set, and each task model is obtained by training the pre-trained model based on the corresponding type of judicial affairs data processed by the task model;
[0023] For each type of parameter module, the distances of the parameter modules of this type in all layer network structures of all task models are averaged to obtain the task relevance score of the parameter modules of this type.
[0024] Furthermore, the parameters of the low-correlation parameter modules of the multiple task models are merged in a static merging manner, and the formula for obtaining the first merged model parameters is:
[0025] ;
[0026] Among them, Θ merge l,low Indicates the first merged model l The merged parameters of the low-correlation parameter modules in the layer network structure, Θ0 l,low Represents the pre-trained model l The parameters of the low correlation parameter module in the layer network structure, λ represents the hyperparameter, T represents the total number of task models, τ t l,low =θ t l,low - Θ0 l,low , Θ t l,low Indicates t The first task model l Parameters of low-correlation parameter modules in the layer network structure.
[0027] Furthermore, merging the parameters of the high-correlation parameter modules of the multiple task models in a dynamic merging manner to obtain the second merged model parameters further includes:
[0028] According to the formula Perform singular value decomposition on the parameters of the high-correlation parameter module of each task model, where Θ t l,high Indicates t The first task model l Parameters of high correlation parameter modules in the layer network structure, U t l,high Indicatest The first task model l The left singular vector of the parameters of the highly correlated parameter modules in the layer network structure, Σ t l,high Indicates t The first task model l The singular value matrix of the parameters of the highly correlated parameter modules in the layer network structure, V t l,high Indicates t The first task model l The right singular vectors of the parameters of the highly correlated parameter modules in the layer network structure;
[0029] Select the first singular value in the singular value matrix after sorting the singular values in descending order. k singular values and the column vectors of the left singular vectors and the column vectors of the right singular vectors corresponding to the selected singular values, forming k rank-one matrices, where the formula for each rank-one matrix is: , where Θ t,j l,rank1 Indicates t The first task model l The first layer of the high correlation parameter module in the network structure j A rank-one matrix, 0< j ≤ k , u t,j l,high Indicates t The first task model l The first layer of the high correlation parameter module in the network structure j The column vector of left singular vectors corresponding to the singular values, s t,j l,high Indicates t The first task model l The first layer of the high correlation parameter module in the network structure j singular values, v t,j l,high Indicates t The first task model l The first layer of the high correlation parameter module in the network structure j The column vector of right singular vectors corresponding to the singular values;
[0030] According to the formula For each task model l The parameters of the high correlation parameter modules in the layer network structure are merged, where Θ merge l,high Represents the second merge modell The merged parameters of the high correlation parameter modules in the layer network structure, Θ0 l,high Represents the pre-trained model l Parameters of high correlation parameter modules in the layer network structure, α t,j l Indicates t The first task model l The first layer of the high correlation parameter module in the network structure j The weights of the singular values.
[0031] Furthermore, the first combined model parameters and the second combined model parameters are integrated to obtain the formula of the judicial macro model:
[0032] ;
[0033] Among them, Θ merge l The first part of the judicial model l A set of parameters for the layer network structure.
[0034] On the other hand, the embodiment of this specification also provides a judicial affairs processing device based on a multi-task merged judicial big model, the device comprising:
[0035] A judicial affairs acquisition unit, used for acquiring judicial affairs to be processed;
[0036] A judicial affairs processing unit is used to input the judicial affairs into a judicial big model obtained by merging multiple task models for analysis and processing, and obtain the processing result of the judicial affairs, wherein the step of obtaining the judicial big model by merging multiple task models includes:
[0037] Analyzing the parameter modules in the model architecture of the plurality of task models to obtain a task relevance score for each type of parameter module, wherein the task models are pre-trained models for processing corresponding types of judicial affairs, and the model architectures of the task models are the same;
[0038] According to the task relevance score of each type of parameter module, determine the low-relevance parameter module and the high-relevance parameter module;
[0039] Merging the parameters of the low-correlation parameter modules of the multiple task models in a static merging manner to obtain first merged model parameters;
[0040] Merging the parameters of the high-correlation parameter modules of the multiple task models in a dynamic merging manner to obtain second merged model parameters;
[0041] The first merged model parameters and the second merged model parameters are integrated to obtain the judicial big model, and the judicial big model is used to handle all types of judicial affairs corresponding to each task model.
[0042] On the other hand, an embodiment of the present specification further provides a computer device, including a memory, a processor, and a computer program stored in the memory, and the processor implements the above method when executing the computer program.
[0043] On the other hand, an embodiment of the present specification further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program implements the above method when executed by a processor.
[0044] By using the embodiments of this specification, independent expert models trained on different judicial data sets or tasks can be efficiently merged into a unified judicial big model. This technical solution has shown significant advantages in terms of resource conservation and performance assurance, including the following aspects: (1) Efficiency improvement: The processing of various types of transactions is completed in one judicial big model, avoiding the hardware overhead and scheduling complexity caused by the parallel deployment of multiple task models. (2) Independent performance close to the original model: Although multiple independent task models are merged into one judicial big model, the embodiments of this specification ensure that the performance of the merged judicial big model is close to or even reaches the level of independent task models through technologies such as task correlation analysis and dynamic selection of rank-one experts. (3) Multi-task processing capability: The judicial big model of the embodiments of this specification can support multiple judicial tasks such as case classification, legal article recommendation, similar case retrieval and judgment prediction, covering the main intelligent needs in the judicial process. (4) Task expansion capability: The judicial big model supports the rapid integration of subsequent new tasks and can adapt to the ever-changing needs in judicial scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0046] Figure 1 It is a schematic diagram of an implementation system of a judicial affairs processing party based on a judicial big model of multi-task merging in an embodiment of this specification;
[0047] Figure 2 The figure shows a flowchart of a judicial affairs processing party based on a judicial big model with multi-task merging in an embodiment of this specification;
[0048] Figure 3 It is a schematic diagram of the model architecture of multiple task models in the embodiments of this specification;
[0049] Figure 4 The figure shows a flow chart of obtaining the parameter module analysis in the model architecture of the plurality of task models and obtaining the task relevance score of each type of parameter module in this specification;
[0050] Figure 5 The figure shows a comparison diagram of the average L2 distances between the ATT module and the MLP in the three ViT architectures in the embodiments of this specification;
[0051] Figure 6 It is a schematic diagram showing the comparison of the total parameters of the merged model under different numbers of tasks in the embodiments of this specification;
[0052] Figure 7 The figure shows a schematic diagram of the structure of a judicial affairs processing device based on a multi-task merged judicial big model in an embodiment of this specification;
[0053] Figure 8 The figure is a schematic diagram of the structure of a computer device in an embodiment of the present specification.
[0054]
Description of reference numerals
[0055] 101. Terminal;
[0056] 102. Server;
[0057] 701, Judicial Affairs Acquisition Unit;
[0058] 702. Judicial affairs processing unit;
[0059] 802. Computer equipment;
[0060] 804. Processing equipment;
[0061] 806. Storage resources;
[0062] 808, driving mechanism;
[0063] 810, input / output module;
[0064] 812. Input device;
[0065] 814. Output device;
[0066] 816. Presentation equipment;
[0067] 818. Graphical user interface;
[0068] 820, network interface;
[0069] 822, communication link;
[0070] 824. Communication bus. DETAILED DESCRIPTION
[0071] The following will be combined with the drawings in the embodiments of this specification to clearly and completely describe the technical solutions in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in the embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the embodiments of this specification.
[0072] It should be noted that the terms "first", "second", etc. in the description and claims of the embodiments of this specification and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the embodiments of this specification described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, device, product or equipment that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0073] It should be noted that the acquisition, storage, use, and processing of data in the technical solutions of the embodiments of this specification comply with the relevant provisions of national laws and regulations.
[0074] It should be noted that in the embodiments of this specification, certain software, components, models and other existing solutions in the industry may be mentioned, and they should be regarded as exemplary. Their purpose is only to illustrate the feasibility of implementing the technical solution of this application, but it does not mean that the applicant has or will necessarily use the solution.
[0075] like Figure 1 The diagram shows a schematic diagram of an implementation system of a judicial affairs processing party based on a multi-task merged judicial big model in an embodiment of the specification, including a terminal 101 and a server 102. The terminal 101 and the server 102 can communicate with each other through a network, and the network can include a local area network (Local Area Network, referred to as LAN), a wide area network (Wide Area Network, referred to as WAN), the Internet or a combination thereof, and is connected to a website, a user device (such as a computing device) and a back-end system.
[0076] The staff can input the judicial affairs to be processed into the server 102 through the terminal 101. The server 102 is deployed with a pre-merged judicial big model, which calls the judicial big model to process the judicial affairs to be processed, obtains the processing results, and provides the processing results to the staff through the terminal 101.
[0077] Among them, judicial affairs may include case classification (automatically identifying case categories based on the cause of the case and legal provisions, providing basic classification support for judges and lawyers), legal provision recommendation (automatically recommending the most relevant legal provisions based on the content of the case and the focus of the dispute, helping judges to quickly locate applicable regulations and reduce the risk of misjudgment or missed judgment), case similarity matching (retrieval of cases with similar characteristics to the current case in the historical case library, providing reference for judges to make judgments, and enhancing the consistency and fairness of judicial decision-making) and judicial decision prediction (combining case background information and historical precedents to predict the potential ruling results of the case and provide judges with data-driven decision-making assistance), etc.
[0078] Alternatively, the server 102 may be a node of a cloud computing system (not shown), or each server may be a separate cloud computing system including a plurality of computers interconnected by a network and operating as a distributed processing system.
[0079] In addition, it should be noted that Figure 1 What is shown is only one application environment provided by the embodiment of this specification. In actual application, other application environments may also be included, and this specification does not limit it.
[0080] In view of the problems existing in the prior art, the embodiments of this specification provide a judicial affairs processing method based on a judicial big model of multi-task merging. Figure 2 The flowchart of the judicial affairs processing party of the judicial big model based on multi-task merging in the embodiment of this specification is shown. This figure describes the process of processing the judicial affairs to be processed and building the judicial big model. The order of steps listed in the embodiment is only one way of executing the order of many steps and does not represent the only execution order. When the system or device product is executed in practice, it can be executed in sequence or in parallel according to the method shown in the embodiment or the accompanying drawings.
[0081] Specific as Figure 2 As shown, the method may include:
[0082] Step 201: Obtain pending judicial affairs;
[0083] Step 202: inputting the judicial affairs into a judicial macro-model obtained by merging multiple task models for analysis and processing, and obtaining the processing result of the judicial affairs;
[0084] The steps of merging multiple task models to obtain the judicial big model include:
[0085] Step 221: Analyze the parameter modules in the model architecture of the plurality of task models to obtain a task relevance score for each type of parameter module, wherein the task models are pre-trained models for processing corresponding types of judicial affairs, and the model architectures of the task models are the same;
[0086] Step 222: Determine low-relevance parameter modules and high-relevance parameter modules according to the task relevance score of each type of parameter module;
[0087] Step 2023: merging the parameters of the low-correlation parameter modules of the multiple task models in a static merging manner to obtain a first merged model parameter;
[0088] Step 2024: merging the parameters of the high-correlation parameter modules of the multiple task models in a dynamic merging manner to obtain second merged model parameters;
[0089] Step 2025: Integrate the first merged model parameters and the second merged model parameters to obtain the judicial big model, which is used to handle all types of judicial affairs corresponding to each task model.
[0090] Through the embodiments of this specification, first, the parameter modules of the task model to be merged are divided into low-correlation groups and high-correlation groups through task correlation analysis. The low-correlation group adopts a static merging method to combine parameters through fixed rules to maximize storage efficiency, while the high-correlation group uses a dynamic merging strategy to retain task-specific characteristics and improve the task adaptability of the model.
[0091] In the highly correlated group, task-specific rank-one experts are extracted through singular value decomposition (SVD), and a rank-one expert pool shared across tasks is constructed. The expert pool provides a shared parameter library with low storage overhead while retaining the core features of each task. In the dynamic merging stage, a dynamic mechanism based on a routing network is designed to select the most relevant rank-one experts for combination in real time according to the characteristics of the input instance.
[0092] In addition, by combining static merging with dynamic merging, both efficiency and flexibility are taken into account. In scenarios where storage and computing resources are limited, the design of rank-one experts and dynamic selection significantly reduces the parameter storage requirements, making it suitable for large-scale multi-task model merging. This solution achieves the best balance between performance and resource efficiency, solving the problems of insufficient performance of existing static methods and excessive resource consumption of dynamic methods.
[0093] Through the above technical solutions, the embodiments of this specification provide an efficient and flexible solution to the problem of multi-task model merging in the field of judicial technology.
[0094] In the embodiments of this specification, in the field of judicial technology, the main function of the task model is to provide a service model for different judicial scenarios, which solves specific judicial problems.
[0095] For example:
[0096] Task model A may be a case classification model, which can automatically identify the category of an input case, such as a criminal, civil, administrative, or economic dispute case.
[0097] Task model B can be a legal provision recommendation model, which automatically recommends the most relevant legal provisions or regulations based on the input case content (such as legal documents and evidence materials) to provide reference for judges and lawyers.
[0098] Task model C can be a case similarity matching model, which can retrieve the case most similar to the current case in the historical case database, provide reference precedents, and ensure consistency and fairness in processing.
[0099] The goal of the model merging scheme involved in the embodiments of this specification is to merge the case classification model A, the legal provision recommendation model B and the case similarity matching model C into a comprehensive judicial model M. The judicial model M can simultaneously complete tasks such as case classification, legal provision recommendation and case similarity matching, as well as draw on the knowledge of the other model.
[0100] According to one embodiment of this specification, Figure 3 As shown, the model architecture of the multiple task models includes an L-layer network structure, and the parameter modules of each layer of the network structure include a normalization (Norm) module, a multi-layer perceptron (MLP) module, and a multi-head attention (ATT) module (that is, Figure 3 The expert model A and the expert model B in the example both include L-layer network results, and each layer of the network structure includes a normalization layer, a multi-layer perceptron layer, and a multi-head attention layer). Since the ATT module and the MLP module contribute the main parameters in the model, the embodiment of this specification focuses on the task relevance analysis of these two modules.
[0101] Specifically, Figure 4 As shown, analyzing the parameter modules in the model architecture of the multiple task models to obtain the task relevance score of each type of parameter module further includes:
[0102] Step 401: Calculate the distance between the parameters of each type of parameter module in each layer of the network structure of the task model and the parameters of the corresponding parameter module of the corresponding layer of the network structure of the pre-trained model corresponding to the task model, wherein the pre-trained model is a deep learning model trained based on a large-scale general data set, and each task model is obtained by training the pre-trained model based on the corresponding type of judicial affairs data processed by the task model;
[0103] Step 402: for each type of parameter module, average the distances of the parameter modules of this type in all layer network structures of all task models to obtain the task relevance score of the parameter modules of this type.
[0104] In the embodiments of this specification, the MLP modules and ATT modules of the multiple task models to be merged are first analyzed, and the parameter modules are divided into low-correlation parameter modules and high-correlation parameter modules according to the task relevance, so as to ensure that the merged model can efficiently utilize the shared features while retaining task-specific knowledge. Specifically, the parameter changes of each task model are analyzed, and its L2 distance with the pre-trained model parameters is calculated, which is formalized as:
[0105]
[0106] in, d (Θ t l,p ) indicates the t The first task model l Layer network structure p The L2 distance of the parameter module, Θ t l,p Indicates t The first task model l Layer network structure p The parameters of the parameter module, Θ0 l,p Represents the pre-trained model l Layer network structure p The parameter module's parameter p Parameters of a parameter module.
[0107] Next, the L2 distances of each type of parameter module in all layers of all task models are averaged to obtain the task relevance score of each type of parameter module:
[0108]
[0109] in, d mean p Indicates pThe task relevance scores of parameter modules of different types are represented by T, L, and the total number of network structure layers of the task model, where the network structure of each task model is the same.
[0110] Therefore, if the parameters of a parameter module change greatly, it means that the parameter module needs to significantly adjust the general knowledge of the pre-trained model to adapt to the needs of a specific judicial task. In contrast, parameter modules with smaller changes usually retain the general characteristics of the pre-trained model and contribute less to the characteristics of specific judicial tasks.
[0111] Then, the task relevance scores of the MLP module and the ATT module are compared, and the module with low task relevance score is taken as the low relevance parameter module, and the module with high task relevance score is taken as the high relevance parameter module.
[0112] like Figure 5 As shown, the embodiment of this specification analyzes the parameter modules of three task models of different architectures, ViT-B / 32, ViT-B / 16 and ViT-L / 14, and obtains the task relevance score of each type of parameter module.
[0113] The model structure is the task relevance score of the MLP module in the task model of ViT-B / 32 (i.e. Figure 5 The L2 distance of the task vector in is 0.53, and the task relevance score of the ATT module is 0.41. Therefore, the MLP module is a high-relevance parameter module and the ATT module is a low-relevance parameter module.
[0114] In the task model with the model structure of ViT-B / 16, the task relevance score of the MLP module is 0.38, and the task relevance score of the ATT module is 0.29. Therefore, the MLP module is a high-relevance parameter module, and the ATT module is a low-relevance parameter module.
[0115] In the task model with the model structure of ViT-B / 14, the task relevance score of the MLP module is 0.63, and the task relevance score of the ATT module is 0.48. Therefore, the MLP module is a high-relevance parameter module, and the ATT module is a low-relevance parameter module.
[0116] Then, the parameters of the low-correlation parameter modules of the multiple task models are merged in a static merging manner to obtain the first merged model parameters, so as to maximize parameter sharing and storage efficiency. This operation retains the general characteristics of the pre-trained model and incorporates the shared knowledge of each task in a weighted manner. The parameters of the high-correlation parameter modules of the multiple task models are merged in a dynamic merging manner to obtain the first merged model parameters, which not only retains task-specific knowledge but also reduces parameter storage requirements.
[0117] Specifically, the parameters of the low-correlation parameter modules of the multiple task models are merged in a static merging manner, and the formula for obtaining the first merged model parameters is:
[0118]
[0119] Among them, Θ merge l,low Indicates the first merged model l The merged parameters of the low-correlation parameter modules in the layer network structure, Θ0 l,low Represents the pre-trained model l The parameters of the low correlation parameter module in the layer network structure, λ represents the hyperparameter, T represents the total number of task models, τ t l,low =θ t l,low - Θ0 l,low , Θ t l,low Indicates t The first task model l Parameters of low-correlation parameter modules in the layer network structure.
[0120] Merging the parameters of the high-correlation parameter modules of the multiple task models in a dynamic merging manner to obtain the second merged model parameters further includes:
[0121] According to the formula:
[0122]
[0123] Perform singular value decomposition (SVD) on the parameters of the high-correlation parameter modules of each task model;
[0124] Among them, Θ t l,high Indicates t The first task model l Parameters of high correlation parameter modules in the layer network structure, U t l,high Indicates t The first task model l The left singular vector of the parameters of the highly correlated parameter modules in the layer network structure, Σ t l,high Indicates t The first task model l The singular value matrix of the parameters of the highly correlated parameter modules in the layer network structure, V tl,high Indicates t The first task model l The right singular vectors of the parameters of the highly correlated parameter modules in the layer network structure;
[0125] Select the first singular value in the singular value matrix after sorting the singular values in descending order. k Singular values ( k is an empirical value or experimental value) and the column vectors of the left singular vectors and the column vectors of the right singular vectors corresponding to the selected singular values, forming k rank-one matrices, where the formula for each rank-one matrix is:
[0126]
[0127] Among them, Θ t,j l,rank1 Indicates t The first task model l The first layer of the high correlation parameter module in the network structure j A rank-one matrix, 0< j ≤ k , u t,j l,high Indicates t The first task model l The first layer of the high correlation parameter module in the network structure j The column vector of left singular vectors corresponding to the singular values, s t,j l,high Indicates t The first task model l The first layer of the high correlation parameter module in the network structure j singular values, v t,j l,high Indicates t The first task model l The first layer of the high correlation parameter module in the network structure j The column vector of right singular vectors corresponding to the singular values;
[0128] Optionally, each task model can be mapped to k Rank-one matrices (only singular vectors need to be stored) are stored in the expert pool p l,high In , a shared task-specific knowledge base is built to facilitate subsequent merging.
[0129] Then a routing network is used to dynamically select from the expert pool according to the input instance. p l,highThe most relevant rank-one experts are selected for merging to achieve dynamic adaptability at the task and sample levels. The routing network is a set of neural network parameters, which is defined as a single-layer fully connected layer in the embodiment of this specification. It is randomly initialized and then updated by gradient descent.
[0130] Specifically, the input of the routing network R(x) is the implicit feature x of the instance, the implicit feature of each layer refers to the input of the neural network layer, and the output is the weight assigned to each task expert α t,j l .
[0131] Then follow the formula:
[0132]
[0133] For each task model l The parameters of the highly correlated parameter modules in the layer network structure are merged;
[0134] Among them, Θ merge l,high Represents the second merge model l The merged parameters of the high correlation parameter modules in the layer network structure, Θ0 l,high Represents the pre-trained model l Parameters of high correlation parameter modules in the layer network structure, α t,j l Indicates t The first task model l The first layer of the high correlation parameter module in the network structure j The weights of the singular values.
[0135] Finally, all task models are integrated through static merging and dynamic merging and used for multi-task reasoning.
[0136] The first combined model parameter and the second combined model parameter are integrated to obtain the formula of the judicial big model:
[0137]
[0138] Among them, Θ merge l The first part of the judicial model l A set of parameters for the layer network structure.
[0139] During the inference phase, the routing network dynamically selects experts based on the input instances without the need for additional optimization steps.
[0140] Below, the embodiments of this specification use the models trained on 8 public data sets (SUN397, Cars, RESISC45, EuroSAT, SVHN, GTSRB, MNIST, DTD) to illustrate the benefits of the proposed solution in terms of performance and parameter efficiency. As shown in Tables 1 and 2, on the two architectures of ViT-B / 32 and ViT-L / 14, a variety of model merging methods are used to merge the expert models of eight tasks, and the performance values are tested. The results show that the method of the embodiment of this specification has significant performance improvements compared to static model merging methods (such as weight averaging, task arithmetic, and parameter symbol alignment merging TIES-Merging). At the same time, the performance of the method of the embodiment of this specification is very close to the weighted hybrid expert integration network WEMoE, and the performance gap is within 1%.
[0141] Table 1 Comparison of multi-task performance when merging eight ViT-B / 32 models
[0142]
[0143] Table 2 Comparison of multi-task performance when merging eight ViT-L / 14 models
[0144]
[0145] It is particularly noteworthy that Figure 6 As shown, the embodiment of this specification measures the parameter efficiency of retaining different ranks when performing SVD decomposition on ViT-B / 32 for highly correlated parameter modules. Wherein, individual represents an independent task model, WEMoE represents a weighted hybrid expert ensemble network, RankOne-MoE-16 represents that the model merging method of the embodiment of this specification merges the task model of the ViT-B / 32 architecture when the retained rank is 16, RankOne-MoE-32 represents that the model merging method of the embodiment of this specification merges the task model of the ViT-B / 32 architecture when the retained rank is 32, RankOne-MoE-64 represents that the model merging method of the embodiment of this specification merges the task model of the ViT-B / 32 architecture when the retained rank is 64, and RankOne-MoE-128 represents that the model merging method of the embodiment of this specification merges the task model of the ViT-B / 32 architecture when the retained rank is 128.
[0146] It can be seen that the parameter cost of the weighted hybrid expert ensemble network WEMoE is much higher than the method (RankOne-MoE) in the embodiment of this specification.
[0147] For example, in the ViT-B / 32 architecture, the total number of parameters of the method of the embodiment of this specification is reduced by about 81.45% compared with the weighted hybrid expert ensemble network WEMoE, which significantly reduces the storage and computing costs. This excellent parameter efficiency makes the method of the embodiment of this specification extremely practical in actual application scenarios with limited resources.
[0148] Based on the same inventive concept, the embodiment of this specification also provides a judicial affairs processing device based on a multi-task merged judicial big model. Figure 7 As shown, the device comprises:
[0149] A judicial affairs acquisition unit 701 is used to acquire judicial affairs to be processed;
[0150] The judicial affairs processing unit 702 is used to input the judicial affairs into the judicial big model obtained by merging multiple task models for analysis and processing, and obtain the processing result of the judicial affairs, wherein the step of obtaining the judicial big model by merging multiple task models includes:
[0151] Analyzing the parameter modules in the model architecture of the plurality of task models to obtain a task relevance score for each type of parameter module, wherein the task models are pre-trained models for processing corresponding types of judicial affairs, and the model architectures of the task models are the same;
[0152] According to the task relevance score of each type of parameter module, determine the low-relevance parameter module and the high-relevance parameter module;
[0153] Merging the parameters of the low-correlation parameter modules of the multiple task models in a static merging manner to obtain first merged model parameters;
[0154] Merging the parameters of the high-correlation parameter modules of the multiple task models in a dynamic merging manner to obtain first merged model parameters;
[0155] The first merged model parameters and the second merged model parameters are integrated to obtain the judicial big model, and the judicial big model is used to handle all types of judicial affairs corresponding to each task model.
[0156] Since the principle of solving the problem by the above device is similar to that of the above method, the implementation of the above system can refer to the implementation of the above method, and the repeated parts will not be repeated.
[0157] like Figure 8The structure diagram of the computer device of the embodiment of this specification is shown. The permeability calculation unit or the conductivity calculation unit in the present invention can be the computer device in this embodiment, which executes the method of the present invention. The computer device 802 may include one or more processing devices 804, such as one or more central processing units (CPUs), each of which may implement one or more hardware threads.
[0158] The computer device 802 may also include any storage resources 806 for storing any kind of information such as code, settings, data, and the like.
[0159] For example, without limitation, the storage resource 806 may include any one or a combination of the following: any type of RAM, any type of ROM, a flash memory device, a hard disk, an optical disk, etc.
[0160] More generally, any storage resource may use any technology to store information.
[0161] Further, any storage resource may provide volatile or non-volatile retention of information.
[0162] Further, any storage resources may represent fixed or removable components of computer device 802 .
[0163] In one embodiment, when the processing device 804 executes the associated instructions stored in any storage resource or combination of storage resources, the computer device 802 can perform any operation of the associated instructions. The computer device 802 also includes one or more drive mechanisms 808 for interacting with any storage resource, such as a hard disk drive mechanism, an optical disk drive mechanism, etc.
[0164] The computer device 802 may also include an input / output module 810 (I / O) for receiving various inputs (via input devices 812) and for providing various outputs (via output devices 814). A specific output mechanism may include a presentation device 816 and an associated graphical user interface (GUI) 818. In other embodiments, the input / output module 810 (I / O), the input device 812, and the output device 814 may not be included, and the computer device 802 may be used as a computer device in a network. The computer device 802 may also include one or more network interfaces 820 for exchanging data with other devices via one or more communication links 822. One or more communication buses 824 couple the components described above together.
[0165] The communication link 822 may be implemented in any manner, for example, through a local area network, a wide area network (e.g., the Internet), a point-to-point connection, etc., or any combination thereof. The communication link 822 may include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc. governed by any protocol or combination of protocols.
[0166] The embodiments of the present specification also provide a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program implements the above method when executed by a processor.
[0167] The embodiments of the present specification also provide a computer-readable instruction, wherein when a processor executes the instruction, the program therein causes the processor to execute the above method.
[0168] It should be understood that in the various embodiments of the present specification, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present specification.
[0169] It should also be understood that in the embodiments of this specification, the term "and / or" is only a description of the association relationship of the associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in the embodiments of this specification generally indicates that the associated objects before and after are in an "or" relationship.
[0170] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the embodiments of this specification can be implemented with electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of this specification.
[0171] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0172] In the several embodiments provided in the embodiments of this specification, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, or it can be an electrical, mechanical or other form of connection.
[0173] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of this specification.
[0174] In addition, each functional unit in each embodiment of the present specification can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of software functional units.
[0175] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of this specification is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the embodiment of this specification. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.
[0176] The embodiments of this specification use specific embodiments to illustrate the principles and implementation methods of the embodiments of this specification. The description of the above embodiments is only used to help understand the methods and core ideas of the embodiments of this specification. At the same time, for those skilled in the art, according to the ideas of the embodiments of this specification, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the embodiments of this specification.
Claims
1. A judicial affairs processing method based on a multi-task merged judicial big model, characterized in that: The method comprises: Obtain pending judicial matters; The judicial affairs are input into a judicial big model obtained by merging multiple task models for analysis and processing to obtain the processing result of the judicial affairs, wherein the step of obtaining the judicial big model by merging multiple task models includes: Analyzing the parameter modules in the model architecture of the plurality of task models to obtain a task relevance score for each type of parameter module, wherein the task models are pre-trained models for processing corresponding types of judicial affairs, and the model architectures of the task models are the same; According to the task relevance score of each type of parameter module, determine the low-relevance parameter module and the high-relevance parameter module; Merging the parameters of the low-correlation parameter modules of the multiple task models in a static merging manner to obtain first merged model parameters; Merging the parameters of the high-correlation parameter modules of the multiple task models in a dynamic merging manner to obtain second merged model parameters; The first merged model parameters and the second merged model parameters are integrated to obtain the judicial big model, and the judicial big model is used to handle all types of judicial affairs corresponding to each task model.
2. The method according to claim 1, characterized in that The parameter modules in the model architecture of the multiple task models include a multi-layer perceptron MLP module and a multi-head attention ATT module.
3. The method according to claim 2, characterized in that According to the task relevance score of each type of parameter module, determine the low-relevance parameter module and the high-relevance parameter module; The task relevance scores of the MLP module and the ATT module are compared, and the module with a low task relevance score is used as a low-relevance parameter module, and the module with a high task relevance score is used as a high-relevance parameter module.
4. The method according to claim 1, characterized in that: Analyzing the parameter modules in the model architecture of the multiple task models to obtain the task relevance score of each type of parameter module further includes: Calculate the distance between the parameters of each type of parameter module in each layer of the network structure of the task model and the parameters of the corresponding parameter module of the corresponding layer of the network structure of the pre-trained model corresponding to the task model, wherein the pre-trained model is a deep learning model trained based on a large-scale general data set, and each task model is obtained by training the pre-trained model based on the corresponding type of judicial affairs data processed by the task model; For each type of parameter module, the distances of the parameter modules of this type in all layer network structures of all task models are averaged to obtain the task relevance score of the parameter modules of this type.
5. The method according to claim 2, characterized in that: The parameters of the low-correlation parameter modules of the multiple task models are merged in a static merging manner, and the formula for obtaining the first merged model parameters is: ; Among them, Θ merge l,low represents the merged parameters of the low-correlation parameter modules in the first merged model layer l network structure, Θ0 l,low represents the parameters of the low-correlation parameter module in the l-th layer network structure of the pre-trained model, λ represents the hyperparameter, T represents the total number of task models, τ t l,low =θ t l,low -Θ0 l,low , Θ t l,low Represents the parameters of the low-correlation parameter module in the l-th layer network structure of the t-th task model.
6. The method according to claim 5, characterized in that Merging the parameters of the high-correlation parameter modules of the multiple task models in a dynamic merging manner to obtain the second merged model parameters further includes: According to the formula Perform singular value decomposition on the parameters of the high-correlation parameter module of each task model, where Θ t l,high represents the parameters of the high correlation parameter module in the l-th layer network structure of the t-th task model, U t l,high Represents the left singular vector of the parameters of the high-correlation parameter module in the l-th layer network structure of the t-th task model, Σ t l,high V represents the singular value matrix of the parameters of the high correlation parameter module in the l-th layer network structure of the t-th task model. t l,high Represents the right singular vector of the parameters of the high-correlation parameter module in the l-th layer network structure of the t-th task model; Select the first k singular values in the singular value matrix after the singular values are sorted in order from large to small, and the column vectors of the left singular vectors and the column vectors of the right singular vectors corresponding to the selected singular values, to form k rank-one matrices, where the formula of each rank-one matrix is: , where Θ t,j l,rank1 represents the jth rank-one matrix of the high-correlation parameter module in the lth layer network structure of the tth task model, 0<j≤k, u t,j l,high The column vector of the left singular vector corresponding to the jth singular value of the high correlation parameter module in the lth layer network structure of the tth task model, s t,j l,high represents the jth singular value of the high correlation parameter module in the lth layer network structure of the tth task model, v t,j l,high Represents the column vector of the right singular vector corresponding to the j-th singular value of the high correlation parameter module in the l-th layer network structure of the t-th task model; According to the formula The parameters of the highly correlated parameter modules in the l-th layer network structure of each task model are merged, where Θ merge l,high represents the merged parameters of the high correlation parameter modules in the lth layer network structure of the second merge model, Θ0 l,high Represents the parameters of the high-correlation parameter module in the l-th layer network structure of the pre-trained model, α t,j l Represents the weight of the jth singular value of the high correlation parameter module in the lth layer network structure of the tth task model.
7. The method according to claim 6, characterized in that The first combined model parameters and the second combined model parameters are integrated to obtain the formula of the judicial macro model: ; Among them, Θ merge l A set of parameters representing the lth layer network structure of the judicial large model.
8. A judicial affairs processing device based on a multi-task merged judicial big model, characterized in that: The device comprises: A judicial affairs acquisition unit, used for acquiring judicial affairs to be processed; A judicial affairs processing unit is used to input the judicial affairs into a judicial big model obtained by merging multiple task models for analysis and processing, and obtain the processing result of the judicial affairs, wherein the step of obtaining the judicial big model by merging multiple task models includes: Analyzing the parameter modules in the model architecture of the plurality of task models to obtain a task relevance score for each type of parameter module, wherein the task models are pre-trained models for processing corresponding types of judicial affairs, and the model architectures of the task models are the same; According to the task relevance score of each type of parameter module, determine the low-relevance parameter module and the high-relevance parameter module; Merging the parameters of the low-correlation parameter modules of the multiple task models in a static merging manner to obtain first merged model parameters; Merging the parameters of the high-correlation parameter modules of the multiple task models in a dynamic merging manner to obtain second merged model parameters; The first merged model parameters and the second merged model parameters are integrated to obtain the judicial big model, and the judicial big model is used to handle all types of judicial affairs corresponding to each task model.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Federal learning bird nest target detection algorithm method based on dynamic gradient encryption
CN117351193A
Fine adjustment method and device for large model in security field and readable storage medium
CN118468928A