Diversified Judicial Instruction Data Selection Method and System Based on Large Model Perception
Through a diversified judicial directive data selection method based on large-model perception, a subset of data containing the most variety of class activation tags is selected, which solves the problem of insufficient data selection diversity and coverage in the prior art, and improves the instruction compliance ability of large language models in downstream tasks in judicial field.
Patent Information
- Application Number
- CN202510370950.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-27
AI Technical Summary
In the selection of large-scale judicial instruction training data, it is difficult to efficiently screen out a subset of data that maintains task diversity and fully covers the model activation mode, resulting in insufficient instruction compliance capabilities in downstream tasks in judicial field.
Using a diversified judicial instruction data selection method based on big model perception, a subset of data containing the most variety of types of activation tags is filtered by converting judicial instruction data into activation tags and using the proportion and frequency characteristics of each layer of the big model activation vector.
The large language model's ability to follow instructions in downstream tasks in multiple judicial fields is realized, ensuring better diversity and coverage of selected data subsets, and reducing computational and time costs.
Smart Images

Figure CN119886229B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of judicial instruction training data selection in large language models, and in particular to a method and system for diverse judicial instruction data selection based on large model perception. Background Art
[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] In practical applications, in order to enable the model to better follow judicial instructions and conform to human preferences, it is necessary to fine-tune the model using judicial instruction data. Therefore, the fine-tuning stage is a key link in improving the performance of the model in a specific field. During the fine-tuning process for the judicial field, the quality and diversity of judicial instruction data have an important impact on the performance of the model.
[0004] In the field of large model judicial instruction training data selection, the core challenge is to efficiently screen out a data subset from a large amount of judicial instruction data that not only maintains task diversity but also comprehensively covers the model activation patterns. Researchers have proposed various data selection methods. These methods can be divided into two categories: data perception methods and large model activation perception methods. Data perception methods first use a pre-trained language model (such as BERT) to obtain data representations, and then select diverse data through methods such as k-means clustering. These methods are independent of the model to be fine-tuned when obtaining data representations, and can only select semantically diverse data based on text features, and cannot use the model itself to guide the data selection process. Large model activation perception methods use a trainable large language model to calculate the necessity of each data point, so as to select diverse data beneficial to the model. These methods tend to select data that is difficult for the current model, but it is difficult to ensure the diversity and coverage of the data, which may lead to sub-optimal results.
[0005] In addition, some methods for selecting diverse data based on model gradients require additional gradient calculations, greatly increasing the computational amount and time cost. In terms of core set selection, the goal is to select a subset from all training data such that the performance of the model trained on this subset is similar to the model trained on the complete data set. In the field of large model training data selection, especially in the scenario of instruction tuning, the existing methods have the following technical problems when performing large model training data selection: (1) Data perception methods cannot utilize the internal activation information of the model, resulting in insufficient semantic diversity; (2) Large model activation perception methods tend to select high-difficulty data, resulting in incomplete coverage of judicial instruction types; (3) Gradient-based screening methods require additional calculations and are difficult to process large-scale judicial instruction data (such as millions of instruction-response pairs), with low efficiency. Summary of the Invention
[0006] To solve the technical problems existing in the above-mentioned background art, the present invention provides a method and system for selecting diverse judicial instruction data based on large model perception. The present invention efficiently screens out an optimal judicial instruction data subset with diversity and coverage from numerous large model judicial instruction fine-tuning training data, maximizing the improvement of the instruction following ability of the large language model in multiple downstream judicial tasks after training.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] The first aspect of the present invention provides a method for selecting diverse judicial instruction data based on large model perception.
[0009] A method for selecting diverse judicial instruction data based on large model perception includes:
[0010] Convert each obtained original judicial instruction data into several tokens, input the tokens into the large model for inference, obtain the activation vectors output by the activation function of each layer, and select the index set of the dimensions in the activation vectors where the activation values are greater than the set value; sort the activation values in the index set according to their magnitudes as the activation labels of the tokens, and take the largest activation value as the value of the activation label of the token; take the union of the activation labels of all tokens of the judicial instruction data to obtain the activation label of the judicial instruction data.
[0011] Obtain the activation vectors of all judicial instruction data in each layer of the large model, calculate the proportion of the dimensions in the activation vectors of each layer of the large model where the activation values are greater than the set value to obtain the average proportion of each layer of the large model; select the layer with the largest proportion, and filter out the activation labels of the judicial instruction data with low frequencies in the layer with the largest proportion to obtain the activation labels of the screened judicial instruction data.
[0012] Based on the activation labels of the screened judicial instruction data, select the judicial instruction data containing the most diverse types of activation labels to obtain the sampled judicial instruction data subset.
[0013] Further, based on the activation labels of the screened judicial instruction data, select the judicial instruction data containing the most diverse types of activation labels to obtain the sampled judicial instruction data subset; the method includes: counting the activation labels covered by the screened original judicial instruction data to obtain the activation label data set, enumerating all activation values corresponding to each activation label in the activation label data set, counting the judicial instruction data corresponding to each activation value, and selecting the judicial instruction data containing the most diverse types of activation labels; sort all the original judicial instruction data in descending order according to the number of types of activation labels contained in the judicial instruction data, and select several judicial instruction data ranked at the front to construct the sampled judicial instruction data subset.
[0014] Further, calculate the proportion of dimensions in the activation vector of each layer of the large model where the activation value is greater than a set value to obtain the average proportion of each layer of the large model. The method includes: summing and averaging the proportion of dimensions in the activation vector of each layer of the large model where the activation value is greater than the set value to obtain the average proportion of each layer of the large model.
[0015] Further, filter out the activation labels of low-frequency judicial instruction data in the layer with the largest proportion; which is represented by the following formula:
[0016] Based on the number of types of activation labels in each selected layer TN Calculate the filtering weight of each layer :
[0017]
[0018] Set a basic frequency filtering threshold , so the filtering threshold of each layer is:
[0019]
[0020] Filter out the activation labels that appear less than times in the corresponding layer, so as to obtain the activation labels of the screened judicial instruction data;
[0021] where, , is the number of types of activation labels included in the th layer among the selected layers.
[0022] Further, select the index set of dimensions in the activation vector where the activation value is greater than the set value; which is represented by the following formula:
[0023]
[0024] where, is the dimension index in the vector , represents the index of the dimension in the vector where the value is greater than 1, represents the index set, and 1 is the set value.
[0025] Further, use the largest activation value as the activation label value of the tag; which is represented by the following formula:
[0026]
[0027] where, is the sorting function, the parameter in is the sorting basis, that is, according to Sort the specified digital sequence in descending order and use the maximum activation value in the activation vector as the activation label value.
[0028] The second aspect of the present invention provides a diversified judicial instruction data selection system based on large model perception.
[0029] A diversified judicial instruction data selection system based on large model perception, comprising:
[0030] An activation label acquisition module configured to: convert each piece of acquired original judicial instruction data into a number of tokens, input the tokens into a large model for inference, obtain the activation vector output by the activation function of each layer, and select the index set of the dimensions in the activation vector where the activation value is greater than a set value; sort the activation values in the index set according to size as the activation label of the token, and use the maximum activation value as the value of the activation label of the token; take the union of the activation labels of all tokens of the judicial instruction data to obtain the activation label of the judicial instruction data;
[0031] An activation label screening module configured to: obtain the activation vectors of all judicial instruction data in each layer of the large model, calculate the proportion of the dimensions in the activation vectors of each layer of the large model where the activation value is greater than a set value to obtain the average proportion of each layer of the large model; select the layer with the largest proportion and filter out the activation labels of the judicial instruction data with low frequency in the layer with the largest proportion to obtain the activation labels of the screened judicial instruction data;
[0032] An activation label-based data sampling module configured to: select the judicial instruction data containing the most types of activation labels based on the activation labels of the screened judicial instruction data to obtain a subset of the sampled judicial instruction data.
[0033] The third aspect of the present invention provides a computer-readable storage medium.
[0034] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the method for selecting diversified judicial instruction data based on large model perception as described in the first aspect above.
[0035] The fourth aspect of the present invention provides a computer device.
[0036] A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the steps in the method for selecting diversified judicial instruction data based on large model perception as described in the first aspect above.
[0037] The fifth aspect of the present invention provides a computer program product or a computer program.
[0038] The present invention provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the method for selecting diverse judicial instruction data based on large model perception as described in the first aspect above.
[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0040] The present invention combines the advantages of data perception and large model activation perception methods, and uses the differences in the activation states of different data during the inference process of the large language model to distinguish different types of judicial instruction data. Specifically, the large language model first calculates the activation states of all neurons included in a judicial dataset, and then filters out a data subset covering all activation states as the core set for fine-tuning the large language model. Compared with the existing methods, the present invention extracts data features at the model level, ensuring that the selected subset of judicial instruction data has better diversity and coverage. In addition, the neuron activation state can be obtained through a single inference without additional training, so the present invention is efficient at the inference level, reducing the computational and time costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The accompanying drawings forming a part of this invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0042] Figure 1 is a flowchart of the method for selecting diverse judicial instruction data based on large model perception shown in the present invention;
[0043] Figure 2 is an example diagram of the method for selecting diverse judicial instruction data based on large model perception shown in the present invention;
[0044] Figure 3 is a structural diagram of the system for selecting diverse judicial instruction data based on large model perception shown in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] The present invention will be further described below in conjunction with the drawings and embodiments.
[0046] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0047] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should also be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0048] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of methods and systems according to various embodiments of the present disclosure. It should be noted that each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code may include one or more executable instructions for implementing the logical functions specified in each embodiment. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the flowchart and / or block diagram, and combinations of blocks in the flowchart and / or block diagram, can be implemented using a dedicated hardware-based system for performing the specified functions or operations, or can be implemented using a combination of dedicated hardware and computer instructions.
[0049] Embodiment 1
[0050] As Figure 1As shown in the figure, this embodiment provides a method for selecting diverse judicial instruction data based on large model perception. This embodiment takes the application of this method to a server as an example. It can be understood that this method can also be applied to a terminal, or to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, web servers, cloud communications, middleware services, domain name services, security services CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions in this regard. In this embodiment, the method includes the following steps:
[0051] Convert each piece of obtained original judicial instruction data into a number of tokens, input the tokens into the large model for inference, obtain the activation vectors output by the activation function of each layer, and select the index set of the dimensions in the activation vectors where the activation values are greater than the set value; sort the activation values in the index set according to their magnitudes as the activation labels of the tokens, and take the largest activation value as the value of the activation label of the token; take the union of the activation labels of all tokens of the judicial instruction data to obtain the activation label of the judicial instruction data.
[0052] Obtain the activation vectors of all judicial instruction data in each layer of the large model, calculate the proportion of the dimensions in the activation vectors of each layer of the large model where the activation values are greater than the set value to obtain the average proportion of each layer of the large model; select the layer with the largest proportion, and filter out the activation labels of the judicial instruction data with low frequencies in the layer with the largest proportion to obtain the activation labels of the filtered judicial instruction data.
[0053] Based on the activation labels of the filtered judicial instruction data, select the judicial instruction data containing the most diverse types of activation labels to obtain the sampled subset of judicial instruction data.
[0054] Now introduce the specific technical solution. The core of the present invention is to propose a method for selecting diverse judicial instruction data based on large model perception (Model-Aware Diverse Core Set Selection for Instruction Tuning, abbreviated as MADS) for judicial instruction data selection, so as to select a subset of judicial instruction data that can cover all activation states as the core set for fine-tuning large language models.
[0055] In some embodiments, each piece of the obtained original judicial instruction data is converted into a number of tokens, and the tokens are input into a large model for inference to obtain the activation vectors output by the activation function for each layer. The index set of the dimensions with activation values greater than the set value in the activation vectors is selected; the activation values in the index set are sorted by size to be used as the activation labels of the tokens, and the largest activation value is used as the value of the activation label of the token; the union of the activation labels of all the tokens of the judicial instruction data is taken to obtain the activation label of the judicial instruction data. This is specifically implemented through the following scheme:
[0056] At this stage, first define as the original judicial instruction data for large model training without screening. Input all the data in into the large language model for inference, and extract the neuron activation states triggered by each piece of judicial instruction data during the inference process.
[0057] Then, process the output of the activation function into activation labels. First, arbitrarily select an open-source large language model such as Llama-3-3B, Llama-3-8B, etc. Before inputting the judicial instruction data into the large language model for inference, the judicial instruction data needs to be input into the corresponding tokenizer to obtain multiple tokens output by the tokenizer , and the activation labels of each token are merged into the activation label of the judicial instruction data.
[0058] Assume that the large language model has L layers, and each layer contains a neural network sub-module with an activation function. The goal of the present invention is to obtain the output of this activation function.
[0059] During the inference process of the large model, L outputs from the activation function will be generated:
[0060]
[0061] Among them, is the selected open-source large language model, represents the th token in the input, represents the th token (i.e., ) in the th layer of to generate a d-dimensional vector, where d is the dimension of the output of the activation layer of the large language model.
[0062] After obtaining the activation function output vector of , process this vector to obtain the activation label The present invention only focuses on the strongly activated dimensions, and the numerical value of each dimension in the d-dimensional activation vector is the activation value of that dimension. The larger the activation value, the stronger the activation of that dimension. Therefore, the dimensions in the activation vector with activation values greater than 1 are obtained. Let represent the index set of the dimensions in the vector with values greater than 1:
[0063]
[0064] where j represents the index of the dimension in the vector , , represents the index of the dimension in the vector with a value greater than 1. For example , has a total of 5 dimensions, then j = 1, 2, 3, 4, 5; among them, the numbers of the 2nd, 4th, and 5th dimensions > 1, then: .
[0065] Considering that the multiple dimensions included in the index set have different activation levels, in order to retain this information in the activation label, the dimensions are sorted in descending order according to their corresponding activation values and used as the activation label. At the same time, the maximum activation value is used as the value of the activation label:
[0066]
[0067] where is the sorting function, the parameter in is the basis for sorting, that is, sorting in descending order according to the numerical sequence specified by represents the numerical value of the dimension in the vector j . For example , has a total of 5 dimensions, then j = 1, 2, 3, 4, 5; among them, the numbers of the 2nd, 4th, and 5th dimensions > 1, then: . is to sort , , then the sorting result of is {4, 5, 2}.
[0068] At the same time, the maximum activation value in the activation vector is used as the activation label The value. The purpose of recording this value is to ensure that in the subsequent data selection stage, the core set contains data with different activation levels of specific activation tags.
[0069]
[0070] Among them, is the activation tag value; is the dimension index in the vector .
[0071] Based on the above, the activation tag corresponding to is , where L is the total number of layers of the large language model, and the activation tag of the data is the union of the tags of the tokens it contains:
[0072]
[0073]
[0074] In this stage, the activation tags of all judicial instruction data in are extracted.
[0075] In some embodiments, the activation vectors of all judicial instruction data in each layer of the large model are obtained, the proportion of the dimensions in the activation vector of each layer of the large model whose activation values are greater than a set value is calculated to obtain the average proportion of each layer of the large model; the layer with the largest proportion is selected, and the activation tags of the low-frequency judicial instruction data in the layer with the largest proportion are filtered out to obtain the activation tags of the screened judicial instruction data; specifically, it is implemented through the following scheme:
[0076] The activation tag of the data is , which contains the activation situation of each token in the judicial instruction data in each layer. Considering the following two factors: (1) There may be a large overlap between the activation tags of adjacent layers. (2) Different layers of the language model capture different features.
[0077] Therefore, aiming at the hierarchical feature distribution characteristics unique to the large model training data, N layers are selected as evenly as possible from low to high, and the activation tags of these N layers are filtered out from . In addition, in order to further determine which layers to select, the following method is adopted: for the I-th layer, obtain The activation vectors of all judicial instruction data in the layer, then calculate the proportion of activation values greater than 1 in each activation vector, and finally obtain the average proportion value. This average value represents the proportion of strongly activated neurons in each layer to all neurons in the layer. Starting from layer 0, this proportion rises in a wave-like manner. The larger the proportion, the more information it may contain, and the layer corresponding to the peak indicates that the proportion of strongly activated neurons in it is locally the largest. Therefore, select the layer corresponding to the peak. After selecting the layer, filter out the low-frequency activation labels in each selected layer. This is because very low-frequency activation labels lack universality and may even be noise. First, count the number of types of activation labels in each selected layer, denoted as .
[0078] Then, based on calculate the filtering weight of each layer :
[0079]
[0080] Among them, is the number of types of activation labels contained in the th layer among the selected layers.
[0081] Set a basic frequency filtering threshold , so the filtering threshold of each layer is:
[0082]
[0083] Finally, filter out the activation labels that appear less than times in the corresponding layer, so as to obtain .
[0084] In some embodiments, based on the activation labels of the screened judicial instruction data, select the judicial instruction data containing the most types of activation labels to obtain a subset of the sampled judicial instruction data; specifically, it is implemented through the following scheme:
[0085] At this stage, the goal is to combine the key indicators in the field of large model training data selection, and select a subset from such that the judicial instruction data in can cover all the activation labels in and contain all the activation values of these labels. First, count the activation labels covered by . Specifically, take the union of the labels contained in all the judicial instruction data in , denoted as :
[0086]
[0087] For each activation tag in , first enumerate all the activation values corresponding to it, then further count which judicial instruction data corresponds to each activation value, and finally select a representative judicial instruction data from these judicial instruction data. When selecting the representative judicial instruction data, select the data with the highest complexity, that is, the judicial instruction data containing the most types of activation tags. Specifically, sort the judicial instruction data in descending order according to the number of types of activation tags contained in it, and select the judicial instruction data ranked first. Finally, the sampled judicial instruction data can be obtained.
[0088] As Figure 2 shown, for each activation tag in (corresponding to the "tag" in Figure 2 ), first enumerate all the activation values corresponding to it (corresponding to the "value" in Figure 2 ), then further count which judicial instruction data (corresponding to the "index" in Figure 2 , here the index of the instruction data represents the corresponding data) corresponds to each activation value, and finally select a representative judicial instruction data from these judicial instruction data. When selecting the representative judicial instruction data, select the data with the highest complexity, that is, the judicial instruction data containing the most types of activation tags (corresponding to the "number of tags" in Figure 2 ). Specifically, sort the judicial instruction data in descending order according to the number of types of activation tags contained in it, and select the judicial instruction data ranked first. Finally, the sampled judicial instruction data can be obtained.
[0089] Embodiment 2
[0090] As Figure 3 shown, this embodiment provides a diversified judicial instruction data selection system based on large model perception, including:
[0091] An activation tag acquisition module, which is configured to: convert each obtained original judicial instruction data into a number of tokens, input the tokens into the large model for inference, obtain the activation vectors output by each layer of the activation function, and select the index set of the dimensions where the activation values are greater than the set value in the activation vectors; sort the activation values in the index set according to their magnitudes, and use them as the activation tags of the tokens, and use the largest activation value as the value of the activation tag of the token; take the union of the activation tags of all the tokens of the judicial instruction data to obtain the activation tags of the judicial instruction data;
[0092] Activate the label screening module, which is configured to: obtain the activation vectors of all judicial instruction data in each layer of the large model, calculate the proportion of the dimensions with activation values greater than the set value in the activation vectors of each layer of the large model to obtain the average proportion of each layer of the large model; select the layer with the largest proportion, and filter out the activation labels of the judicial instruction data with low frequency in the layer with the largest proportion to obtain the activation labels of the screened judicial instruction data;
[0093] The data sampling module based on activation labels, which is configured to: based on the activation labels of the screened judicial instruction data, select the judicial instruction data containing the most types of activation labels to obtain the sampled subset of judicial instruction data.
[0094] In some embodiments, the data sampling module based on activation labels is specifically configured to: count the activation labels covered by the screened original judicial instruction data to obtain an activation label dataset. For each activation label in the activation label dataset, enumerate all the corresponding activation values, count the judicial instruction data corresponding to each activation value, and select the judicial instruction data containing the most types of activation labels; sort all the original judicial instruction data in descending order according to the number of activation label types contained in the judicial instruction data, and select several judicial instruction data at the front to construct the sampled subset of judicial instruction data.
[0095] In some embodiments, the activation label screening module is specifically configured to: sum and average the proportion of the dimensions with activation values greater than the set value in the activation vectors of each layer of the large model to obtain the average proportion of each layer of the large model.
[0096] In some embodiments, the activation label screening module is specifically configured to filter out the activation labels of the judicial instruction data with low frequency in the layer with the largest proportion by using the following formula:
[0097] Based on the number of activation label types in each selected layer TN Calculate the filtering weight for each layer :
[0098]
[0099] Set a basic frequency filtering threshold , so the filtering threshold for each layer is:
[0100]
[0101] Filter out the activation labels that appear less than times in the corresponding layer, so as to obtain ;
[0102] wherein, , An activation tag representing the screened judicial instruction data.
[0103] In some embodiments, the activation tag acquisition module is specifically configured to: select the index set of the dimensions in the activation vector whose activation values are greater than the set value by using the following formula:
[0104]
[0105] where is the dimension index in the vector , represents the index of the dimension in the vector whose value is greater than 1, represents the index set, and 1 is the set value.
[0106] In some embodiments, the activation tag acquisition module is specifically configured to: use the following formula to take the maximum activation value as the value of the activation tag of this mark:
[0107]
[0108] where is the sorting function, the parameter in is the sorting basis, that is, sort in descending order according to the digital sequence specified by , and use the maximum activation value in the activation vector as the activation tag value.
[0109] Embodiment III
[0110] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps in the method for selecting diverse judicial instruction data based on large model perception as described in Embodiment I above are implemented.
[0111] Embodiment IV
[0112] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps in the method for selecting diverse judicial instruction data based on large model perception as described in Embodiment I above are implemented.
[0113] Embodiment V
[0114] This embodiment provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the method for selecting diverse judicial instruction data based on large model perception described in Embodiment 1 above.
[0115] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) that contain computer-usable program code.
[0116] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one Figure 1 process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0117] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one Figure 1 process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0118] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0119] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0120] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for selecting diversified judicial instruction data based on large model perception, characterized in that: include: Each piece of raw judicial instruction data is converted into a number of tags, and the tags are input into the large model for inference to obtain the activation vector output by each layer of activation function, and the index set of the dimension whose activation value in the activation vector is greater than the set value is selected; Sort the activation values in the index set by size to use them as the marked activation labels, and use the largest activation value as the value of the marked activation label; take the union of all marked activation labels of the judicial instruction data to obtain the activation label of the judicial instruction data; Obtain the activation vectors of all judicial instruction data at each layer of the large model, calculate the proportion of dimensions whose activation values are greater than the set value in the activation vectors of each layer of the large model, so as to obtain the average proportion of each layer of the large model; select the layer with the largest proportion, and filter out the activation labels of the judicial instruction data with low frequency in the layer with the largest proportion, so as to obtain the activation labels of the filtered judicial instruction data; Based on the activation labels of the screened judicial instruction data, the judicial instruction data containing the most types of activation labels are selected to obtain a sampled judicial instruction data subset.
2. The method for selecting diversified judicial instruction data based on large model perception according to claim 1 is characterized in that: The activation labels based on the screened judicial instruction data select the judicial instruction data containing the most types of activation labels to obtain a sampled judicial instruction data subset; the method includes: counting the activation labels covered by the screened original judicial instruction data to obtain an activation label data set, for each activation label in the activation label data set, enumerating all activation values corresponding to it, counting the judicial instruction data corresponding to each activation value, and selecting the judicial instruction data containing the most types of activation labels; sorting all the original judicial instruction data in descending order according to the number of activation label types contained in the judicial instruction data, selecting several judicial instruction data ranked in the front, and constructing a sampled judicial instruction data subset.
3. The method for selecting diversified judicial instruction data based on large model perception according to claim 1 is characterized in that: The method comprises: calculating the proportion of dimensions in the activation vector of each layer of the large model whose activation values are greater than the set value to obtain the average proportion of each layer of the large model; and the method comprises: summing and averaging the proportion of dimensions in the activation vector of each layer of the large model whose activation values are greater than the set value to obtain the average proportion of each layer of the large model.
4. The method for selecting diversified judicial instruction data based on large model perception according to claim 1 is characterized in that: The activation labels of the judicial instruction data with low frequency in the layer with the largest proportion are filtered out; It is expressed by the following formula: Based on the number of active label categories in each selected layer TN Calculate the filter weights for each layer : Set a basic frequency filter threshold , so the filtering threshold of each layer for: Filter out the numbers that appear less than 's activation tag, thereby obtaining the activation tag of the filtered judicial instruction data; in, , The selected layer The number of active label types contained in the layer.
5. The method for selecting diversified judicial instruction data based on large model perception according to claim 1 is characterized in that: The index set of the dimension whose activation value in the selected activation vector is greater than the set value; It is expressed by the following formula: in, is a vector The dimension index in , Representation vector The indices of the dimensions whose median values are greater than 1, Represents an index set, 1 is the set value.
6. The method for selecting diversified judicial instruction data based on large model perception according to claim 1 is characterized in that: The maximum activation value is used as the value of the activation label of the mark; It is expressed by the following formula: in, is the sorting function, Parameters in As the sorting basis, that is, according to Sort the specified sequence of numbers in descending order, using the activation vector The maximum activation value in is taken as the activation label The value of .
7. A diversified judicial instruction data selection system based on large model perception, characterized in that: include: An activation label acquisition module is configured to: convert each acquired original judicial instruction data into a number of labels, input the labels into the large model for inference, obtain the activation vector output by each layer of activation function, and select the index set of the dimension whose activation value in the activation vector is greater than the set value; Sort the activation values in the index set by size to use them as the marked activation labels, and use the largest activation value as the value of the marked activation label; take the union of all marked activation labels of the judicial instruction data to obtain the activation label of the judicial instruction data; An activation label screening module is configured to: obtain activation vectors of all judicial instruction data at each layer of the large model, calculate the proportion of dimensions whose activation values are greater than a set value in the activation vectors of each layer of the large model, so as to obtain an average proportion of each layer of the large model; select the layer with the largest proportion, and filter out activation labels of judicial instruction data with low frequency in the layer with the largest proportion, so as to obtain activation labels of filtered judicial instruction data; The activation tag-based data sampling module is configured to: based on the activation tags of the screened judicial instruction data, select the judicial instruction data containing the most types of activation tags to obtain a sampled judicial instruction data subset.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in the method for selecting diversified judicial instruction data based on large model perception as described in any one of claims 1 to 6 are implemented.
9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the method for selecting diversified judicial instruction data based on large model perception as described in any one of claims 1-6 are implemented.
10. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements the steps in the method for selecting diversified judicial instruction data based on large model perception as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Judicial data processing system based on artificial intelligence
CN118839160A
Judicial domain class case retrieval method and system based on large model
CN119474409A