Model Training Method, Apparatus, Storage Medium, and Computer Device
By clustering the security task description and combining security reference data to generate Q&A data, it is used to fine-tune the security task model, solving the problem of poor accuracy in large models in specific security fields, achieving higher answer accuracy.
Patent Information
- Application Number
- CN202510459014.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The lack of high-quality data in the field in the prior art has led to poor accuracy in answers to large models in specific security fields.
By generating multiple security task descriptions, clustering and partitioning to obtain a security task description cluster, generating input instructions in combination with security reference data, using the generative model to generate question-and-answer data, and using it to fine-tune the security task model.
This improves the accuracy of the model's answers in a specific security field and solves the problem of poor answer accuracy caused by the lack of high-quality data.
Smart Images

Figure CN119988985B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a model training method, device, storage medium, and computer device. Background Art
[0002] As a key department in network security, the Security Operations Center (SOC) is responsible for detecting, analyzing, and responding to network security incidents using various tools, technologies, and processes. However, the center currently faces many severe challenges, such as a lack of security expertise, spending a large amount of time investigating alarms, and slow response to advanced threats. With the emergence of large models, it has become possible to utilize their powerful understanding and generation capabilities to assist in security operations. It should be noted that network security is a highly specialized field, and dealing with security issues requires profound professional knowledge and skills.
[0003] In related technologies, in order to build a question-and-answer large model suitable for the security operation scenario, it is particularly necessary to fine-tune the large model with domain-specific data. However, the prominent problem currently faced is the lack of high-quality domain data, and most of the existing large model fine-tuning technologies focus on the general domain and do not fully consider the construction of fine-tuning data for specific domains, resulting in poor accuracy of the model's answers for specific domains. Therefore, related technologies urgently need to propose a model training method to solve the above technical problems. Summary of the Invention
[0004] The main purpose of this application is to provide a model training method, device, storage medium, and computer device, which can fine-tune the model with question-and-answer data in a specific domain to improve the accuracy of the model's answers for specific domains.
[0005] In a first aspect, an embodiment of this application provides a model training method, including:
[0006] Generating a plurality of first security task descriptions, clustering and dividing the plurality of first security task descriptions to obtain a plurality of security task description clusters;
[0007] Obtaining a plurality of security reference data, and determining the similarity between each security reference data and any first security task description in each security task description cluster;
[0008] Determining the target security task description corresponding to each security reference data according to the similarity between each security reference data and the corresponding first security task description;
[0009] Generating a first input instruction based on each security reference data and the corresponding target security task description, where the first input instruction is an input prompt for commanding to generate question-and-answer data according to the security reference data and the corresponding target security task description;
[0010] Input the first input instruction into the trained generation model to generate Q&A data corresponding to the first input instruction;
[0011] Input the Q&A data into the trained security task model to fine-tune the trained security task model, and obtain a fine-tuned security task model.
[0012] In a second aspect, an embodiment of the present application provides a model training device, including:
[0013] A clustering and partitioning unit, configured to generate multiple first security task descriptions, perform clustering and partitioning on the multiple first security task descriptions, and obtain multiple security task description clusters;
[0014] A first determination unit, configured to obtain multiple security reference data, and determine the similarity between each piece of security reference data and any first security task description in each security task description cluster;
[0015] A second determination unit, configured to determine a target security task description corresponding to each piece of security reference data according to the similarity between each piece of security reference data and the corresponding first security task description;
[0016] A generation unit, configured to generate a first input instruction based on each piece of security reference data and the corresponding target security task description, where the first input instruction is an input prompt for commanding to generate Q&A data according to the security reference data and the corresponding target security task description;
[0017] An input unit, configured to input the first input instruction into the trained generation model to generate Q&A data corresponding to the first input instruction;
[0018] A fine-tuning unit, configured to input the Q&A data into the trained security task model to fine-tune the trained security task model, and obtain a fine-tuned security task model.
[0019] In a third aspect, an embodiment of the present application provides a storage medium. The computer-readable storage medium stores multiple instructions, and these instructions are suitable for being loaded by a processor to execute the model training method as described in any one of the above.
[0020] In a fourth aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the model training method as described in any one of the above is implemented.
[0021] In an embodiment of the present application, by generating a plurality of first security task descriptions, clustering and partitioning the plurality of first security task descriptions to obtain a plurality of security task description clusters; obtaining a plurality of security reference data, and determining the similarity between each security reference data and any one of the first security task descriptions in each security task description cluster; determining a target security task description corresponding to each security reference data according to the similarity between each security reference data and the corresponding first security task description; generating a first input instruction based on each security reference data and the corresponding target security task description, where the first input instruction is an input prompt for commanding to generate question-and-answer data according to the security reference data and the corresponding target security task description; inputting the first input instruction into a trained generation model to generate question-and-answer data corresponding to the first input instruction; inputting the question-and-answer data into a trained security task model to fine-tune the trained security task model to obtain a fine-tuned security task model. Compared with the related art, in which the accuracy of answers for a specific field is poor due to the lack of high-quality data in the field, in the embodiment of the present application, by clustering and partitioning a plurality of first security task descriptions to obtain a security task description cluster for a specific field, and combining reference data to generate a first input instruction, question-and-answer data for model training in a specific field is obtained through the generation model, and the trained security task model that needs to be fine-tuned is trained through the question-and-answer data, thereby improving the accuracy of the model's answers for a specific field.
[0022] Other features and advantages of the present disclosure will be described in the following specification, and in part, will be obvious from the specification, or will be understood by implementing the present disclosure. The objectives and other advantages of the present disclosure can be realized and obtained by the structures particularly pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of this specification. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0024] Figure 1 It is a schematic diagram of the scenario of the model training system provided by the embodiment of the present application.
[0025] Figure 2 It is a schematic flowchart of the model training method provided by the embodiment of the present application.
[0026] Figure 3 It is a schematic structural diagram of the model training device provided by the embodiment of the present application.
[0027] Figure 4 This is a schematic structural diagram of the computer device provided by the embodiment of the present application. Specific embodiments
[0028] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0029] It should be noted that in some processes described in the specification, claims and the above-mentioned drawings, there are multiple steps that appear in a specific order, but it should be clearly understood that these steps can be executed not in the order in which they appear in this article or in parallel. The step numbers are only used to distinguish different steps, and the numbers themselves do not represent any execution order. In addition, descriptions such as "first", "second" or "target" in this article are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.
[0030] Please refer to Figure 1 , Figure 1 This is a schematic diagram of the scenario of the model training system provided by the embodiment of the present application. It includes a terminal 140, the Internet 130, a gateway 120, a server 110, etc.
[0031] The terminal 140 includes, but is not limited to, a pre-configured laptop computer, or a tablet computer, a desktop computer, etc., which are electronic devices with the ability to submit data. In addition, it can be a single device or a collection of multiple devices. The terminal 140 can communicate with the Internet 130 in a wired or wireless manner to exchange data.
[0032] The terminal 140 refers to a computer system that can submit data to the server 110. Compared with ordinary terminals, the server 110 has higher requirements in terms of stability, security, performance, etc. The server 110 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a part (such as a virtual machine) allocated from a high-performance computer, a combination of parts (such as virtual machines) allocated from multiple high-performance computers, etc.
[0033] The gateway 120 is also called an internetwork connector or a protocol converter. The gateway realizes network interconnection at the transport layer and is a computer system or device that acts as a converter. The gateway is a translator between two systems that use different communication protocols, data formats or languages, or even completely different architectures. At the same time, the gateway can also provide filtering and security functions. The message sent by the terminal 140 to the server 110 must be sent to the corresponding server 110 through the gateway 120. The message sent by the server 110 to the terminal 140 must also be sent to the corresponding terminal 140 through the gateway 120.
[0034] The model training method of the embodiment of the present disclosure can be implemented on the server 110 .
[0035] It should be noted that Figure 1 The scenario diagram of the model training system shown is only an example. The model training system and scenario described in the embodiment of the present application are intended to more clearly illustrate the technical solution of the embodiment of the present application, and do not constitute a limitation on the technical solution provided in the embodiment of the present application. A person of ordinary skill in the art can know that with the evolution of image processing technology and the emergence of new business scenarios, the technical solution provided in the embodiment of the present application is also applicable to similar technical problems.
[0036] In this embodiment, the model training device will be described from the perspective of the model training device, which can be specifically integrated into a computer device having a storage unit and a microprocessor installed therein and having computing capabilities.
[0037] See also Figure 2 , Figure 2 A flow chart of a model training method provided in an embodiment of the present application. The model training method comprises:
[0038] In step 201, a plurality of first safety task descriptions are generated, and the plurality of first safety task descriptions are clustered to obtain a plurality of safety task description clusters.
[0039] The first security task description refers to a textual description of a task in a specific security field, such as "detecting whether there is malware intrusion in the network", "evaluating the system's defense capabilities when it is attacked by DDoS", etc. It details the specific tasks that need to be completed in the security field. The security task description cluster is a set obtained by clustering multiple first security task descriptions. The first security task descriptions in the same cluster belong to the same security field. For example, the task descriptions in the security field such as network attack detection are divided into one cluster.
[0040] In some implementations, clustering the plurality of first security task descriptions to obtain a plurality of security task description clusters includes:
[0041] (1) Perform vector transformation on each of the first security task descriptions to obtain a vector representation of each of the first security task descriptions;
[0042] (2) Select a first preset number of candidate security task descriptions from each of the first security task descriptions as initial clustering centers for different initial security task description clusters;
[0043] (3) Calculate the distances between the vector representations of each of the other security task descriptions and the vector representations of each of the initial clustering centers;
[0044] (4) Divide each of the other security task descriptions into the initial security task description clusters corresponding to the initial clustering centers with the closest distances to obtain multiple updated security task description clusters;
[0045] (5) Determine the mean of the vector representations of the other security task descriptions in each of the updated security task description clusters and the vector representations of the corresponding initial clustering centers to obtain updated clustering centers for each of the updated security task description clusters;
[0046] (6) When the clustering division termination condition is not satisfied, determine the updated clustering centers as the initial clustering centers, determine the updated security task description clusters as the initial security task description clusters, and return to execute the step of calculating the distances between the vector representations of the other security task descriptions and the vector representations of each of the initial clustering centers until the clustering division termination condition is satisfied, and determine the multiple obtained updated security task description clusters as security task description clusters.
[0047] Among them, vector transformation is to convert the first security task description in text form into a numerical vector form that can be processed by a computer for subsequent distance calculation and clustering analysis. An embedding model (such as Word2Vec, BERT, etc.) is used to convert each task description into a vector representation. Suppose there are n task descriptions, denoted as , after being processed by the embedding model, each task description is converted into a -dimensional vector , where .
[0048] Specifically, the first preset quantity is a preset value used to determine the number of initial clustering centers selected. The setting of this quantity affects the clustering result and calculation efficiency, and usually needs to be reasonably adjusted according to the quantity and characteristics of the first security task description. The candidate security task description is selected from the first security task description set and serves as the security task description of the initial clustering center. The initial security task description cluster is, at the beginning of clustering, centered on each initial clustering center, and subsequently, other security task descriptions will be divided into these clusters to form a preliminary clustering result. The initial clustering center is the core of each initial security task description cluster. It is a vector representing the initial characteristics of the cluster. The other security task descriptions are the other first security task descriptions in the first security task description except those selected as the initial clustering centers. The distance is used to measure the similarity or difference degree between two vectors. Common distance measurement methods include Euclidean distance (Euclidean distance), etc.
[0049] Specifically, initialize the cluster and randomly select data points as the initial clustering centers, denoted as where , .
[0050] For each of the other security task descriptions , calculate its distance from each clustering center , usually using the Euclidean distance:
[0051] ;
[0052] where , is the distance between the other security task description and each clustering center . is the vector difference of the same vector dimension between the other security task description and each clustering center .
[0053] Based on the distances calculated in the previous step, for each of the other security task descriptions, find the initial clustering center with the closest distance to it, and then divide this description into the initial security task description cluster corresponding to this initial clustering center. After the division operation on all other security task descriptions, multiple updated security task description clusters are obtained. For each updated security task description cluster, aggregate (here it is to find the mean value) the vector representations of all other security task descriptions in it with the vector representation of the corresponding initial clustering center. In this way, calculate the updated clustering center of each updated security task description cluster.
[0054] Among them, the clustering division termination condition is a pre-set condition for judging whether the clustering process can end, and it is checked whether the current clustering result meets the clustering division termination condition. If it does not meet, the updated clustering center is used as the new initial clustering center, and the updated security task description cluster is used as the new initial security task description cluster, and then the loop is executed, that is, the distance between other security task descriptions and the new initial clustering center is recalculated, and the description division and clustering center update are performed. This process is continuously repeated until the clustering division termination condition is met. Finally, the multiple updated security task description clusters obtained when the termination condition is met are determined as the final security task description clusters.
[0055] Specifically, for each cluster, calculate the mean value of all data points in the cluster to update the clustering center. Let the th security task description cluster be , then the updated clustering center is:
[0056] ;
[0057] Among them, represents the number of data points in the security task description cluster , that is, the updated clustering center is determined by taking the mean value. Repeat the above steps, and the first security task description is divided into clusters, denoted as , and each cluster contains a group of first security task descriptions in the same security domain.
[0058] In this way, through clustering division, many first security task descriptions are divided into different security task description clusters according to similarity. The originally scattered and disordered security task descriptions are integrated to form organized clusters with specific characteristics. In the subsequent security task model training, the clustered security task description clusters can be used as high-quality training data. The task descriptions within the same cluster have similarity, and the model can better learn the patterns and characteristics of this type of task, improving the training effect and generalization ability of the model. For example, when training a model for detecting network attacks, using the security task description cluster related to network attack detection as training data, the model can focus more on learning the knowledge in this field, improving the accuracy and efficiency of detection.
[0059] In some embodiments, after obtaining the updated clustering center of each of the updated security task description clusters, it further includes:
[0060] (1) Calculate the difference value between the vector representation of the updated clustering center of each updated security task description cluster and the vector representation of the corresponding initial clustering center;
[0061] (2) Compare the difference value of each updated security task description cluster with a preset difference value;
[0062] (3) When there is an updated security task description cluster whose corresponding difference value is greater than or equal to the preset difference value, it is determined that the termination condition of clustering division is not satisfied;
[0063] (4) When the difference value of each updated security task description cluster is less than the preset difference value, it is determined that the termination condition of clustering division is satisfied.
[0064] Among them, for each updated security task description cluster, a suitable difference measurement method (such as Euclidean distance, cosine distance, etc.) is used to calculate the difference value between the vector representation of the updated cluster center and the vector representation of the corresponding initial cluster center. The preset difference value is a preset threshold used as a standard for judging the difference degree between the updated cluster center and the initial cluster center. The setting of this value usually needs to be adjusted according to actual clustering requirements and data characteristics. Compare the calculated difference value of each updated security task description cluster with the preset difference value.
[0065] Specifically, check the comparison results of the difference values of all updated security task description clusters with the preset difference value. If there is at least one updated security task description cluster whose difference value is greater than or equal to the preset difference value, it is considered that the current clustering result is not stable enough and the change of the cluster center is relatively large. At this time, it is determined that the termination condition of clustering division is not satisfied and clustering operations need to be continued. When the difference values of all updated security task description clusters are less than the preset difference value, it means that after this round of clustering operations, the difference between the updated cluster center and the initial cluster center of each cluster is relatively small and the cluster center has been relatively stable. At this time, it can be considered that the clustering result has achieved the expected effect, the termination condition of clustering division is satisfied, the clustering process can be stopped, and the multiple updated security task description clusters obtained currently are determined as the final security task description clusters.
[0066] Thus, by comparing the difference value between the updated cluster center and the initial cluster center and contrasting it with the preset difference value, it can effectively avoid prematurely stopping clustering when the cluster center is not yet stable. If such a judgment is not made, it may lead to inaccurate clustering results because the cluster center has not fully converged to a position that can accurately represent the data characteristics within the cluster. For example, when the difference value is large, it indicates that the cluster center is still changing to a large extent. If clustering is stopped at this time, some similar security task descriptions may be divided into different clusters, or different security task descriptions may be wrongly divided into the same cluster. Through this judgment mechanism, it is only when the difference values of all clusters are small that the cluster center is considered stable, thereby obtaining a more accurate and reliable clustering result.
[0067] In some embodiments, generating a plurality of first security task descriptions includes:
[0068] (1) Selecting a second preset number of security task descriptions from a first security task description set and a constructed second security task description set;
[0069] (2) Generating a second input instruction based on the selected security task descriptions, where the second input instruction is used to command to generate a third preset number of generated security task descriptions according to the selected security task descriptions;
[0070] (3) Inputting the second input instruction into the trained generation model to generate corresponding generated security task descriptions;
[0071] (4) Determining the generated security task descriptions that meet the security task description conditions as the first security task descriptions and allocating them to the first security task description set;
[0072] (5) When the total number of first security task descriptions in the first security task description set does not reach the fourth preset number, returning to execute the step of selecting a second preset number of security task descriptions from the constructed second security task description set and the first security task description set until the total number of first security task descriptions in the first security task description set reaches the fourth preset number, obtaining a plurality of first security task descriptions in the first security task description set.
[0073] Among them, the first security task description set is used to store the generated first security tasks, the initial set is an empty set, the second security task description set is a task set constructed by technicians according to the requirements of the security field, and the set includes a plurality of second security task descriptions under at least one security field. The second preset number is a preset value used to specify the total number of security task descriptions selected from the first security task description set and the second security task description set. Security task descriptions are randomly selected from the first security task description set and the second security task description set so that the total number of selected descriptions reaches the second preset number, and the number of security task descriptions selected from the second security task description set is more than the number of security task descriptions selected from the first security task description set.
[0074] Specifically, the second input instruction is an instruction generated based on the selected security task descriptions, and its function is to convey specific task requirements to the trained generation model, that is, to let the model use these selected descriptions as examples to generate new security task descriptions. For example, the selected security task descriptions are: Task 1, Task 2, and Task 3. Task 1 is: Threat entity extraction, and the task description is: Please extract entities related to network attack events such as threat organizations, malware, and attack means from the given text. It is required to extract comprehensive and accurate information and label the category of each entity. Task 2 and Task 3 are similar to Task 1 and will not be elaborated here. The third preset quantity is not less than 3, so the second input instruction generated based on Task 1, Task 2, and Task 3 is:
[0075] "Please generate no less than 3 tasks and task descriptions related to security tasks:
[0076] Task 1: Threat entity extraction
[0077] Task description: Please extract entities related to network attack events such as threat organizations, malware, and attack means from the given text. It is required to extract comprehensive and accurate information and label the category of each entity.
[0078] Task 2: XXX
[0079] Task 3: XXXX"
[0080] Thus, the generated second input instruction is input into a trained generation model such as GPT-4o, ChatGPT, etc. to generate a generated security task description corresponding to the second input instruction through the trained generation model.
[0081] Since the specific content of the generated security task description may be irrelevant to the security task, only the generated security task description that meets the security task description conditions is determined as the first security task description and assigned to the first security task description set. The total number of first security task descriptions in the first security task description set is detected. When the total number does not reach the pre-set fourth preset quantity, return to execute the step of selecting the second preset quantity of security task descriptions from the constructed second security task description set and the first security task description set until the total number of first security task descriptions in the first security task description set reaches the fourth preset quantity, and obtain multiple first security task descriptions in the first security task description set.
[0082] Therefore, in the field of security, high-quality task description data is often limited. By means of the trained generation model, the embodiments of the present application can generate new descriptions based on the existing security task descriptions, effectively expanding the scale of the first security task description set. As the number of iterations increases, the number of task descriptions in the set continues to increase, providing richer data support for subsequent security task analysis, model training, and other tasks. The generation model can expand, modify, and combine the selected security task descriptions to varying degrees, thereby generating diverse new descriptions. These new descriptions may cover different security scenarios, task types, and expression methods, making the data in the first security task description set more rich and diverse.
[0083] Moreover, with the development of security services, the scope of security tasks may continue to expand or be adjusted. This solution allows technicians to construct and update the second security task description set according to actual needs, enabling the first security task description set to flexibly adapt to changes in the task scope through the iterative generation process. For example, when an enterprise decides to expand its security services into new areas, it can add task descriptions of the new areas to the second security task description set, and then incorporate the relevant task descriptions into the first security task description set.
[0084] In addition, in the embodiments of the present application, security task description conditions are set to screen the generated descriptions, and only the descriptions that meet the conditions are determined as the first security task descriptions and added to the set. This screening mechanism can effectively exclude descriptions that are irrelevant to security tasks or of low quality, ensuring the quality of the data in the first security task description set.
[0085] In some embodiments, determining the generated security task description that meets the security task description conditions as the first security task description includes:
[0086] (1) Generating a third input instruction based on each generated security task description, where the third input instruction is used to command to answer whether the corresponding generated security task description is relevant to the security task;
[0087] (2) Inputting each third input instruction into the trained generation model to generate an answer result;
[0088] (3) Determining the generated security task description corresponding to the third input instruction with the answer result being "yes" as the first security task description.
[0089] Among them, the third input instruction is a specific instruction generated according to each generated security task description. Its main function is to enable the trained generation model to determine whether the description is relevant to the security task. This instruction clarifies that the task is to judge the relevance of the generated security task description. For each generated security task description, it is integrated into an instruction to form the third input instruction. Each generated third input instruction is sequentially input into the trained generation model. After receiving the instruction, the model understands and analyzes the instruction, and based on the knowledge and patterns it has learned, judges the relevance between the generated security task description and the security task, and generates a corresponding answer result.
[0090] Specifically, check the answer result corresponding to each third input instruction. For the case where the answer result is "yes", determine the generated security task description corresponding to this third input instruction as the first security task description. Then allocate these determined first security task descriptions to the first security task description set.
[0091] For example, the generated security task description is: Malicious IP and domain name identification, and the task description is to identify malicious IP addresses and domain names from the given network traffic or DNS query logs. According to the existing threat intelligence database, judge whether these IPs and domain names are related to known attackers or maliciousness. Then the third input instruction is:
[0092] "Please judge whether the following content is related to the security task:
[0093] Malicious IP and domain name identification
[0094] Task description: Identify malicious IP addresses and domain names from the given network traffic or DNS query logs. According to the existing threat intelligence database, judge whether these IPs and domain names are related to known attackers or maliciousness.
[0095] Please output "yes" for relevant and "no" for irrelevant.
[0096] Output: XXX"
[0097] Therefore, in the process of generating security task descriptions, some content that is not relevant or weakly relevant to security tasks may be generated. By having the trained generation model judge the relevance of each generated security task description, it is possible to effectively screen out the descriptions that are truly relevant to security tasks. Using the trained generation model to judge relevance realizes the automation of the judgment process. Compared with manually judging the relevance of generated security task descriptions one by one, this automated method greatly improves work efficiency and saves labor costs. The generation model can quickly process a large number of third input instructions and generate corresponding answer results, especially suitable for situations where the number of generated security task descriptions is large. In this example, "XXX" represents the answer result output by the generation model, specifically "yes" or "no".
[0098] In step 202, obtain multiple security reference data, and determine the similarity between each security reference data and any first security task description in each security task description cluster.
[0099] Among them, security reference data are data related to the security field, including network traffic data, system log data, security vulnerability report data, etc. These data are the basis for subsequent analysis and generation of Q&A data.
[0100] Specifically, collect security reference data from multiple channels such as att&ck, cve, cwe, and open-source threat intelligence, convert the security reference data and the first security task description into vector form, and then use a similarity calculation method (such as cosine similarity, etc.) to calculate the similarity between each security reference data and any first security task description in each security task description cluster. In this way, the similarity values between each security reference data and each first security task description can be obtained.
[0101] Specifically, obtain a dataset of security reference data through steps such as cleaning, long text partitioning, and deduplication , for each security reference data , sample one threat intelligence task description from each of the clusters in T, and calculate the similarity with .
[0102] In step 203, according to the similarity between each security reference data and the corresponding first security task description, determine the target security task description corresponding to each security reference data.
[0103] Among them, for each security reference data, obtain the s first security task descriptions with the highest similarity from the first security task descriptions (s is a parameter customized by the fine-tuning data constructor), and determine these s first security task descriptions as the corresponding target security task descriptions.
[0104] In step 204, based on each security reference data and the corresponding target security task description, a first input instruction is generated. The first input instruction is an input prompt for commanding to generate Q&A data according to the security reference data and the corresponding target security task description.
[0105] Among them, the first input instruction is a clear instruction, which is used to tell the subsequent generation model that it needs to generate Q&A data according to specific security reference data and the corresponding target security task description.
[0106] For example, the security reference data is the specific data content of network traffic data. According to the specific data content of the network traffic data and the target security task description of detecting whether there is an intrusion behavior of malware in the network, Q&A data on how to detect malware intrusion based on the traffic data is generated.
[0107] Specifically, the specific content of each security reference data and its corresponding target security task description are integrated, and the first input instruction is generated in a certain format. The format of the instruction can be designed according to actual needs, but usually it will clearly point out the input data and task description, and require to generate Q&A data.
[0108] Since the target security task description is a security task description in a specific field, the Q&A data generated according to the first input instruction is also Q&A data related to this specific field.
[0109] The first input instruction can be as shown in the following example:
[0110] "Task: Malicious IP and Domain Name Identification
[0111] Text: The security team found the following suspicious activities when analyzing network logs:
[0112] IP address 173.XXX initiated more than 10,000 connection requests to the internal server in the past 24 hours, and all were HTTPS requests. This IP also tried to access multiple common management background paths such as / admin, etc.
[0113] At the same time, it was found that the domain name hack.XXX was resolved to this IP address, and this domain name was newly registered 48 hours ago.
[0114] Query through the security platform shows that this IP address has been marked as a malicious scanning source by multiple security vendors.
[0115] Please generate a task-related question and answer based on the text content, in the format of:
[0116] {
[0117] Question:
[0118] Thought process:
[0119] Answer:
[0120] }"。
[0121] In step 205, the first input instruction is input into the trained generation model to generate Q&A data corresponding to the first input instruction.
[0122] Among them, the generated first input instruction is input into the trained generation model. The model understands and processes the first input instruction according to the knowledge and patterns it has learned, and generates corresponding Q&A data. During the generation process, the model will consider the input safety reference data and the target safety task description, and generate Q&A data that meets the requirements as much as possible.
[0123] In step 206, the Q&A data is input into the trained safety task model to fine-tune the trained safety task model, and a fine-tuned safety task model is obtained.
[0124] Among them, the trained safety task model is a model that has been trained to a certain extent for safety tasks, but it still needs to be fine-tuned for specific safety fields to improve its performance in specific safety fields. Fine-tuning techniques are not limited to LoRA, QLoRA, P-Tuning v2, Prefix Tuning, Prompt Tuning, etc., and open-source models are not limited to LLaMA, Falcon, BLOOM, ChatGLM, Baichuan, InternLM, etc. The specific selection is based on actual hardware resources and comparison of fine-tuning effects for optimization.
[0125] Specifically, the trained safety task model is fine-tuned with Q&A data related to specific fields, and finally a fine-tuned safety task model with better performance in specific safety fields is obtained.
[0126] In some embodiments, the Q&A data includes questions, thought processes, and answers. The step of inputting the Q&A data into the trained safety task model to fine-tune the trained safety task model to obtain a fine-tuned safety task model includes:
[0127] (1) Input the question into the trained safety model to obtain the first prediction probability of each token in the thought process and the second prediction probability of each non-stopword token in the answer;
[0128] (2) Determine the first loss value based on the first prediction probability of each token in the thought process and the second prediction probability of each non-stopword token in the answer;
[0129] (3) Fine-tune the trained security task model according to the first loss value to obtain a security task model after the first-stage fine-tuning;
[0130] (4) Input the question and the thinking process into the security task model with adjustable parameters after the first-stage fine-tuning to obtain the third prediction probability of each token in the answer predicted by the security task model with adjustable parameters after the first-stage fine-tuning, and the first distribution probability of the answer;
[0131] (5) Input the question and the thinking process into the security task model with frozen parameters after the first-stage fine-tuning to obtain the second distribution probability of the answer;
[0132] (6) Determine the second loss value based on the third prediction probability of each token in the answer, the first distribution probability of the answer, and the second distribution probability of the answer;
[0133] (7) Fine-tune the security task model with adjustable parameters after the first-stage fine-tuning according to the second loss value to obtain a fine-tuned security task model.
[0134] Among them, the question-and-answer data specifically includes: questions, thinking processes, and answers. A question is a specific question raised in the question-and-answer data; a thinking process is a step-by-step description of the analysis, reasoning, and thinking of the question; an answer is the content of the answer given to the question. A token is the basic unit in text processing, usually a word or a meaningful character segment.
[0135] The existing technical defects in the fine-tuning of current large models include: (1) When jointly training the thinking process and the answer, the model does not pay enough attention to the key answer information; (2) Directly removing the thinking process for fine-tuning will lead to the phenomenon of knowledge forgetting; (3) Traditional single loss functions cannot balance the quality control of long text generation. Therefore, the embodiments of the present application propose to adopt a staged loss function framework when fine-tuning and training the model, divide the training process into an inference enhancement stage (the first-stage fine-tuning) and a refinement and optimization stage (the second-stage fine-tuning), and improve the model's inference ability and output conciseness through a multi-stage differential training strategy. For the inference enhancement stage, it is mainly to improve the inference ability of the security task model; for the refinement and optimization stage, it is mainly to further improve the accuracy of the security task model for the answer on the premise that the inference ability of the security task model has been improved in the first-stage fine-tuning.
[0136] Specifically, for the first-stage fine-tuning, the input and output are defined as:
[0137] "Input sequence: X = [CLS] Question: {question} [SEP]
[0138] Target output: Y = [STA] Reasoning process: {reasoning} [SEP] Answer: {answer} ".
[0139] Among them, [CLS] is a classification token used to aggregate the feature information of the entire input sequence, [STA] is the sequence start token, [SEP] is the paragraph separator, and is the sequence termination token. Since the output of the first-stage fine-tuning includes two parts: the reasoning process and the answer, the loss function of the first-stage fine-tuning is composed of the sub-loss values of these two parts. For the sub-loss value of the reasoning process, it is determined by predicting the first prediction probability of each token in the reasoning process through the trained security model. For the sub-loss value of the answer, it is determined by predicting the second prediction probability of each non-stopword token in the reasoning process through the trained security model. Finally, based on the first prediction probability of each token in the reasoning process and the second prediction probability of each non-stopword token in the answer, the first loss value of the first-stage fine-tuning is determined, and thus the trained security task model is fine-tuned in the first stage according to the first loss value to obtain the security task model after the first-stage fine-tuning.
[0140] For the second-stage fine-tuning, the input and output are defined as:
[0141] "Input sequence: X' = [CLS] Question: {question} [SEP] Reasoning process: {reasoning}[SEP]
[0142] Target output: Y' = [STA] Answer: {answer} ".
[0143] The output of the second-stage fine-tuning only includes the answer part. However, since the second-stage fine-tuning needs to prevent the model from forgetting the knowledge learned previously during the process of optimizing the answer, a knowledge distillation sub-loss value is also required to balance answer generation and knowledge retention. Therefore, the loss function of the second-stage fine-tuning consists of the sub-loss values of these two parts: the answer and knowledge distillation. For the sub-loss value regarding the answer, the question and the thinking process are input into the safety task model with adjustable parameters after the first-stage fine-tuning, and the third prediction probability of each token in the answer predicted by the safety task model with adjustable parameters after the first-stage fine-tuning is used to determine it. For the knowledge distillation sub-loss value, the question and the thinking process are respectively input into the safety task model with adjustable parameters after the first-stage fine-tuning and the safety task model with frozen parameters after the first-stage fine-tuning, and the first distribution probability of the predicted answer and the second distribution probability of the predicted answer are obtained, and it is determined through the first distribution probability and the second distribution probability. Finally, based on the third prediction probability of each token in the answer, the first distribution probability of the answer, and the second distribution probability of the answer, the second loss value is determined, and the second loss value of the second-stage fine-tuning is determined. Thus, according to the second loss value, the trained safety task model is fine-tuned in the second stage to obtain the fine-tuned safety task model.
[0144] Therefore, in the inference enhancement stage (the first-stage fine-tuning), the input sequence is set to only contain the question, and the target output includes the thinking process and the answer. Such a setting makes the model more focused on reasoning starting from the question and generating a reasonable thinking process. Thus, the reasoning ability of the model is effectively improved, and the problem of over-focusing on key answer information and ignoring reasoning during the joint training of the thinking process and the answer is avoided. And since the loss function of the first-stage fine-tuning consists of the sub-loss values of the thinking process and the answer, by respectively determining the sub-loss value according to the first prediction probability of each token in the thinking process and the second prediction probability of each non-stopword token in the answer, the model can be guided to learn more comprehensively. This prompts the model to not only pay attention to the accuracy of the answer but also attach importance to the rationality of the thinking process, further strengthening the training of the reasoning ability.
[0145] In addition, in the refinement and optimization stage (the second-stage fine-tuning), the safety task model with frozen parameters after the first-stage fine-tuning is introduced. The question and the thinking process are input into this model and the safety task model with adjustable parameters, and the knowledge distillation sub-loss value is calculated to balance answer generation and knowledge retention. This can ensure that the model does not forget the knowledge learned in the first stage during the process of optimizing the answer. It guarantees that the generated answer not only meets the current task requirements but also retains the knowledge learned previously, effectively avoiding the phenomenon of knowledge forgetting caused by directly removing the thinking process for fine-tuning.
[0146] Through the synergy of two stages of fine-tuning, the first stage lays the reasoning ability and knowledge foundation for the second stage, and the second stage further optimizes the answer while retaining knowledge. This multi-stage training method enables the model to gradually learn and consolidate knowledge, improving the model's memory and application ability of knowledge.
[0147] In some embodiments, determining the first loss value based on the first prediction probability of each token in the thinking process and the second prediction probability of each non-stopword token in the answer includes:
[0148] (1.1) Obtain the position weight of each token in the thinking process;
[0149] (1.2) Determine the first sub-loss value of the thinking process based on the position weight of each token in the thinking process and the corresponding first prediction probability;
[0150] (1.3) Determine the answer sub-loss value based on the second prediction probability of each non-stopword token in the answer;
[0151] (1.4) Perform weighted summation on the first sub-loss value and the answer sub-loss value to obtain the first loss value.
[0152] Among them, the position weight is a numerical value used to represent the importance of each token in the thinking process at its position. Tokens in different positions have different impacts on the overall thinking logic and results, and the position weight is used to measure this impact degree. For example, tokens at the beginning of the thinking process may play an important role in guiding the direction, and their position weights may be relatively high; while tokens in the middle or at the end also have corresponding weight settings according to their roles in the reasoning process.
[0153] Specifically, the determination method of the position weight can refer to the following formula:
[0154] ;
[0155] Among them, is the length of the token sequence representing the thinking process, t represents the position index, and is used to traverse each token in the token sequence composed of multiple tokens in the thinking process. By calculating the position where the current token is located, is a hyperparameter for calculating the position weight and can be set to 0.2. Through this formula, the position weight of each token in the thinking process can be calculated.
[0156] Specifically, the first sub-loss value of the thinking process can be calculated with reference to the following loss function:
[0157]
[0158] Among them, represents the true token (or the target token) associated with the predicted probability of the token at position t during the thinking process. For example, when generating the sequence of the thinking process, the model has a predicted token distribution for each position t, which is the token that should actually appear at this position and is used to calculate the difference between the model prediction and the actual situation. represents the sequence composed of all tokens before position t during the thinking process. That is to say, it contains all token information from the start of the sequence to position t - 1. When the model predicts the token at position t , it will refer to the token information that has been generated previously, that is, , to calculate the predicted probability of the token at the current position , where is the input sequence, that is, the question. In this way, the model can use context information to make more accurate predictions.
[0159] Specifically, the sub - loss value of the answer can be calculated with reference to the following loss function:
[0160] ;
[0161] Among them, is the total number of all tokens composed of the thinking process and the answer. is the token that does not belong to the stop words, is the indicator function. When , = 1; when , = 0. is the true value corresponding to the token at position t in the answer part. When calculating the answer loss function , it represents the correct token that the model should predict at position t in the answer part and is used to measure the accuracy of the model's prediction of each token in the answer. represents the sequence composed of all tokens before position t when generating the answer. Similar to in the thinking process, but here it is for the answer part. When the model predicts the token at position t in the answer, it will calculate the predicted probability of based on the previously generated answer token information and the input sequence , that is,
[0162] After obtaining the first sub - loss value of the thinking process and the answer sub-loss value of the answer After that, calculate the first loss value of the first-stage fine-tuning through the following formula :
[0163]
[0164] where is the weight of the first sub-loss value and is the weight of the answer sub-loss value. Calculate the first loss value through weighted summation . .
[0165] In this way, by assigning position weights to the tokens in the thinking process, the importance of tokens at different positions in the thinking process can be highlighted, avoiding the model treating all tokens equally when processing the thinking process. Instead, it learns with emphasis according to position and importance, improving the quality of the thinking process. When calculating the answer sub-loss value, a stop-word masking mechanism is introduced (through the indicator function ), and the loss calculation is only performed on the prediction probabilities of non-stop words (entity words). This can prevent the model from over-focusing on stop words during training and instead concentrate its attention on entity words with actual semantics, enhancing the attention to entity words and helping to generate more meaningful and accurate answers. Similarly, based on the answer token information before position t and the input sequence X calculate the prediction probability of the answer token, and use the context information to guide the answer generation, enabling the model to generate more logical and semantic answers according to the previous information, improving the accuracy and relevance of the answers.
[0166] In some embodiments, determining the second loss value based on the third prediction probability of each token in the answer, the first distribution probability of the answer, and the second distribution probability of the answer includes:
[0167] (1.1) Determine the second sub-loss value based on the third prediction probability of each token in the answer;
[0168] (1.2) Determine the KL divergence between the first distribution probability of the answer and the second distribution probability of the answer to obtain the knowledge distillation sub-loss value;
[0169] (1.3) Perform weighted summation on the second sub-loss value and the knowledge distillation sub-loss value to obtain the second loss value.
[0170] Among them, the second sub-loss value is a numerical value used to measure the degree of difference between the accuracy of the token prediction in the answer by the adjustable-parameter security task model after the first-stage fine-tuning and the actual situation, and is a part of calculating the second loss value.
[0171] Specifically, the second sub-loss value can be calculated through the following loss function:
[0172] ;
[0173] Among them, is the input sequence for the second-stage fine-tuning, that is, the question and the thinking process, is the total number of tokens included in the answer, is a conditional probability, indicating that given the input sequence , the probability that the token at position t in the model-predicted answer is . represents the true token at position t in the answer.
[0174] Since it is necessary to avoid the phenomenon of knowledge forgetting caused by directly removing the thinking process for fine-tuning, the knowledge distillation sub-loss value needs to be determined by the KL divergence between the first distribution probability of the answer and the second distribution probability of the answer. The knowledge distillation sub-loss value can be calculated through the following formula:
[0175] ;
[0176] Among them, is the first distribution probability of the answer predicted by the security task model with adjustable parameters after the first-stage fine-tuning, are adjustable parameters, is the second distribution probability of the answer predicted by the security task model with frozen parameters after the first-stage fine-tuning, are frozen parameters. By determining the KL divergence between the first distribution probability and the second distribution probability of the answer, the knowledge distillation sub-loss value for preventing knowledge forgetting is obtained.
[0177] Among them, the second loss value can be calculated with reference to the following formula:
[0178] ;
[0179] Among them, is the weight value of the second sub-loss value , is the knowledge distillation sub-loss value The weight value is obtained by weighted summation of the second sub-loss value and the knowledge distillation sub-loss value to obtain the final second loss value .
[0180] Thus, the difference between the accuracy of the safety task model with adjustable parameters after the first-stage fine-tuning and the actual situation in predicting the tokens in the answer is measured by the second sub-loss value. By evaluating the prediction probability of each token, the model can more accurately learn the correct expression of each token in the answer. The knowledge distillation sub-loss value is determined by calculating the KL divergence between the first distribution probability and the second distribution probability. The first distribution probability comes from the model with adjustable parameters, and the second distribution probability comes from the model with frozen parameters. In this way, the model with adjustable parameters can refer to the knowledge of the model with frozen parameters during the process of optimizing the answer, avoiding the phenomenon of knowledge forgetting caused by directly removing the thinking process for fine-tuning. The KL divergence measures the difference between two probability distributions. By minimizing , the model with adjustable parameters can be as close as possible to the prediction distribution of the model with frozen parameters, thus retaining the knowledge learned by the model in the first stage. This enables the model to not lose the useful information accumulated previously when performing the second-stage fine-tuning due to excessive focus on the optimization of the current answer, ensuring the stability and generalization ability of the model.
[0181] In some embodiments, after obtaining the fine-tuned safety task model, it is necessary to evaluate and test the model performance of the fine-tuned safety task model. To ensure the fairness of the evaluation, a third-party evaluation set can be selected for evaluation. Different quantitative calculation methods are adopted for different types of tasks, and the influence of reference examples is fully considered. For the situation of evaluating the performance of the fine-tuned safety task model in the case of no-reference-example mode and with-reference-example mode, in the evaluation method with reference examples, reference examples are sampled from the data set of the same task in the evaluation set. When assembling the evaluation input prompt, the reference examples are input into the model to be evaluated as a part of it; for classification-type tasks, the accuracy needs to be calculated after rule matching according to the actual situation. For generation-type tasks, the similarity between the output of the fine-tuned model and the standard output is calculated. The higher the similarity, the better the generation effect. The pre-trained Transform model can be used to vectorize the two input texts to be compared respectively. The BERT model can map natural language texts to a high-dimensional vector space, enabling the semantic information of the texts to be fully expressed. In the embodiments of the present application, the SecBERT model fine-tuned in the field of network security can be selected. After obtaining the vector representations corresponding to the two texts, the cosine similarity between these two vectors is calculated. Specifically, let the vector representation of the output of the fine-tuned model be A, and the vector representation of the standard output be B. The formula is expressed as:
[0182] ;
[0183] wherein, represents the dot product of two vectors, and respectively represent the Euclidean norms of two vectors. The higher the similarity, the better the output effect of the fine-tuned model. In specific implementation, the following two points can be noted:
[0184] (1) Vector extraction strategy: The output vector corresponding to the [CLS] token can be used as the representation of the whole sentence or the average value of the output vectors of all tokens can be used. It is recommended to use the output features of the last layer or the last few layers.
[0185] (2) Calculation optimization: For long texts, it is recommended to perform segmented processing and then take the average value; it is recommended to normalize the vectors.
[0186] As can be seen from the above, in the embodiment of the present application, by generating a plurality of first security task descriptions, clustering and dividing the plurality of first security task descriptions to obtain a plurality of security task description clusters; obtaining a plurality of security reference data, and determining the similarity between each security reference data and any one of the first security task descriptions in each security task description cluster; according to the similarity between each security reference data and the corresponding first security task description, determining the target security task description corresponding to each security reference data; based on each security reference data and the corresponding target security task description, generating a first input instruction, where the first input instruction is an input prompt for commanding to generate question-and-answer data according to the security reference data and the corresponding target security task description; inputting the first input instruction into the trained generation model to generate the question-and-answer data corresponding to the first input instruction; inputting the question-and-answer data into the trained security task model to fine-tune the trained security task model to obtain a fine-tuned security task model. Compared with the related art, due to the lack of high-quality data in the field, the accuracy of answers for specific fields is relatively poor. In the embodiment of the present application, by clustering and dividing a plurality of first security task descriptions to obtain security task description clusters in a specific field, and combining reference data to generate a first input instruction, question-and-answer data for model training in a specific field is obtained through the generation model, and the trained security task model that needs to be fine-tuned is trained through the question-and-answer data, thereby improving the accuracy of the model's answers for specific fields.
[0187] For the specific implementation of each of the above steps, reference can be made to the previous embodiments, and details are not described herein again.
[0188] To facilitate better implementation of the model training method provided in the embodiments of the present application, the embodiments of the present application further provide a device based on the above model training method. The meanings of the nouns are the same as those in the above model training method, and the specific implementation details can be referred to the descriptions in the method embodiments.
[0189] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of the model training device provided in the embodiments of the present application. The model training device is applied to a computer device. The model training device may include a clustering and partitioning unit 601, a first determination unit 602, a second determination unit 603, a generation unit 604, an input unit 605, and a fine-tuning unit 606, etc.
[0190] The clustering and partitioning unit 601 is configured to generate a plurality of first security task descriptions, and perform clustering and partitioning on the plurality of first security task descriptions to obtain a plurality of security task description clusters;
[0191] The first determination unit 602 is configured to obtain a plurality of security reference data, and determine the similarity between each security reference data and any one of the first security task descriptions in each security task description cluster;
[0192] The second determination unit 603 is configured to determine the target security task description corresponding to each security reference data according to the similarity between each security reference data and the corresponding first security task description;
[0193] The generation unit 604 is configured to generate a first input instruction based on each security reference data and the corresponding target security task description. The first input instruction is an input prompt for commanding to generate question-and-answer data according to the security reference data and the corresponding target security task description;
[0194] The input unit 605 is configured to input the first input instruction into the trained generation model to generate question-and-answer data corresponding to the first input instruction;
[0195] The fine-tuning unit 606 is configured to input the question-and-answer data into the trained security task model to fine-tune the trained security task model to obtain a fine-tuned security task model.
[0196] In some embodiments, the clustering and partitioning unit 601 includes:
[0197] A conversion sub-unit is configured to perform vector conversion on each first security task description to obtain a vector representation of each first security task description;
[0198] A first selection subunit, configured to select a first preset number of candidate security task descriptions from each of the first security task descriptions as initial clustering centers for different initial security task description clusters;
[0199] A first calculation subunit, configured to calculate the distances between the vector representations of each of the other security task descriptions and the vector representations of each of the initial clustering centers;
[0200] A division subunit, configured to divide each of the other security task descriptions into the initial security task description clusters corresponding to the initial clustering centers with the closest distances, to obtain a plurality of updated security task description clusters;
[0201] A first determination subunit, configured to determine the means of the vector representations of the other security task descriptions in each of the updated security task description clusters and the vector representations of the corresponding initial clustering centers, to obtain updated clustering centers for each of the updated security task description clusters;
[0202] A first execution subunit, configured to, when the clustering division termination condition is not satisfied, determine the updated clustering centers as initial clustering centers, determine the updated security task description clusters as initial security task description clusters, and return to execute the step of calculating the distances between the vector representations of each of the other security task descriptions and the vector representations of each of the initial clustering centers, until the clustering division termination condition is satisfied, and determine the obtained plurality of updated security task description clusters as security task description clusters.
[0203] In some embodiments, the clustering division unit 601 further includes:
[0204] A second calculation subunit, configured to calculate the difference values between the vector representations of the updated clustering centers of each of the updated security task description clusters and the vector representations of the corresponding initial clustering centers;
[0205] A comparison sub-element, configured to compare the difference values of each of the updated security task description clusters with a preset difference value;
[0206] A second determination subunit, configured to, when there is an updated security task description cluster whose corresponding difference value is greater than or equal to the preset difference value, determine that the clustering division termination condition is not satisfied;
[0207] A third determination subunit, configured to, when the difference values of each of the updated security task description clusters are all less than the preset difference value, determine that the clustering division termination condition is satisfied.
[0208] In some embodiments, the clustering division unit 601 further includes:
[0209] A second selection subunit, configured to select a second preset number of security task descriptions from the first set of security task descriptions and the constructed second set of security task descriptions;
[0210] A generation subunit, configured to generate a second input instruction based on the selected security task descriptions, where the second input instruction is used to command to generate a third preset number of generated security task descriptions according to the selected security task descriptions;
[0211] A first input subunit, configured to input the second input instruction into the trained generation model to generate corresponding generated security task descriptions;
[0212] A fourth determination subunit, configured to determine the generated security task descriptions that meet the security task description conditions as the first security task descriptions and allocate them to the first set of security task descriptions;
[0213] A second execution subunit, configured to, when the total number of the first security task descriptions in the first set of security task descriptions does not reach the fourth preset number, return to execute the step of selecting a second preset number of security task descriptions from the constructed second set of security task descriptions and the first set of security task descriptions until the total number of the first security task descriptions in the first set of security task descriptions reaches the fourth preset number, so as to obtain a plurality of first security task descriptions in the first set of security task descriptions.
[0214] In some embodiments, the fourth determination subunit is configured to:
[0215] Generate a third input instruction based on each of the generated security task descriptions, where the third input instruction is used to command to answer whether the corresponding generated security task description is related to the security task;
[0216] Input each of the third input instructions into the trained generation model to generate an answer result;
[0217] Determine the generated security task descriptions corresponding to the third input instructions with the answer result being "yes" as the first security task descriptions.
[0218] In some embodiments, the question-and-answer data includes a question, a thinking process, and an answer. The fine-tuning unit 606 includes:
[0219] A second input subunit, configured to input the question into the trained security model to obtain a first prediction probability of each token in the thinking process and a second prediction probability of each non-stopword token in the answer;
[0220] A fifth determination subunit, configured to determine a first loss value based on the first prediction probability of each token in the thinking process and the second prediction probability of each non-stopword token in the answer;
[0221] A first fine-tuning subunit, configured to perform first-stage fine-tuning on the trained security task model according to the first loss value to obtain a first-stage fine-tuned security task model;
[0222] A third input subunit, configured to input the question and the thinking process into the security task model with adjustable parameters after the first-stage fine-tuning, to obtain the third prediction probability of each token in the answer predicted by the security task model with adjustable parameters after the first-stage fine-tuning, and the first distribution probability of the answer;
[0223] A fourth input subunit, configured to input the question and the thinking process into the security task model with frozen parameters after the first-stage fine-tuning, to obtain the second distribution probability of the answer;
[0224] A sixth determination subunit, configured to determine a second loss value based on the third prediction probability of each token in the answer, the first distribution probability of the answer, and the second distribution probability of the answer;
[0225] A second fine-tuning subunit, configured to perform second-stage fine-tuning on the security task model with adjustable parameters after the first-stage fine-tuning according to the second loss value to obtain a fine-tuned security task model.
[0226] In some embodiments, the fifth determination subunit is configured to:
[0227] Obtain the position weight of each token in the thinking process;
[0228] Determine a first sub-loss value of the thinking process based on the position weight of each token in the thinking process and the corresponding first prediction probability;
[0229] Determine an answer sub-loss value based on the second prediction probability of each non-stopword token in the answer;
[0230] Perform weighted summation on the first sub-loss value and the answer sub-loss value to obtain a first loss value.
[0231] In some embodiments, the sixth determination subunit is configured to:
[0232] Determine a second sub-loss value based on the third prediction probability of each token in the answer;
[0233] Determine the KL divergence between the first distribution probability and the second distribution probability of the answer to obtain a knowledge distillation sub-loss value;
[0234] Perform a weighted sum of the second sub-loss value and the knowledge distillation sub-loss value to obtain a second loss value.
[0235] For the specific implementation of each of the above units, reference may be made to the previous embodiments, which will not be elaborated herein.
[0236] As can be seen from the above, in the embodiment of the present application, the clustering division unit 601 generates a plurality of first security task descriptions, performs clustering division on the plurality of first security task descriptions to obtain a plurality of security task description clusters; the first determination unit 602 obtains a plurality of security reference data and determines the similarity between each security reference data and any first security task description in each security task description cluster; the second determination unit 603 determines the target security task description corresponding to each security reference data according to the similarity between each security reference data and the corresponding first security task description; the generation unit 604 generates a first input instruction based on each security reference data and the corresponding target security task description, and the first input instruction is an input prompt for commanding to generate question-and-answer data according to the security reference data and the corresponding target security task description; the input unit 605 inputs the first input instruction into the trained generation model to generate the question-and-answer data corresponding to the first input instruction; the fine-tuning unit 606 inputs the question-and-answer data into the trained security task model to fine-tune the trained security task model to obtain a fine-tuned security task model. Compared with the related art, in which the accuracy of answers for a specific domain is poor due to the lack of high-quality domain data, in the embodiment of the present application, a plurality of first security task descriptions are clustered and divided to obtain security task description clusters for a specific domain, and a first input instruction is generated in combination with reference data. The question-and-answer data for model training for a specific domain is obtained through the generation model, and the trained security task model that needs to be fine-tuned is trained through the question-and-answer data, thereby improving the accuracy of the model's answers for a specific domain.
[0237] For the specific implementation of each of the above units, reference may be made to the previous embodiments, which will not be elaborated herein.
[0238] Refer to Figure 4 , Figure 4Structural block diagram of a part of computer device 1000 for implementing the embodiments of the present disclosure. Computer device 1000 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 622 (for example, one or more processors) and a memory 632, and one or more storage media 630 (for example, one or more mass storage devices) for storing application programs 642 or data 644. Among them, memory 632 and storage medium 630 may be transient storage or persistent storage. The programs stored in storage medium 630 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on computer device 1000. Further, central processing unit 622 may be configured to communicate with storage medium 630 and execute a series of instruction operations in storage medium 630 on computer device 1000.
[0239] Computer device 1000 may further include one or more power supplies 626, one or more wired or wireless network interfaces 650, one or more input / output interfaces 658, and / or one or more operating systems 641, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, and so on.
[0240] The central processing unit 622 in computer device 1000 may be used to execute the model training method of the embodiments of the present disclosure, for example:
[0241] Generate a plurality of first security task descriptions, perform clustering on the plurality of first security task descriptions to obtain a plurality of security task description clusters;
[0242] Obtain a plurality of security reference data, and determine the similarity between each security reference data and any first security task description in each security task description cluster;
[0243] Determine the target security task description corresponding to each security reference data according to the similarity between each security reference data and the corresponding first security task description;
[0244] Generate a first input instruction based on each security reference data and the corresponding target security task description, and the first input instruction is an input prompt for commanding to generate question-and-answer data according to the security reference data and the corresponding target security task description;
[0245] Input the first input instruction into the trained generation model to generate the question-and-answer data corresponding to the first input instruction;
[0246] Input the Q&A data into the trained security task model to fine-tune the trained security task model and obtain a fine-tuned security task model.
[0247] The embodiments of the present disclosure further provide a computer-readable storage medium for storing program codes for executing the model training methods of the foregoing various embodiments.
[0248] The embodiments of the present disclosure further provide a computer program product including a computer program. The processor of the computer device reads and executes the computer program, enabling the computer device to execute the model training method described above. For example:
[0249] Generate a plurality of first security task descriptions, perform clustering on the plurality of first security task descriptions to obtain a plurality of security task description clusters;
[0250] Obtain a plurality of security reference data, and determine the similarity between each security reference data and any first security task description in each security task description cluster;
[0251] Determine the target security task description corresponding to each security reference data according to the similarity between each security reference data and the corresponding first security task description;
[0252] Generate a first input instruction based on each security reference data and the corresponding target security task description, where the first input instruction is an input prompt for generating Q&A data according to the security reference data and the corresponding target security task description;
[0253] Input the first input instruction into the trained generation model to generate the Q&A data corresponding to the first input instruction;
[0254] Input the Q&A data into the trained security task model to fine-tune the trained security task model and obtain a fine-tuned security task model.
[0255] In addition, the terms "including" and "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0256] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one)" or a similar expression below refers to any combination of these items, including any combination of a single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0257] It should be understood that in the description of the embodiments of this application, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as "greater than", "less than", and "exceeding" do not include the present number, and understandings such as "above", "below", and "within" include the present number.
[0258] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.
[0259] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0260] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0261] When an integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0262] It should also be understood that the various embodiments provided in the embodiments of this application can be combined arbitrarily to achieve different technical effects.
[0263] In the embodiments of this application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of this module or unit.
[0264] The above is a specific description of the embodiments of this application, but this application is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of this application, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.
Claims
1. A model training method, characterized in that: include: Generate a plurality of first safety task descriptions, and cluster the plurality of first safety task descriptions to obtain a plurality of safety task description clusters; Acquire a plurality of safety reference data, and determine the similarity between each of the safety reference data and any first safety task description in each of the safety task description clusters; Determining a target security task description corresponding to each of the security reference data according to a similarity between each of the security reference data and the corresponding first security task description; Based on each of the security reference data and the corresponding target security task description, a first input instruction is generated, wherein the first input instruction is a command to generate an input prompt for question and answer data according to the security reference data and the corresponding target security task description; Inputting the first input instruction into the trained generation model to generate question and answer data corresponding to the first input instruction, wherein the question and answer data includes a question, a thinking process, and an answer; Inputting the question into the trained security task model to obtain a first prediction probability of each word unit in the thinking process and a second prediction probability of each non-stop word word unit in the answer; Determining a first loss value based on a first predicted probability of each word-gram in the thought process and a second predicted probability of each non-stop word word-gram in the answer; Performing a first-stage fine-tuning on the trained safety task model according to the first loss value to obtain a safety task model after first-stage fine-tuning; Input the question and the thinking process into the security task model with adjustable parameters after fine-tuning in the first stage, so as to obtain the security task model with adjustable parameters after fine-tuning in the first stage to predict the third prediction probability of each word in the answer and the first distribution probability of the answer; Inputting the question and the thinking process into the safety task model with frozen parameters fine-tuned in the first stage to obtain a second distribution probability for predicting the answer; determining a second loss value based on a third predicted probability of each word-gram in the answer, the first distribution probability of the answer, and the second distribution probability of the answer; The second stage fine-tuning is performed on the safety task model with adjustable parameters after the first stage fine-tuning according to the second loss value to obtain the safety task model after fine-tuning.
2. The model training method according to claim 1, characterized in that: The clustering of the plurality of first safety task descriptions to obtain a plurality of safety task description clusters includes: Performing vector conversion on each of the first safety task descriptions to obtain a vector representation of each of the first safety task descriptions; Selecting a first preset number of candidate safety task descriptions from each of the first safety task descriptions as initial clustering centers of different initial safety task description clusters; Calculating the distance between the vector representation of each other security task description and the vector representation of each of the initial cluster centers; Dividing each of the other safety task descriptions into an initial safety task description cluster corresponding to the nearest initial clustering center, to obtain a plurality of updated safety task description clusters; Determine the mean of the vector representations of other safety task descriptions in each of the updated safety task description clusters and the vector representation of the corresponding initial cluster center to obtain the updated cluster center of each of the updated safety task description clusters; When the clustering termination condition is not met, the updated cluster center is determined as the initial cluster center, the updated safety task description cluster is determined as the initial safety task description cluster, and the step of calculating the distance between the vector representation of each other safety task description and the vector representation of each initial cluster center is returned to execute until the clustering termination condition is met, and the obtained multiple updated safety task description clusters are determined as safety task description clusters.
3. The model training method according to claim 2, characterized in that: After obtaining the updated cluster center of each updated safety task description cluster, the method further includes: Calculating a difference value between a vector representation of an updated cluster center of each updated safety task description cluster and a vector representation of a corresponding initial cluster center; Comparing the difference value of each updated security task description cluster with a preset difference value; When there is an updated safety task description cluster whose corresponding difference value is greater than or equal to the preset difference value, determining that the clustering termination condition is not satisfied; When the difference value of each of the updated safety task description clusters is less than the preset difference value, it is determined that the clustering termination condition is met.
4. The model training method according to claim 1, characterized in that: The generating of a plurality of first safety task descriptions comprises: Selecting a second preset number of safety task descriptions from the first safety task description set and the constructed second safety task description set; Generate a second input instruction based on the selected safety task description, the second input instruction is used to command the generation of a third preset number of generated safety task descriptions according to the selected safety task description; Inputting the second input instruction into the trained generation model to generate a corresponding generation safety task description; Determine the generated safety task description that meets the safety task description condition as the first safety task description, and assign it to the first safety task description set; When the total number of first security task descriptions in the first security task description set does not reach a fourth preset number, return to execute the step of selecting a second preset number of security task descriptions from the constructed second security task description set and the first security task description set until the total number of first security task descriptions in the first security task description set reaches a fourth preset number, so as to obtain multiple first security task descriptions in the first security task description set.
5. The model training method according to claim 4, characterized in that: The step of determining the generated safety task description that meets the safety task description condition as the first safety task description includes: Generate a third input instruction based on each of the generated safety task descriptions, the third input instruction being used to command an answer as to whether the corresponding generated safety task description is related to the safety task; Input each of the third input instructions into the trained generation model to generate an answer result; The generated safety task description corresponding to the third input instruction for which the answer result is yes is determined as the first safety task description.
6. The model training method according to claim 1, characterized in that: The determining of the first loss value based on the first predicted probability of each word-gram in the thinking process and the second predicted probability of each non-stop word-gram in the answer comprises: Obtaining the position weight of each word in the thinking process; Determine a first sub-loss value of the thinking process based on the position weight of each word element in the thinking process and the corresponding first prediction probability; determining a sub-loss value for an answer based on a second predicted probability of a word-gram for each non-stop word in the answer; The first sub-loss value and the answer sub-loss value are weightedly summed to obtain a first loss value.
7. The model training method according to claim 1, characterized in that: The determining of the second loss value based on the third predicted probability of each word in the answer, the first distribution probability of the answer, and the second distribution probability of the answer comprises: determining a second sub-loss value based on a third predicted probability of each word-gram in the answer; Determine a KL divergence between a first distribution probability of the answer and a second distribution probability of the answer to obtain a knowledge distillation sub-loss value; A weighted sum is performed on the second sub-loss value and the knowledge distillation sub-loss value to obtain a second loss value.
8. A model training device, characterized in that: include: A clustering division unit, used to generate a plurality of first safety task descriptions, and cluster the plurality of first safety task descriptions to obtain a plurality of safety task description clusters; A first determining unit, configured to obtain a plurality of safety reference data, and determine a similarity between each of the safety reference data and any first safety task description in each of the safety task description clusters; A second determining unit, configured to determine a target security task description corresponding to each security reference data according to a similarity between each security reference data and the corresponding first security task description; A generating unit, configured to generate a first input instruction based on each of the security reference data and the corresponding target security task description, wherein the first input instruction is a command to generate an input prompt for question and answer data according to the security reference data and the corresponding target security task description; An input unit, configured to input the first input instruction into a trained generation model to generate question-and-answer data corresponding to the first input instruction, wherein the question-and-answer data includes a question, a thinking process, and an answer; Fine-tuning unit, including: A second input subunit is used to input the question into the trained security task model to obtain a first prediction probability of each word unit in the thinking process and a second prediction probability of each non-stop word word unit in the answer; a fifth determination subunit, configured to determine a first loss value based on a first prediction probability of each word-gram in the thinking process and a second prediction probability of each non-stop word word-gram in the answer; A first fine-tuning subunit, configured to perform a first-stage fine-tuning on the trained safety task model according to the first loss value to obtain a safety task model after first-stage fine-tuning; A third input subunit is used to input the question and the thinking process into the security task model with adjustable parameters after fine-tuning in the first stage, so as to obtain a third prediction probability of each word in the answer predicted by the security task model with adjustable parameters after fine-tuning in the first stage, and to predict a first distribution probability of the answer; A fourth input subunit is used to input the question and the thinking process into the safety task model of the frozen parameters after fine-tuning in the first stage, so as to obtain a second distribution probability for predicting the answer; a sixth determining subunit, configured to determine a second loss value based on a third predicted probability of each word in the answer, the first distribution probability of the answer, and the second distribution probability of the answer; The second fine-tuning subunit is used to perform a second-stage fine-tuning on the safety task model with adjustable parameters after the first-stage fine-tuning according to the second loss value to obtain a safety task model after the fine-tuning.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, which are suitable for loading by a processor to execute the model training method described in any one of claims 1 to 7.
10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the model training method described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Knowledge question and answer model training method, knowledge question and answer method and related device
CN118520294A
Question and answer task processing model training method and device, equipment and storage medium
CN119493849A