Model training method and device, storage medium and computer equipment

By clustering the security task description and combining security reference data to generate Q&A data, the security task model is fine-tuned, which solves the problem of poor answer accuracy caused by the lack of high-quality data in the field in the prior art, and improves the accuracy of the model's answers in a specific security field.

CN119988985AActive Publication Date: 2025-05-13PENG CHENG LAB

Patent Information

Application Number
CN202510459014.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

The lack of high-quality data in the field in the prior art has led to poor accuracy in answers to large models in specific security fields.

Method used

By generating multiple security task descriptions, clustering and partitioning is performed to obtain security task description clusters in specific fields, and input instructions are generated based on security reference data, and the generative model is used to generate question-and-answer data to fine-tune the security task model.

Benefits of technology

Improve the accuracy of the model's answers in specific security areas, and solves the problem of poor answer accuracy caused by the lack of high-quality data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988985A_ABST
    Figure CN119988985A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model training method and device, a storage medium and computer equipment, and the method comprises the steps: generating a plurality of first security task descriptions, and carrying out the clustering division to obtain a security task description cluster; and obtaining security reference data, determining the similarity between the security reference data and the description in the cluster, and determining a target security task description corresponding to each reference data according to the similarity. A first input instruction is generated based thereon, the instruction requiring generation of question and answer data in accordance with the security reference data and the target security task description. And inputting the instruction into the trained generation model to obtain corresponding question and answer data. And inputting the question and answer data into the trained security task model for fine tuning so as to obtain the fine-tuned security task model. According to the process, security task description is firstly clustered, then a reference data generation instruction is combined, question and answer data is obtained by utilizing the generation model, finally fine adjustment of the security task model is realized, and the accuracy of answering in a specific field by the model is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a model training method, device, storage medium and computer equipment. Background Art

[0002] As a key department in network security, the security operation center is responsible for using various tools, technologies and processes to detect, analyze and respond to network security incidents. However, the center currently faces many severe challenges, such as lack of security expertise, a lot of time spent on investigating alerts, and slow response to advanced threats. With the emergence of big models, it has become possible to use their powerful understanding and generation capabilities to assist security operations. However, it should be noted that network security is a highly professional field, and dealing with security issues requires deep professional knowledge and skills.

[0003] In related technologies, in order to build a large question-answering model that fits the security operation scenario, it is particularly necessary to fine-tune the large model with domain-specific data. However, the outstanding problem currently faced is the lack of high-quality data in the field, and most of the existing large model fine-tuning technologies are concentrated in general fields, without fully considering the construction of fine-tuning data in specific fields, resulting in poor accuracy of the model's answers in specific fields. Therefore, related technologies urgently need to propose a model training method to solve the above technical problems. Summary of the invention

[0004] The main purpose of this application is to provide a model training method, device, storage medium and computer equipment, which can fine-tune the model through question and answer data in a specific field to improve the accuracy of the model's answers in a specific field.

[0005] In a first aspect, an embodiment of the present application provides a model training method, comprising: Generate a plurality of first safety task descriptions, and cluster the plurality of first safety task descriptions to obtain a plurality of safety task description clusters; Acquire a plurality of safety reference data, and determine the similarity between each of the safety reference data and any first safety task description in each of the safety task description clusters; Determining a target security task description corresponding to each security reference data according to a similarity between each security reference data and a corresponding first security task description; Based on each of the security reference data and the corresponding target security task description, a first input instruction is generated, wherein the first input instruction is a command to generate an input prompt for question and answer data according to the security reference data and the corresponding target security task description; Inputting the first input instruction into the trained generation model to generate question and answer data corresponding to the first input instruction; The question-and-answer data is input into the trained safety task model to fine-tune the trained safety task model to obtain a fine-tuned safety task model.

[0006] In a second aspect, an embodiment of the present application provides a model training device, comprising: A clustering division unit, used to generate a plurality of first safety task descriptions, and cluster the plurality of first safety task descriptions to obtain a plurality of safety task description clusters; A first determining unit, configured to obtain a plurality of safety reference data, and determine a similarity between each of the safety reference data and any first safety task description in each of the safety task description clusters; A second determining unit, configured to determine a target security task description corresponding to each security reference data according to a similarity between each security reference data and the corresponding first security task description; A generating unit, configured to generate a first input instruction based on each of the security reference data and the corresponding target security task description, wherein the first input instruction is a command to generate an input prompt for question and answer data according to the security reference data and the corresponding target security task description; An input unit, used to input the first input instruction into the trained generation model to generate question and answer data corresponding to the first input instruction; A fine-tuning unit is used to input the question and answer data into the trained safety task model to fine-tune the trained safety task model to obtain a fine-tuned safety task model.

[0007] In a third aspect, an embodiment of the present application provides a storage medium, wherein the computer-readable storage medium stores a plurality of instructions, which are suitable for loading by a processor to execute any of the above model training methods.

[0008] In a fourth aspect, an embodiment of the present application provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any one of the above model training methods when executing the computer program.

[0009] In an embodiment of the present application, a plurality of first security task descriptions are generated, and the plurality of first security task descriptions are clustered to obtain a plurality of security task description clusters; a plurality of security reference data are acquired, and the similarity between each of the security reference data and any first security task description in each of the security task description clusters is determined; a target security task description corresponding to each of the security reference data is determined based on the similarity between each of the security reference data and the corresponding first security task description; a first input instruction is generated based on each of the security reference data and the corresponding target security task description, and the first input instruction is a command to generate an input prompt for question and answer data based on the security reference data and the corresponding target security task description; the first input instruction is input into a training data unit; The trained generation model generates question and answer data corresponding to the first input instruction; the question and answer data is input into the trained security task model to fine-tune the trained security task model to obtain a fine-tuned security task model. Compared with the problem of poor accuracy of answers in specific fields due to lack of high-quality data in the field in related technologies, the embodiment of the present application describes multiple first security tasks and clusters them to obtain a cluster of security task descriptions in a specific field, and generates a first input instruction in combination with reference data, obtains question and answer data for model training in a specific field through the generation model, and trains the security task model that needs to be fine-tuned after training through the question and answer data, thereby improving the accuracy of the model's answers in specific fields.

[0010] Other features and advantages of the present disclosure will be described in the following description, and partly become apparent from the description, or be understood by practicing the present disclosure. The purpose and other advantages of the present disclosure can be realized and obtained by the structures particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0012] Figure 1 A schematic diagram of a scenario of a model training system provided in an embodiment of the present application.

[0013] Figure 2 A flowchart of the model training method provided in an embodiment of the present application.

[0014] Figure 3 A schematic diagram of the structure of a model training device provided in an embodiment of the present application.

[0015] Figure 4 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0016] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0017] It should be noted that in some processes described in the specification, claims and the above-mentioned drawings, multiple steps appearing in a specific order are included, but it should be clearly understood that these steps may not be executed in the order in which they appear in this document or may be executed in parallel. The step numbers are only used to distinguish different steps, and the numbers themselves do not represent any execution order. In addition, descriptions such as "first", "second" or "target" in this document are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0018] See also Figure 1 , Figure 1 A schematic diagram of a scenario of a model training system provided in an embodiment of the present application, which includes a terminal 140, the Internet 130, a gateway 120, a server 110, etc.

[0019] The terminal 140 includes but is not limited to a pre-configured laptop, tablet computer, desktop computer or other electronic device with data reporting capability. In addition, it can be a single device or a collection of multiple devices. The terminal 140 can communicate with the Internet 130 in a wired or wireless manner to exchange data.

[0020] The terminal 140 refers to a computer system that can report data to the server 110. Compared with ordinary terminals, the server 110 has higher requirements in terms of stability, security, performance, etc. The server 110 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a part of a high-performance computer (such as a virtual machine), a combination of parts of multiple high-performance computers (such as virtual machines), etc.

[0021] The gateway 120 is also called an internetwork connector or a protocol converter. The gateway realizes network interconnection at the transport layer and is a computer system or device that acts as a converter. The gateway is a translator between two systems that use different communication protocols, data formats or languages, or even completely different architectures. At the same time, the gateway can also provide filtering and security functions. The message sent by the terminal 140 to the server 110 must be sent to the corresponding server 110 through the gateway 120. The message sent by the server 110 to the terminal 140 must also be sent to the corresponding terminal 140 through the gateway 120.

[0022] The model training method of the embodiment of the present disclosure can be implemented on the server 110 .

[0023] It should be noted that Figure 1 The scenario diagram of the model training system shown is only an example. The model training system and scenario described in the embodiment of the present application are intended to more clearly illustrate the technical solution of the embodiment of the present application, and do not constitute a limitation on the technical solution provided in the embodiment of the present application. A person of ordinary skill in the art can know that with the evolution of image processing technology and the emergence of new business scenarios, the technical solution provided in the embodiment of the present application is also applicable to similar technical problems.

[0024] In this embodiment, the model training device will be described from the perspective of the model training device, which can be specifically integrated into a computer device having a storage unit and a microprocessor installed therein and having computing capabilities.

[0025] See also Figure 2 , Figure 2 A flow chart of a model training method provided in an embodiment of the present application. The model training method comprises: In step 201, a plurality of first safety task descriptions are generated, and the plurality of first safety task descriptions are clustered to obtain a plurality of safety task description clusters.

[0026] The first security task description refers to a textual description of a task in a specific security field, such as "detecting whether there is malware intrusion in the network", "evaluating the system's defense capabilities when it is attacked by DDoS", etc. It details the specific tasks that need to be completed in the security field. The security task description cluster is a set obtained by clustering multiple first security task descriptions. The first security task descriptions in the same cluster belong to the same security field. For example, the task descriptions in the security field such as network attack detection are divided into one cluster.

[0027] In some implementations, clustering the plurality of first security task descriptions to obtain a plurality of security task description clusters includes: (1) performing vector conversion on each of the first safety task descriptions to obtain a vector representation of each of the first safety task descriptions; (2) selecting a first preset number of candidate safety task descriptions from each of the first safety task descriptions as initial clustering centers of different initial safety task description clusters; (3) calculating the distance between the vector representation of each of the other security task descriptions and the vector representation of each of the initial cluster centers; (4) dividing each of the other safety task descriptions into an initial safety task description cluster corresponding to the initial cluster center that is closest to the other safety task description, to obtain a plurality of updated safety task description clusters; (5) determining the mean of the vector representations of other safety task descriptions in each of the updated safety task description clusters and the vector representations of the corresponding initial cluster centers, to obtain the updated cluster centers of each of the updated safety task description clusters; (6) When the clustering termination condition is not met, the updated cluster center is determined as the initial cluster center, the updated safety task description cluster is determined as the initial safety task description cluster, and the step of calculating the distance between the vector representation of other safety task descriptions and the vector representation of each of the initial cluster centers is returned to execute until the clustering termination condition is met, and the multiple updated safety task description clusters obtained are determined as safety task description clusters.

[0028] The vector conversion converts the first security task description in text form into a numerical vector form that can be processed by the computer for subsequent distance calculation and cluster analysis. Each task description is converted into a vector representation using an embedding model (such as Word2Vec, BERT, etc.). Assume that there are n task descriptions, denoted as , after being processed by the embedding model, each task description is converted into a Dimensional vector ,in .

[0029] Specifically, the first preset number is a pre-set value used to determine the number of selected initial clustering centers. The setting of this number will affect the clustering results and computational efficiency, and usually needs to be reasonably adjusted according to the number and characteristics of the first safety task descriptions. The candidate safety task description is a safety task description selected from the first safety task description set as the initial clustering center. The initial safety task description cluster is at the beginning of clustering, with each initial clustering center as the core, and other safety task descriptions will be divided into these clusters later to form a preliminary clustering result. The initial clustering center is the core of each initial safety task description cluster, which is a vector representing the initial characteristics of the cluster. Other safety task descriptions are other first safety task descriptions in the first safety task description except those first safety task descriptions selected as the initial clustering center. Distance is used to measure the similarity or difference between two vectors. Common distance measurement methods include Euclidean distance (Euclidean distance) and the like.

[0030] Specifically, initialize the cluster and randomly select data points as the initial cluster centers, denoted as in, , .

[0031] For each additional security task description , calculate its relationship with each cluster center The distance is usually the Euclidean distance: ; in, , Other security tasks are described With each cluster center distance. Description for other security tasks With each cluster center The difference of vectors of the same dimension in .

[0032] According to the distance calculated in the previous step, for each other safety task description, find the initial cluster center closest to it, and then divide the description into the initial safety task description cluster corresponding to this initial cluster center. After the division operation of all other safety task descriptions, multiple updated safety task description clusters are obtained. For each updated safety task description cluster, the vector representations of all other safety task descriptions in it are aggregated with the vector representation of the corresponding initial cluster center (here the average is calculated). In this way, the updated cluster center of each updated safety task description cluster is calculated.

[0033] The clustering termination condition is a pre-set condition for determining whether the clustering process can be terminated. Check whether the current clustering result meets the clustering termination condition. If not, the updated cluster center is used as the new initial cluster center, and the updated safety task description cluster is used as the new initial safety task description cluster. Then the process is executed in a loop, that is, the distance between other safety task descriptions and the new initial cluster center is recalculated, and the descriptions are divided and the cluster center is updated. This process is repeated until the clustering termination condition is met. Finally, the multiple updated safety task description clusters obtained when the termination condition is met are determined as the final safety task description cluster.

[0034] Specifically, for each cluster, the mean of all data points in the cluster is calculated to update the cluster center. The safety task description cluster is , then the updated cluster center for: ; in, Represents a safety task description cluster The number of data points in the cluster, that is, the updated cluster center is determined by averaging Repeat the above steps to divide the first safety task description into clusters, denoted as ,Each cluster contains a set of first security task descriptions in the same security domain.

[0035] In this way, through clustering, the numerous first security task descriptions are divided into different security task description clusters according to similarity. The originally scattered and disordered security task descriptions are integrated to form organized clusters with specific characteristics. In the subsequent security task model training, the clustered security task description clusters can be used as high-quality training data. The task descriptions within the same cluster are similar, and the model can better learn the patterns and characteristics of this type of task, thereby improving the training effect and generalization ability of the model. For example, when training a model for detecting network attacks, using the security task description cluster related to network attack detection as training data, the model can learn the knowledge in this field more attentively and improve the accuracy and efficiency of detection.

[0036] In some implementations, after obtaining the updated cluster center of each updated safety task description cluster, the method further includes: (1) calculating the difference between the vector representation of the updated cluster center of each updated safety task description cluster and the vector representation of the corresponding initial cluster center; (2) comparing the difference value of each updated security task description cluster with a preset difference value; (3) when there is an updated safety task description cluster whose corresponding difference value is greater than or equal to the preset difference value, determining that the clustering termination condition is not satisfied; (4) When the difference value of each of the updated safety task description clusters is less than the preset difference value, it is determined that the clustering termination condition is met.

[0037] Among them, for each updated safety task description cluster, use a suitable difference measurement method (such as Euclidean distance, cosine distance, etc.) to calculate the difference between the vector representation of the updated cluster center and the vector representation of the corresponding initial cluster center. The preset difference value is a pre-set threshold used as a standard for judging the degree of difference between the updated cluster center and the initial cluster center. The setting of this value usually needs to be adjusted according to the actual clustering requirements and data characteristics. Compare the calculated difference value of each updated safety task description cluster with the preset difference value.

[0038] Specifically, the comparison results of the difference values ​​of all updated safety task description clusters and the preset difference values ​​are checked. If there is at least one updated safety task description cluster whose difference value is greater than or equal to the preset difference value, it is considered that the current clustering result is not stable enough and the change of the cluster center is still relatively large. At this time, it is determined that the clustering division termination condition is not met and the clustering operation needs to be continued. When the difference values ​​of all updated safety task description clusters are less than the preset difference value, it means that after this round of clustering operation, the difference between the updated cluster center and the initial cluster center of each cluster is relatively small, and the cluster center has been relatively stable. At this time, it can be considered that the clustering result has achieved the expected effect, and it is determined that the clustering division termination condition is met. The clustering process can be stopped, and the multiple updated safety task description clusters currently obtained are determined as the final safety task description clusters.

[0039] In this way, by comparing the difference between the updated cluster center and the initial cluster center and comparing it with the preset difference value, it is possible to effectively avoid stopping clustering prematurely when the cluster center is not yet stable. If such a judgment is not made, the clustering results may be inaccurate because the cluster center has not fully converged to a position that can accurately represent the characteristics of the data within the cluster. For example, when the difference value is large, it means that the cluster center is still changing to a large extent. If clustering is stopped at this time, some similar safety task descriptions may be divided into different clusters, or different safety task descriptions may be mistakenly divided into the same cluster. Through this judgment mechanism, the cluster center is considered stable only when the difference values ​​of all clusters are small, thereby obtaining a more accurate and reliable clustering result.

[0040] In some implementations, generating a plurality of first safety task descriptions includes: (1) selecting a second preset number of safety task descriptions from the first safety task description set and the constructed second safety task description set; (2) generating a second input instruction based on the selected safety task description, wherein the second input instruction is used to command the generation of a third preset number of generated safety task descriptions according to the selected safety task description; (3) inputting the second input instruction into the trained generation model to generate a corresponding generation safety task description; (4) determining the generated security task description that meets the security task description condition as a first security task description, and assigning it to the first security task description set; (5) When the total number of first security task descriptions in the first security task description set does not reach a fourth preset number, return to execute the step of selecting a second preset number of security task descriptions from the constructed second security task description set and the first security task description set until the total number of first security task descriptions in the first security task description set reaches a fourth preset number, thereby obtaining multiple first security task descriptions in the first security task description set.

[0041] The first security task description set is used to store the generated first security task, the initial set is an empty set, and the second security task description set is a task set constructed by the technician according to the needs of the security field, and the set includes multiple second security task descriptions under at least one security field. The second preset number is a pre-set value used to specify the total number of security task descriptions selected from the first security task description set and the second security task description set. Security task descriptions are randomly selected from the first security task description set and the second security task description set, so that the total number of selected descriptions reaches the second preset number, and the number of security task descriptions selected from the second security task description set is greater than the number of security task descriptions selected from the first security task description set.

[0042] Specifically, the second input instruction is an instruction generated based on the selected security task description, and its function is to convey specific task requirements to the trained generative model, that is, to let the model generate new security task descriptions using the selected descriptions as examples. For example, the selected security task descriptions are: Task 1, Task 2, and Task 3. Task 1 is: Threat Entity Extraction, and the task description is: Please extract entities related to network attack incidents such as threat organizations, malware, and attack methods from the given text. The extracted information is required to be comprehensive and accurate, and the category of each entity must be marked. Task 2 and Task 3 are similar to Task 1 and will not be repeated here. The third preset number is not less than 3, so the second input instruction generated based on Task 1, Task 2, and Task 3 is: “Please generate no less than 3 tasks and task descriptions related to security tasks: Task 1: Threat Entity Extraction Task description: Please extract entities related to cyber attack incidents such as threat organizations, malware, and attack methods from the given text. The extracted information must be comprehensive and accurate, and the category of each entity must be marked.

[0043] Mission 2: XXX Mission 3: XXXX" In this way, the generated second input instruction is input into a trained generation model such as GPT-4o, ChatGPT, etc., so as to generate a generation security task description corresponding to the second input instruction through the trained generation model.

[0044] Since the specific content of the generated security task description may be irrelevant to the security task, only the generated security task description of the security task description condition is determined as the first security task description and allocated to the first security task description set. The total number of first security task descriptions in the first security task description set is detected. When the total number of first security task descriptions in the first security task description set is not reached, the step of selecting a second preset number of security task descriptions from the constructed second security task description set and the first security task description set is returned to execute until the total number of first security task descriptions in the first security task description set reaches the fourth preset number, thereby obtaining multiple first security task descriptions in the first security task description set.

[0045] Therefore, in the field of security, high-quality task description data is often limited. The embodiment of the present application generates new descriptions based on existing security task descriptions with the help of a trained generation model, which can effectively expand the size of the first security task description set. As the number of iterations increases, the number of task descriptions in the set continues to increase, providing richer data support for subsequent security task analysis, model training and other work. The generation model can expand, modify and combine the selected security task descriptions to varying degrees to generate new descriptions with diversity. These new descriptions may cover different security scenarios, task types and expressions, making the data in the first security task description set more diverse.

[0046] Moreover, as security services develop, the scope of security tasks may continue to expand or adjust. This solution allows technicians to build and update the second security task description set according to actual needs, and through the iterative generation process, the first security task description set can flexibly adapt to changes in the scope of tasks. For example, when an enterprise decides to expand its security business to a new field, it can add task descriptions of the new field to the second security task description set, and then incorporate the relevant task descriptions into the first security task description set.

[0047] In addition, in the embodiment of the present application, a security task description condition is set to screen the generated descriptions, and only the descriptions that meet the condition are determined as the first security task descriptions and added to the set. This screening mechanism can effectively exclude descriptions that are irrelevant to the security task or of low quality, and ensure the quality of the data in the first security task description set.

[0048] In some implementations, determining the generated security task description that meets the security task description condition as the first security task description includes: (1) generating a third input instruction based on each of the generated safety task descriptions, the third input instruction being used to command an answer as to whether the corresponding generated safety task description is related to the safety task; (2) inputting each of the third input instructions into the trained generation model to generate an answer result; (3) The generated safety task description corresponding to the third input instruction for which the answer result is yes is determined as the first safety task description.

[0049] Among them, the third input instruction is a specific instruction generated according to each generated safety task description, and its main function is to let the trained generation model determine whether the description is related to the safety task. This instruction clarifies that the task is to make a relevance judgment on the generated safety task description. For each generated safety task description, it is integrated into an instruction to form a third input instruction. Each generated third input instruction is input into the trained generation model in turn. After receiving the instruction, the model understands and analyzes the instruction, and judges the relevance of the generated safety task description to the safety task based on the knowledge and patterns it has learned, and generates a corresponding answer result.

[0050] Specifically, the answer result corresponding to each third input instruction is checked. If the answer result is "yes", the generated safety task description corresponding to the third input instruction is determined as the first safety task description. Then, these determined first safety task descriptions are allocated to the first safety task description set.

[0051] For example, the generated security task description is: Malicious IP and domain name identification, and the task description is to identify malicious IP addresses and domain names from a given network traffic or DNS query log. According to the existing threat intelligence database, determine whether these IPs and domain names are related to known attackers or malicious or related. Then the third input instruction is: "Please determine whether the following is relevant to the security task: Malicious IP and domain name identification Task description: Identify malicious IP addresses and domain names from given network traffic or DNS query logs. Based on the existing threat intelligence database, determine whether these IP addresses and domain names are related to known attackers or malicious actors.

[0052] Please output "yes" if relevant and "no" if not relevant.

[0053] Output: XXX" In this way, in the process of generating safety task descriptions, some content that is irrelevant or weakly relevant to the safety task may be generated. By allowing the trained generation model to perform relevance judgment on each generated safety task description, the descriptions that are truly relevant to the safety task can be effectively screened out. The use of the trained generation model for relevance judgment realizes the automation of the judgment process. Compared with manually judging the relevance of generated safety task descriptions one by one, this automated method greatly improves work efficiency and saves labor costs. The generation model can quickly process a large number of third input instructions and generate corresponding answer results. It is especially suitable for situations where a large number of generated safety task descriptions are generated. In this example, "XXX" represents the answer result output by the generation model, specifically "yes" or "no".

[0054] In step 202, a plurality of safety reference data are obtained, and the similarity between each safety reference data and any first safety task description in each safety task description cluster is determined.

[0055] Among them, security reference data refers to data related to the security field, including network traffic data, system log data, security vulnerability report data, etc. These data are the basis for subsequent analysis and generation of question and answer data.

[0056] Specifically, security reference data is collected from multiple channels such as ATT&CK, CVE, CWE, and open source threat intelligence, and the security reference data and the first security task description are converted into vector form. Then, a similarity calculation method (such as cosine similarity, etc.) is used to calculate the similarity between each security reference data and any first security task description in each security task description cluster. In this way, the similarity value between each security reference data and each first security task description can be obtained.

[0057] Specifically, a data set of security reference data is obtained through steps such as cleaning, long text segmentation, and deduplication. ,against Every safety reference data , from T Each cluster samples a threat intelligence task description and Perform similarity calculation.

[0058] In step 203, the target security task description corresponding to each security reference data is determined according to the similarity between each security reference data and the corresponding first security task description.

[0059] Among them, for each security reference data, from The s first safety task descriptions with the highest similarity are obtained from the first safety task descriptions (s is a parameter customized by the fine-tuning data builder), and these s first safety task descriptions are determined as the corresponding target safety task descriptions.

[0060] In step 204, based on each security reference data and the corresponding target security task description, a first input instruction is generated, and the first input instruction is a command to generate an input prompt for question and answer data according to the security reference data and the corresponding target security task description.

[0061] Among them, the first input instruction is a clear instruction, which is used to tell the subsequent generation model that it is necessary to generate question and answer data based on specific security reference data and the corresponding target security task description.

[0062] For example, the security reference data is the specific data content of the network traffic data. Based on the specific data content of the network traffic data and the target security task description of detecting whether there is malware intrusion behavior in the network, question and answer data on how to detect malware intrusion based on the traffic data is generated.

[0063] Specifically, the specific content of each security reference data and its corresponding target security task description are integrated to generate the first input instruction in a certain format. The format of the instruction can be designed according to actual needs, but usually it clearly indicates the input data and task description, as well as the requirement to generate question and answer data.

[0064] Since the target security task description is a security task description in a specific field, the question and answer data generated according to the first input instruction is also question and answer data related to the specific field.

[0065] The first input instruction may be as shown in the following example: “Mission: Malicious IP and domain name identification Text: The security team discovered the following suspicious activity while analyzing network logs: The IP address 173.XXX has initiated more than 10,000 connection requests to the internal server in the past 24 hours, all of which are HTTPS requests. The IP address also attempted to access multiple common management backend paths such as / admin.

[0066] At the same time, it was found that the domain name hack.XXX was being resolved to the IP address, and the domain name was newly registered 48 hours ago.

[0067] A query on the security platform showed that the IP address had been marked as a malicious scanning source by multiple security vendors.

[0068] Please generate a task-related question and answer based on the text content in the following format: { question: Thought Process: answer: }".

[0069] In step 205, the first input instruction is input into the trained generation model to generate question and answer data corresponding to the first input instruction.

[0070] The generated first input instruction is input into the trained generation model, and the model understands and processes the first input instruction based on the knowledge and patterns it has learned, and generates corresponding question and answer data. During the generation process, the model will consider the input security reference data and the target security task description, and generate question and answer data that meets the requirements as much as possible.

[0071] In step 206 , the question-and-answer data is input into the trained safety task model to fine-tune the trained safety task model to obtain a fine-tuned safety task model.

[0072] Among them, the trained security task model is a model that has been trained for security tasks to a certain extent, but it still needs to be fine-tuned for specific security fields to improve performance in specific security fields. Fine-tuning technologies are not limited to LoRA, QLoRA, P-Tuning v2, Prefix Tuning, Prompt Tuning, etc., and open source models are not limited to LLaMA, Falcon, BLOOM, ChatGLM, Baichuan, InternLM, etc. The specific selection is based on the actual hardware resources and fine-tuning effect comparison.

[0073] Specifically, the trained security task model is fine-tuned through relevant question and answer data in a specific field, and finally a fine-tuned security task model with better performance in a specific security field is obtained.

[0074] In some embodiments, the question-and-answer data includes questions, thinking processes, and answers, and the question-and-answer data is input into the trained safety task model to fine-tune the trained safety task model to obtain the fine-tuned safety task model, including: (1) inputting the question into the trained security model to obtain a first prediction probability of each word in the thinking process and a second prediction probability of each non-stop word word in the answer; (2) determining a first loss value based on a first predicted probability of each word-gram in the thinking process and a second predicted probability of each non-stop word-gram in the answer; (3) performing a first-stage fine-tuning on the trained safety task model according to the first loss value to obtain a safety task model after the first-stage fine-tuning; (4) inputting the question and the thinking process into the security task model with adjustable parameters after fine-tuning in the first stage, so that the security task model with adjustable parameters after fine-tuning in the first stage predicts a third predicted probability of each word in the answer and predicts a first distribution probability of the answer; (5) inputting the question and the thinking process into the safety task model with frozen parameters fine-tuned in the first stage to obtain a second distribution probability for predicting the answer; (6) determining a second loss value based on a third predicted probability of each word in the answer, the first distribution probability of the answer, and the second distribution probability of the answer; (7) Performing a second-stage fine-tuning on the safety task model with adjustable parameters after the first-stage fine-tuning according to the second loss value to obtain a safety task model after fine-tuning.

[0075] The question-and-answer data specifically includes: questions, thinking processes, and answers. Questions are specific questions raised in the question-and-answer data; thinking processes are descriptions of the steps of analyzing, reasoning, and thinking about the questions; and answers are the answers given to the questions. A token is the basic unit in text processing, usually a word or a meaningful character fragment.

[0076] The technical defects of current large model fine-tuning include: (1) when jointly training the thinking process and the answer, the model does not pay enough attention to the key answer information; (2) directly removing the thinking process for fine-tuning will lead to knowledge forgetting; (3) the traditional single loss function cannot balance the quality control of long text generation. Therefore, the embodiment of the present application proposes to use a staged loss function framework when fine-tuning the model, divide the training process into a reasoning enhancement stage (first stage fine-tuning) and a refinement and optimization stage (second stage fine-tuning), and improve the model's reasoning ability and output simplicity through a multi-stage differentiated training strategy. For the reasoning enhancement stage, it is mainly to improve the reasoning ability of the security task model; for the refinement and optimization stage, it is mainly to further improve the accuracy of the security task model for answers on the premise that the first stage fine-tuning improves the reasoning ability of the security task model.

[0077] Specifically, for the first stage of fine-tuning, the input and output are defined as: "Input sequence: X = [CLS] Question: {question} [SEP] Target output: Y = [STA] Thinking process: {reasoning} [SEP] Answer: {answer} ".

[0078] Among them, [CLS] is a classification tag, which is used to aggregate the feature information of the entire input sequence, [STA] is a sequence start tag, [SEP] is a paragraph separator, and is a sequence terminator. Since the output of the first stage fine-tuning includes two parts, the thinking process and the answer, the loss function of the first stage fine-tuning is composed of the sub-loss values ​​of the two parts, the thinking process and the answer. For the sub-loss value of the thinking process, the first prediction probability of each word in the thinking process is predicted by the trained security model. For the sub-loss value of the answer, the second prediction probability of each non-stop word word in the thinking process is predicted by the trained security model. Finally, according to the first prediction probability of each word in the thinking process and the second prediction probability of each non-stop word word in the answer, the first loss value of the first stage fine-tuning is determined, so that the trained security task model is fine-tuned in the first stage according to the first loss value to obtain the security task model after the first stage fine-tuning.

[0079] For the second stage fine-tuning, the input and output are defined as: “Input sequence: X' = [CLS] Question: {question} [SEP] Thinking process: {reasoning} [SEP] Target output: Y' = [STA] Answer: {answer} ".

[0080] The output of the second stage fine-tuning only includes the answer part, but because the second stage fine-tuning needs to prevent the model from forgetting the previously learned knowledge in the process of optimizing the answer, a knowledge distillation sub-loss value is also required to balance the answer generation and knowledge retention. Therefore, the loss function of the second stage fine-tuning is composed of the sub-loss values ​​of the answer and knowledge distillation. For the sub-loss value of the answer, the question and the thinking process are input into the adjustable parameter security task model after the first stage fine-tuning, and the third predicted probability of each word in the answer is obtained by the adjustable parameter security task model after the first stage fine-tuning to determine. For the knowledge distillation sub-loss value, the question and the thinking process are input into the adjustable parameter security task model after the first stage fine-tuning and the frozen parameter security task model after the first stage fine-tuning, respectively, to obtain the first distribution probability of the predicted answer and the second distribution probability of the predicted answer, which are determined by the first distribution probability and the second distribution probability. Finally, the second loss value is determined based on the third predicted probability of each word in the answer, the first distribution probability of the answer, and the second distribution probability of the answer, and the second loss value for the second stage of fine-tuning is determined, so that the trained security task model is fine-tuned in the second stage according to the second loss value to obtain the fine-tuned security task model.

[0081] Therefore, in the reasoning enhancement stage (first stage fine-tuning), the input sequence is set to contain only questions, and the target output contains the thinking process and answers. This setting enables the model to focus more on reasoning from the questions and generate a reasonable thinking process. This effectively improves the reasoning ability of the model and avoids the problem of over-focusing on key answer information and ignoring reasoning when jointly training the thinking process and answers. And because the loss function of the first stage fine-tuning is composed of sub-loss values ​​of the thinking process and the answer, the sub-loss value is determined according to the first predicted probability of each word in the thinking process and the second predicted probability of each non-stop word word in the answer, which can guide the model learning more comprehensively. This prompts the model to pay attention not only to the accuracy of the answer, but also to the rationality of the thinking process, further strengthening the training of reasoning ability.

[0082] In addition, in the refinement and optimization phase (second-phase fine-tuning), the safety task model with frozen parameters after fine-tuning in the first phase is introduced. The questions and thinking process are input into the model and the safety task model with adjustable parameters, and the answer generation and knowledge retention are balanced by calculating the knowledge distillation sub-loss value. This ensures that the model does not forget the knowledge learned in the first phase during the process of optimizing the answer. It ensures that the generated answer not only meets the current task requirements but also retains the previously learned knowledge, effectively avoiding the phenomenon of knowledge forgetting caused by directly removing the thinking process for fine-tuning.

[0083] Through the two-stage fine-tuning, the first stage laid the foundation for the reasoning ability and knowledge of the second stage, and the second stage further optimized the answer while retaining the knowledge. This multi-stage training method enables the model to gradually learn and consolidate knowledge, improving the model's ability to remember and apply knowledge.

[0084] In some implementations, determining a first loss value based on a first predicted probability of each word-gram in the thought process and a second predicted probability of each non-stop word word-gram in the answer includes: (1.1) Obtaining the position weight of each word in the thinking process; (1.2) determining a first sub-loss value of the thinking process based on the position weight of each word unit in the thinking process and the corresponding first prediction probability; (1.3) determining a sub-loss value for the answer based on a second predicted probability of each non-stop word in the answer; (1.4) Perform a weighted sum of the first sub-loss value and the answer sub-loss value to obtain a first loss value.

[0085] Among them, the position weight is used to represent the numerical value of the importance of each word in its position during the thinking process. Words in different positions have different influences on the overall thinking logic and results, and the position weight is used to measure the degree of this influence. For example, the word at the beginning of the thinking process may play an important role in guiding the direction, and its position weight may be relatively high; while the words in the middle or end will also have corresponding weight settings according to their role in the reasoning process.

[0086] Specifically, the position weight can be determined by referring to the following formula: ; in, is the length of the token in the thinking process, and t is the position index, which is used to traverse each token in the token sequence composed of multiple tokens in the thinking process. Calculating the position of the current token is a hyperparameter used to calculate the position weight, which can be set to 0.2. Through this formula, the position weight of each token in the thinking process can be calculated.

[0087] Specifically, the first sub-loss value of the thinking process The loss function can be calculated as follows:

[0088] in, Indicates the real word (or target word) associated with the predicted probability of the word at position t in the thinking process. For example, when generating a sequence of thinking processes, the model has a predicted word distribution for each position t. This is the word that should actually appear at this position, which is used to calculate the difference between the model prediction and the actual situation. represents the sequence of all the words before position t in the thinking process. In other words, it contains all the word information from the beginning of the sequence to position t−1. The model predicts position t When generating a word unit, it will refer to the word unit information that has been generated before, that is, , to calculate the prediction probability of the current position word ,in is the input sequence, i.e., the question. In this way, the model is able to leverage contextual information to make more accurate predictions.

[0089] Specifically, the answer sub-loss value of the answer The loss function can be calculated as follows: ; in, The total number of all words that make up the thinking process and the answer. is a word that is not a stop word, is the indicator function, when hour, =1; when hour, =0. is the true value of the word at position t in the answer part. , it represents the correct word that the model should predict at position t in the answer part, and is used to measure the accuracy of the model's prediction of each word in the answer. It represents the sequence of all the words before position t when generating the answer. Similar, but this is for the answer part. When the model predicts the word at position t in the answer, it will use the previously generated answer word information And the input sequence To calculate The predicted probability of , exploiting contextual information to improve the accuracy of answer generation.

[0090] After getting the first sub-loss value of the thinking process And the answer sub-loss value of the answer After that, the first loss value of the first stage fine-tuning is calculated by the following formula :

[0091] in, is the first sub-loss value The weight of The answer sub-loss value The weight of the first loss value is calculated by weighted summation .

[0092] In this way, by assigning position weights to the words in the thinking process, the importance of words at different positions in the thinking process can be highlighted, avoiding the model treating all words equally when processing the thinking process. Instead, it conducts focused learning based on position and importance, improving the quality of the thinking process. When calculating the answer sub-loss value, a stop word shielding mechanism is introduced (through the indicator function ), only the predicted probability of non-stop words (entity words) is calculated. This can prevent the model from paying too much attention to stop words during training, and instead focus on entity words with actual semantics, thereby increasing the attention to entity words and helping to generate more meaningful and accurate answers. Similarly, based on the answer word information before position t and the input sequence X To calculate the predicted probability of the answer word , using contextual information to guide answer generation, so that the model can generate more logical and semantic answers based on previous information, improving the accuracy and relevance of the answers.

[0093] In some implementations, determining the second loss value based on the third predicted probability of each word-gram in the answer, the first distribution probability of the answer, and the second distribution probability of the answer includes: (1.1) determining a second sub-loss value based on a third predicted probability of each word in the answer; (1.2) determining the KL divergence between the first distribution probability of the answer and the second distribution probability of the answer to obtain a knowledge distillation sub-loss value; (1.3) Perform a weighted sum of the second sub-loss value and the knowledge distillation sub-loss value to obtain a second loss value.

[0094] The second sub-loss value is used to measure the degree of difference between the accuracy of the word unit prediction in the answer of the security task model with adjustable parameters after fine-tuning in the first stage and the actual situation, and is part of the calculation of the second loss value.

[0095] Specifically, the second sub-loss value It can be calculated by the following loss function: ; in, The input sequence for the second stage of fine-tuning is the question and the thinking process. is the total number of tokens included in the answer, is a conditional probability, indicating that given an input sequence In the case of , the model predicts the word at position t in the answer to be probability. represents the real word at position t in the answer.

[0096] Since it is necessary to avoid the phenomenon of knowledge forgetting caused by directly removing the thinking process for fine-tuning, it is necessary to determine the knowledge distillation sub-loss value by the KL divergence between the first distribution probability of the answer and the second distribution probability of the answer , knowledge distillation loss value It can be calculated by the following formula: ; in, is the first distribution probability of the answer predicted by the security task model with adjustable parameters after fine-tuning in the first stage, is an adjustable parameter, The second distribution probability of the answer is predicted for the safety task model of the frozen parameters after fine-tuning in the first stage, is a frozen parameter. Determine the KL divergence between the first distribution probability and the second distribution probability of the answer, and obtain the knowledge distillation loss value to prevent knowledge forgetting .

[0097] Among them, the second loss value It can be calculated according to the following formula: ; in, is the second sub-loss value The weight value of is the knowledge distillation loss value The weight value of the second sub-loss value And the knowledge distillation loss value Perform weighted summation to obtain the final second loss value .

[0098] In this way, the second sub-loss value is used to measure the difference between the accuracy of the prediction of the word units in the answer by the security task model with adjustable parameters after fine-tuning in the first stage and the actual situation. By evaluating the predicted probability of each word unit, the model can more accurately learn the correct expression of each word unit in the answer. The knowledge distillation sub-loss value is determined by calculating the KL divergence between the first distribution probability and the second distribution probability. The first distribution probability comes from the model with adjustable parameters, and the second distribution probability comes from the model with frozen parameters. In this way, the model with adjustable parameters can refer to the knowledge of the frozen parameter model in the process of optimizing the answer, avoiding the phenomenon of knowledge forgetting caused by directly removing the thinking process for fine-tuning. The KL divergence measures the difference between the two probability distributions, and minimizes it. , the model with adjustable parameters can be as close as possible to the predicted distribution of the frozen parameter model, thereby retaining the knowledge learned by the model in the first stage. This ensures that when the model is fine-tuned in the second stage, it will not lose the previously accumulated useful information due to excessive focus on the optimization of the current answer, thus ensuring the stability and generalization ability of the model.

[0099] In some embodiments, after obtaining the fine-tuned security task model, it is necessary to evaluate and test the model performance of the fine-tuned security task model. To ensure the fairness of the evaluation, a third-party evaluation set can be selected for evaluation. Different quantitative calculation methods are used for different types of tasks, and the influence of reference samples is fully considered. In the case of evaluating the performance of the fine-tuned security task model in the mode without reference samples and in the mode with reference samples, the reference sample is sampled from the data set of the same task in the evaluation set in the evaluation mode with reference samples. When assembling the evaluation input prompt, the reference sample is input into the evaluated model as part of it; for classification type tasks, it is necessary to calculate the accuracy after matching the rules according to the actual situation. For generation tasks, the similarity between the output of the fine-tuned model and the standard output is calculated, and the higher the similarity, the better the generation effect. The pre-trained Transform model can be used to vectorize the two input texts to be compared. The BERT model can map natural language text to a high-dimensional vector space, so that the semantic information of the text is fully expressed. In an embodiment of the present application, a SecBERT model that has been fine-tuned in the field of network security can be selected. After obtaining the vector representations corresponding to the two texts, calculate the cosine similarity between the two vectors. Specifically, let the vector of the output of the fine-tuned model be represented as A, and the vector of the standard output be represented as B. The formula is: ; in, represents the dot product of two vectors, and Represent the Euclidean norm of two vectors respectively. The higher the similarity, the better the output effect of the fine-tuned model. In the specific implementation, you can pay attention to the following two points: (1) Vector extraction strategy: The output vector corresponding to the [CLS] tag can be used as the representation of the entire sentence or the average value of all token output vectors can be used. It is recommended to use the output features of the last layer or the last few layers.

[0100] (2) Computational optimization: For long texts, it is recommended to segment them and then take the average value; it is recommended to normalize the vectors.

[0101] As can be seen from the above, the embodiment of the present application generates multiple first security task descriptions, clusters the multiple first security task descriptions, and obtains multiple security task description clusters; obtains multiple security reference data, determines the similarity between each of the security reference data and any first security task description in each of the security task description clusters; determines the target security task description corresponding to each of the security reference data based on the similarity between each of the security reference data and the corresponding first security task description; generates a first input instruction based on each of the security reference data and the corresponding target security task description, the first input instruction is a command to generate an input prompt for question and answer data based on the security reference data and the corresponding target security task description; inputs the first input instruction into to a trained generation model to generate question and answer data corresponding to the first input instruction; the question and answer data is input into the trained security task model to fine-tune the trained security task model to obtain a fine-tuned security task model. Compared with the problem of poor accuracy of answers in specific fields due to lack of high-quality data in the field in related technologies, the embodiment of the present application describes multiple first security tasks and clusters them to obtain a cluster of security task descriptions in a specific field, and generates a first input instruction in combination with reference data, obtains question and answer data for model training in a specific field through the generation model, and trains the security task model that needs to be fine-tuned after training through the question and answer data, thereby improving the accuracy of the model's answers in specific fields.

[0102] The specific implementation of the above steps can be found in the previous embodiments, which will not be described in detail here.

[0103] In order to facilitate better implementation of the model training method provided in the embodiment of the present application, the embodiment of the present application also provides a device based on the above model training method. The meanings of the nouns are the same as those in the above model training method, and the specific implementation details can refer to the description in the method embodiment.

[0104] See also Figure 3 , Figure 3A schematic diagram of the structure of a model training device provided in an embodiment of the present application, wherein the model training device is applied to a computer device. The model training device may include a clustering division unit 601, a first determination unit 602, a second determination unit 603, a generation unit 604, an input unit 605, and a fine-tuning unit 606, etc.

[0105] A clustering division unit 601 is used to generate a plurality of first safety task descriptions, and cluster the plurality of first safety task descriptions to obtain a plurality of safety task description clusters; A first determining unit 602 is configured to obtain a plurality of security reference data and determine a similarity between each of the security reference data and any first security task description in each of the security task description clusters; A second determining unit 603 is used to determine a target security task description corresponding to each security reference data according to a similarity between each security reference data and the corresponding first security task description; A generating unit 604 is used to generate a first input instruction based on each of the security reference data and the corresponding target security task description, wherein the first input instruction is an instruction to generate an input prompt for question and answer data according to the security reference data and the corresponding target security task description; An input unit 605, used to input the first input instruction into the trained generation model to generate question and answer data corresponding to the first input instruction; The fine-tuning unit 606 is used to input the question-answer data into the trained safety task model to fine-tune the trained safety task model to obtain a fine-tuned safety task model.

[0106] In some embodiments, the clustering division unit 601 includes: a conversion subunit, configured to perform vector conversion on each of the first safety task descriptions to obtain a vector representation of each of the first safety task descriptions; A first selection subunit, configured to select a first preset number of candidate safety task descriptions from each of the first safety task descriptions as initial clustering centers of different initial safety task description clusters; A first calculation subunit, configured to calculate a distance between a vector representation of each of the other safety task descriptions and a vector representation of each of the initial cluster centers; A partitioning subunit, used to partition each of the other safety task descriptions into an initial safety task description cluster corresponding to the nearest initial clustering center, to obtain a plurality of updated safety task description clusters; A first determination subunit is used to determine the mean of the vector representation of other safety task descriptions in each of the updated safety task description clusters and the vector representation of the corresponding initial cluster center to obtain the updated cluster center of each of the updated safety task description clusters; The first execution subunit is used to, when the clustering termination condition is not met, determine the updated cluster center as the initial cluster center, determine the updated safety task description cluster as the initial safety task description cluster, and return to execute the step of calculating the distance between the vector representation of other safety task descriptions and the vector representation of each of the initial cluster centers until the clustering termination condition is met, and determine the multiple updated safety task description clusters obtained as safety task description clusters.

[0107] In some embodiments, the cluster division unit 601 further includes: A second calculation subunit is used to calculate the difference between the vector representation of the updated cluster center of each updated security task description cluster and the vector representation of the corresponding initial cluster center; A comparison bullet element is used to compare the difference value of each updated security task description cluster with a preset difference value; A second determining subunit is used to determine that the clustering termination condition is not satisfied when there is an updated safety task description cluster whose corresponding difference value is greater than or equal to the preset difference value; The third determining subunit is used to determine that the clustering termination condition is satisfied when the difference value of each of the updated safety task description clusters is less than the preset difference value.

[0108] In some embodiments, the cluster division unit 601 further includes: A second selection subunit is used to select a second preset number of safety task descriptions from the first safety task description set and the constructed second safety task description set; A generating subunit, configured to generate a second input instruction based on the selected safety task description, wherein the second input instruction is used to command the generation of a third preset number of generated safety task descriptions according to the selected safety task description; A first input subunit, used for inputting the second input instruction into the trained generation model to generate a corresponding generation safety task description; a fourth determining subunit, configured to determine the generated security task description that meets the security task description condition as a first security task description, and assign it to the first security task description set; The second execution sub-unit is used to return to execute the step of selecting a second preset number of security task descriptions from the constructed second security task description set and the first security task description set when the total number of first security task descriptions in the first security task description set does not reach a fourth preset number, until the total number of first security task descriptions in the first security task description set reaches a fourth preset number, so as to obtain multiple first security task descriptions in the first security task description set.

[0109] In some embodiments, the fourth determining subunit is configured to: Generate a third input instruction based on each of the generated safety task descriptions, the third input instruction being used to command an answer as to whether the corresponding generated safety task description is related to the safety task; Input each of the third input instructions into the trained generation model to generate an answer result; The generated safety task description corresponding to the third input instruction for which the answer result is yes is determined as the first safety task description.

[0110] In some embodiments, the question-answer data includes questions, thinking processes, and answers. The fine-tuning unit 606 includes: A second input subunit is used to input the question into the trained security model to obtain a first prediction probability of each word in the thinking process and a second prediction probability of each non-stop word word in the answer; a fifth determination subunit, configured to determine a first loss value based on a first prediction probability of each word-gram in the thinking process and a second prediction probability of each non-stop word word-gram in the answer; A first fine-tuning subunit, configured to perform a first-stage fine-tuning on the trained safety task model according to the first loss value to obtain a safety task model after first-stage fine-tuning; A third input subunit is used to input the question and the thinking process into the security task model with adjustable parameters after fine-tuning in the first stage, so as to obtain a third prediction probability of each word in the answer predicted by the security task model with adjustable parameters after fine-tuning in the first stage, and to predict a first distribution probability of the answer; A fourth input subunit is used to input the question and the thinking process into the safety task model of the frozen parameters after fine-tuning in the first stage, so as to obtain a second distribution probability for predicting the answer; a sixth determining subunit, configured to determine a second loss value based on a third predicted probability of each word in the answer, the first distribution probability of the answer, and the second distribution probability of the answer; The second fine-tuning subunit is used to perform a second-stage fine-tuning on the safety task model with adjustable parameters after the first-stage fine-tuning according to the second loss value to obtain the safety task model after the fine-tuning.

[0111] In some embodiments, the fifth determining subunit is configured to: Obtaining the position weight of each word in the thinking process; Determine a first sub-loss value of the thinking process based on the position weight of each word element in the thinking process and the corresponding first prediction probability; determining a sub-loss value for an answer based on a second predicted probability of a word-gram for each non-stop word in the answer; The first sub-loss value and the answer sub-loss value are weightedly summed to obtain a first loss value.

[0112] In some embodiments, the sixth determining subunit is configured to: determining a second sub-loss value based on a third predicted probability of each word-gram in the answer; Determine a KL divergence between a first distribution probability of the answer and a second distribution probability of the answer to obtain a knowledge distillation sub-loss value; A weighted sum is performed on the second sub-loss value and the knowledge distillation sub-loss value to obtain a second loss value.

[0113] The specific implementation of each of the above units can be found in the previous embodiments, which will not be described in detail here.

[0114] As can be seen from the above, the embodiment of the present application generates multiple first security task descriptions through a clustering division unit 601, clusters the multiple first security task descriptions, and obtains multiple security task description clusters; the first determination unit 602 obtains multiple security reference data, and determines the similarity between each of the security reference data and any first security task description in each of the security task description clusters; the second determination unit 603 determines the target security task description corresponding to each of the security reference data according to the similarity between each of the security reference data and the corresponding first security task description; the generation unit 604 generates a first input instruction based on each of the security reference data and the corresponding target security task description, and the first input instruction is an input prompt for commanding to generate question and answer data according to the security reference data and the corresponding target security task description; the input unit 605 inputs the first input instruction into the trained generation model to generate question and answer data corresponding to the first input instruction; the fine-tuning unit 606 inputs the question and answer data into the trained security task model to fine-tune the trained security task model to obtain the fine-tuned security task model. Compared with the problem in related technologies of poor accuracy of answers in specific fields due to lack of high-quality data in the field, the embodiment of the present application obtains a cluster of security task descriptions in a specific field by clustering multiple first security task descriptions, generates a first input instruction in combination with reference data, obtains question and answer data for model training in the specific field by generating a model, and trains the security task model that needs fine-tuning after training with the question and answer data, thereby improving the accuracy of the model's answers in specific fields.

[0115] The specific implementation of each of the above units can be found in the previous embodiments, which will not be described in detail here.

[0116] Reference Figure 4 , Figure 4 The block diagram of the structure of part of the computer device 1000 for implementing the embodiment of the present disclosure. The computer device 1000 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPUs) 622 (for example, one or more processors) and a memory 632, and one or more storage media 630 (for example, one or more mass storage devices) storing application programs 642 or data 644. Among them, the memory 632 and the storage medium 630 can be temporary storage or permanent storage. The program stored in the storage medium 630 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the computer device 1000. Furthermore, the central processing unit 622 can be configured to communicate with the storage medium 630 to execute a series of instruction operations in the storage medium 630 on the computer device 1000.

[0117] The computer device 1000 may also include one or more power supplies 626, one or more wired or wireless network interfaces 650, one or more input and output interfaces 658, and / or one or more operating systems 641, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc.

[0118] The central processor 622 in the computer device 1000 can be used to execute the model training method of the embodiment of the present disclosure, for example: Generate a plurality of first safety task descriptions, and cluster the plurality of first safety task descriptions to obtain a plurality of safety task description clusters; Acquire a plurality of safety reference data, and determine the similarity between each of the safety reference data and any first safety task description in each of the safety task description clusters; Determining a target security task description corresponding to each security reference data according to a similarity between each security reference data and a corresponding first security task description; Based on each of the security reference data and the corresponding target security task description, a first input instruction is generated, wherein the first input instruction is a command to generate an input prompt for question and answer data according to the security reference data and the corresponding target security task description; Inputting the first input instruction into the trained generation model to generate question and answer data corresponding to the first input instruction; The question-and-answer data is input into the trained safety task model to fine-tune the trained safety task model to obtain a fine-tuned safety task model.

[0119] The embodiments of the present disclosure also provide a computer-readable storage medium, which is used to store program code, and the program code is used to execute the model training methods of the aforementioned embodiments.

[0120] The present disclosure also provides a computer program product, which includes a computer program. The processor of the computer device reads and executes the computer program, so that the computer device executes the above-mentioned model training method. For example: Generate a plurality of first safety task descriptions, and cluster the plurality of first safety task descriptions to obtain a plurality of safety task description clusters; Acquire a plurality of safety reference data, and determine the similarity between each of the safety reference data and any first safety task description in each of the safety task description clusters; Determining a target security task description corresponding to each security reference data according to a similarity between each security reference data and a corresponding first security task description; Based on each of the security reference data and the corresponding target security task description, a first input instruction is generated, wherein the first input instruction is a command to generate an input prompt for question and answer data according to the security reference data and the corresponding target security task description; Inputting the first input instruction into the trained generation model to generate question and answer data corresponding to the first input instruction; The question-and-answer data is input into the trained safety task model to fine-tune the trained safety task model to obtain a fine-tuned safety task model.

[0121] In addition, the terms "comprises" and "includes" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus that includes a series of steps or elements is not necessarily limited to those steps or elements explicitly listed, but may include other steps or elements not explicitly listed or inherent to such process, method, product or apparatus.

[0122] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0123] It should be understood that in the description of the embodiments of the present application, the meaning of multiple (or multiple items) is more than two, greater than, less than, exceed, etc. are understood to not include the number, and above, below, within, etc. are understood to include the number.

[0124] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0125] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0126] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0127] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store program codes.

[0128] It should also be understood that the various implementations provided in the embodiments of the present application can be combined arbitrarily to achieve different technical effects.

[0129] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0130] The above is a specific description of the implementation method of the present application, but the present application is not limited to the above-mentioned implementation method. Technical personnel familiar with the field can also make various equivalent modifications or substitutions without violating the spirit of the present application. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present application.

Claims

1. A model training method, characterized in that: include: Generate a plurality of first safety task descriptions, and cluster the plurality of first safety task descriptions to obtain a plurality of safety task description clusters; Acquire a plurality of safety reference data, and determine the similarity between each of the safety reference data and any first safety task description in each of the safety task description clusters; Determining a target security task description corresponding to each of the security reference data according to a similarity between each of the security reference data and the corresponding first security task description; Based on each of the security reference data and the corresponding target security task description, a first input instruction is generated, wherein the first input instruction is a command to generate an input prompt for question and answer data according to the security reference data and the corresponding target security task description; Inputting the first input instruction into the trained generation model to generate question and answer data corresponding to the first input instruction; The question-and-answer data is input into the trained safety task model to fine-tune the trained safety task model to obtain a fine-tuned safety task model.

2. The model training method according to claim 1, characterized in that: The clustering of the plurality of first safety task descriptions to obtain a plurality of safety task description clusters includes: Performing vector conversion on each of the first safety task descriptions to obtain a vector representation of each of the first safety task descriptions; Selecting a first preset number of candidate safety task descriptions from each of the first safety task descriptions as initial clustering centers of different initial safety task description clusters; Calculating the distance between the vector representation of each of the other security task descriptions and the vector representation of each of the initial cluster centers; Dividing each of the other safety task descriptions into an initial safety task description cluster corresponding to the nearest initial clustering center, to obtain a plurality of updated safety task description clusters; Determine the mean of the vector representations of other safety task descriptions in each of the updated safety task description clusters and the vector representation of the corresponding initial cluster center to obtain the updated cluster center of each of the updated safety task description clusters; When the clustering termination condition is not met, the updated cluster center is determined as the initial cluster center, the updated safety task description cluster is determined as the initial safety task description cluster, and the step of calculating the distance between the vector representation of other safety task descriptions and the vector representation of each of the initial cluster centers is returned to execute until the clustering termination condition is met, and the obtained multiple updated safety task description clusters are determined as safety task description clusters.

3. The model training method according to claim 2, characterized in that: After obtaining the updated cluster center of each updated safety task description cluster, the method further includes: Calculating a difference value between a vector representation of an updated cluster center of each updated safety task description cluster and a vector representation of a corresponding initial cluster center; Comparing the difference value of each updated security task description cluster with a preset difference value; When there is an updated safety task description cluster whose corresponding difference value is greater than or equal to the preset difference value, determining that the clustering termination condition is not satisfied; When the difference value of each of the updated safety task description clusters is less than the preset difference value, it is determined that the clustering termination condition is met.

4. The model training method according to claim 1, characterized in that: The generating of a plurality of first safety task descriptions comprises: Selecting a second preset number of safety task descriptions from the first safety task description set and the constructed second safety task description set; Generate a second input instruction based on the selected safety task description, the second input instruction is used to command the generation of a third preset number of generated safety task descriptions according to the selected safety task description; Inputting the second input instruction into the trained generation model to generate a corresponding generation safety task description; Determine the generated safety task description that meets the safety task description condition as the first safety task description, and assign it to the first safety task description set; When the total number of first security task descriptions in the first security task description set does not reach a fourth preset number, return to execute the step of selecting a second preset number of security task descriptions from the constructed second security task description set and the first security task description set until the total number of first security task descriptions in the first security task description set reaches a fourth preset number, so as to obtain multiple first security task descriptions in the first security task description set.

5. The model training method according to claim 4, characterized in that: The step of determining the generated safety task description that meets the safety task description condition as the first safety task description includes: Generate a third input instruction based on each of the generated safety task descriptions, the third input instruction being used to command an answer as to whether the corresponding generated safety task description is related to the safety task; Input each of the third input instructions into the trained generation model to generate an answer result; The generated safety task description corresponding to the third input instruction for which the answer result is yes is determined as the first safety task description.

6. The model training method according to claim 1, characterized in that: The question-and-answer data includes questions, thinking processes, and answers. The question-and-answer data is input into the trained safety task model to fine-tune the trained safety task model to obtain a fine-tuned safety task model, including: Inputting the question into the trained security model to obtain a first prediction probability of each word unit in the thinking process and a second prediction probability of each non-stop word word unit in the answer; Determining a first loss value based on a first predicted probability of each word-gram in the thought process and a second predicted probability of each non-stop word word-gram in the answer; Performing a first-stage fine-tuning on the trained safety task model according to the first loss value to obtain a safety task model after first-stage fine-tuning; Input the question and the thinking process into the security task model with adjustable parameters after fine-tuning in the first stage, so as to obtain the security task model with adjustable parameters after fine-tuning in the first stage to predict the third prediction probability of each word in the answer and the first distribution probability of the answer; Inputting the question and the thinking process into the safety task model with frozen parameters fine-tuned in the first stage to obtain a second distribution probability for predicting the answer; determining a second loss value based on a third predicted probability of each word-gram in the answer, the first distribution probability of the answer, and the second distribution probability of the answer; The second stage fine-tuning is performed on the safety task model with adjustable parameters after the first stage fine-tuning according to the second loss value to obtain the safety task model after fine-tuning.

7. The model training method according to claim 6, characterized in that: The determining of the first loss value based on the first predicted probability of each word-gram in the thinking process and the second predicted probability of each non-stop word-gram in the answer comprises: Obtaining the position weight of each word in the thinking process; Determine a first sub-loss value of the thinking process based on the position weight of each word element in the thinking process and the corresponding first prediction probability; determining a sub-loss value for an answer based on a second predicted probability of a word-gram for each non-stop word in the answer; The first sub-loss value and the answer sub-loss value are weightedly summed to obtain a first loss value.

8. The model training method according to claim 6, characterized in that: The determining of the second loss value based on the third predicted probability of each word in the answer, the first distribution probability of the answer, and the second distribution probability of the answer comprises: determining a second sub-loss value based on a third predicted probability of each word-gram in the answer; Determine a KL divergence between a first distribution probability of the answer and a second distribution probability of the answer to obtain a knowledge distillation sub-loss value; A weighted sum is performed on the second sub-loss value and the knowledge distillation sub-loss value to obtain a second loss value.

9. A model training device, characterized in that: include: A clustering division unit, used to generate a plurality of first safety task descriptions, and cluster the plurality of first safety task descriptions to obtain a plurality of safety task description clusters; A first determining unit, configured to obtain a plurality of safety reference data, and determine a similarity between each of the safety reference data and any first safety task description in each of the safety task description clusters; A second determining unit, configured to determine a target security task description corresponding to each security reference data according to a similarity between each security reference data and the corresponding first security task description; A generating unit, configured to generate a first input instruction based on each of the security reference data and the corresponding target security task description, wherein the first input instruction is a command to generate an input prompt for question and answer data according to the security reference data and the corresponding target security task description; An input unit, used to input the first input instruction into the trained generation model to generate question and answer data corresponding to the first input instruction; A fine-tuning unit is used to input the question and answer data into the trained safety task model to fine-tune the trained safety task model to obtain a fine-tuned safety task model.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, which are suitable for loading by a processor to execute the model training method described in any one of claims 1 to 8.

11. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the model training method described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Knowledge question and answer model training method, knowledge question and answer method and related device

    CN118520294A

  • Question and answer model training method, object analysis method and related equipment

    CN119474272A

  • Question and answer task processing model training method and device, equipment and storage medium

    CN119493849A

Cited By

  • Financial mixed data intelligent question and answer method and system based on large language model

    CN120723866A

  • Model training method, text processing method and related device

    CN121188470A