A data flywheel fine-tuning method and device

By introducing security and usability assessments into large-scale language models through the data flywheel method, building training datasets and fine-tuning the models, we solved the problem of large-scale language models outputting discriminatory and negative speech, achieved a balance between the security and usability of the model output, and improved the credibility and security of the model.

CN120234408BActive Publication Date: 2025-09-16SHENZHEN HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510705943.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-16
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

Large language models learn biased information from internet data, resulting in the output of discriminatory and negative speech. Complex internal parameter interactions make them difficult to trace and repair, affecting the security and credibility of the output content during the model's life cycle.

Method used

Through the data flywheel method, answers are generated in online and offline modes, security and usability assessments are performed, training datasets are constructed, target models are fine-tuned, security assessment models and usability assessment models are introduced, and model outputs are optimized to meet security and usability requirements.

Benefits of technology

It achieves a balance between the security and usability of model output content, ensures that the model output meets user preferences and security requirements, avoids the generation of harmful content, and improves the credibility and security of model output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234408B_ABST
    Figure CN120234408B_ABST
Patent Text Reader

Abstract

The present application provides a data flywheel fine-tuning method and device, which relate to the field of artificial intelligence technology. The data flywheel fine-tuning method includes: obtaining online answers and offline answers based on online questions and target models. The online answer is the answer generated by the target model to the online question in the online retrieval mode, and the offline answer is the answer generated by the target model to the online question in the offline retrieval mode. The security of the online answer is evaluated, and the first online answer is determined based on the evaluation result. A training data set is constructed using the first online answer and the offline answer. The target model is fine-tuned using the training data set. The method introduces security assessment into the data production, model training, data screening and enhancement, and feedback loop of the data flywheel to guide the content output by the model to meet security and availability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence (AI) technology, and in particular to a data flywheel fine-tuning method and device. Background Art

[0002] Large language models learn from internet data and may inherit biased, misleading, or harmful information, leading to harmful responses such as discriminatory and negative comments. However, as a "black box" model, large language models have complex internal parameter interactions, and the same input can produce different responses. This makes it difficult to trace and repair harmful responses, thus compromising the security and credibility of the content they output throughout their lifecycle. Summary of the Invention

[0003] This application provides a data flywheel fine-tuning method and device, which introduces security assessment into the data production, model training, data screening and enhancement, and feedback loop of the data flywheel, and guides the content of the model output to meet security and usability requirements.

[0004] In the first aspect, a data flywheel fine-tuning method is provided. In this method, a target model is deployed in an online service application. Users enter questions into the online service application. The target model generates offline answers to the online questions in offline retrieval mode and simultaneously generates multiple online answers in online retrieval mode. A security assessment is performed on the content of the multiple online answers. Based on the assessment results, a first online answer is determined from the multiple online answers. A training dataset is constructed using the first online answer and the offline answers, and the target model is fine-tuned using the training dataset.

[0005] In this way, by obtaining the answers output by the target model in offline mode and online mode, online data is collected in the data production link of the data flywheel, rather than just obtaining the answers output in offline mode. By performing a security assessment on the content of the online answers, the security of the content of the online answers is screened, and incremental data fine-tuning of the model is achieved on the Internet to ensure that the output content of the model meets security requirements and avoids the generation of harmful content.

[0006] In one possible implementation, the method further includes obtaining quality and security ratings for the first online answer and the offline answer, and determining the order of the first online answer and the offline answer in a training dataset based on the quality ratings of the first online answer and the offline answer. The training dataset includes the online question, the first online answer, the offline answer, the security rating of the first online answer, and the security rating of the offline answer. In this way, the user's quality rating of the online and offline answers is used to obtain user preferences. The training dataset thus incorporates user preferences for answers and their security ratings. The resulting training dataset indicates user content preferences and security requirements. This training data is then used to fine-tune the model to ensure that the model output meets user preferences and security requirements.

[0007] In one possible implementation, the method further includes: evaluating the security of multiple online answers to obtain a security risk index value and a comprehensive uncertainty for each online answer, and determining a first online answer based on the security risk index value and the comprehensive uncertainty. By evaluating the security risk of the content of multiple online answers, obtaining a probability value for each online answer presenting a security risk, and introducing security uncertainty and availability uncertainty assessments, this method can address the distribution drift between security and availability that can easily occur during the security assessment of online answers, guiding the data flywheel to balance the availability and security of the model output content during operation.

[0008] In one possible implementation, the method further includes: evaluating the security of the online answer using an evaluation model to obtain a security prediction value and a security uncertainty of the online answer. The evaluation model includes a security evaluation model for evaluating the security of the answer, and the security uncertainty indicates the security uncertainty of the answer.

[0009] In one possible implementation, the evaluation model also includes a usability evaluation model, which is used to evaluate the usability of the answer. The method also includes: evaluating the online answer using the usability evaluation model to obtain a usability prediction value and usability uncertainty for the online answer. In this way, the online answer is evaluated using the security evaluation model and the usability evaluation model, respectively. By evaluating the security uncertainty and usability uncertainty in the evaluation results, deviations in the security and usability prediction results of the evaluation model that may occur at different stages or with different data distributions are resolved.

[0010] In one possible implementation, the method further includes: obtaining a security risk index value based on the security prediction value and the security uncertainty; and obtaining a combined uncertainty value based on the availability uncertainty and the security uncertainty. Thus, the security risk index value is obtained using the security prediction value and the security uncertainty representing the degree of doubt about the security conclusion. Compared to predicting security risk using only the security prediction value, this method provides a more accurate online answer to the security risk.

[0011] In one possible implementation, the method further includes: inputting online questions from a training dataset into a target model to obtain a training output answer. The security of the training output answer is evaluated to obtain a usability prediction value and a security prediction value for the training output answer. A loss function is determined using the evaluation results of the training output answer and the training dataset. Based on the loss function, the target model is iteratively trained, where the loss function includes a usability loss function and a security loss function. Thus, using the training dataset and the security prediction value and usability prediction value of the training output answer, the usability loss function and the security loss function are determined. During the model iteration phase, multi-objective optimization of usability and security is used to guide the model output content to improve usability within security constraints.

[0012] In a possible implementation, the method further includes: evaluating the security of the training output answer using an evaluation model to obtain a security prediction value of the training output answer, wherein the evaluation model includes a security evaluation model.

[0013] In one possible implementation, the evaluation model also includes a usability evaluation model. The method further includes: evaluating the usability of the training output answer using the usability evaluation model to obtain a usability prediction value for the training output answer. In this way, using the usability evaluation model and the security evaluation model, during the model fine-tuning phase, the model is guided to output content that is beneficial to usability within security constraints.

[0014] In one possible implementation, the method further includes fine-tuning the evaluation model using the training dataset. Thus, the evaluation model is also fine-tuned using the training dataset, and the output of the target model is evaluated using the fine-tuned evaluation model, resulting in a more accurate result.

[0015] In a second aspect, a data flywheel fine-tuning device is provided, comprising a dataset expansion module and a model fine-tuning module. The dataset expansion module is used to obtain online and offline answers based on online questions and a target model. The online answers are the answers generated by the target model to online questions in online retrieval mode, and the offline answers are the answers generated by the target model to online questions in offline retrieval mode. The security of the online answers is evaluated, and based on the evaluation results, a first online answer is determined. A training dataset is constructed using the first online and offline answers. The model fine-tuning module is used to fine-tune the target model using the training dataset.

[0016] In one possible implementation, the dataset expansion module specifically obtains evaluation results of the first online answer and the offline answer, wherein the evaluation results include quality evaluation values ​​and safety evaluation values ​​of the first online answer and the offline answer. Based on the evaluation results, a training dataset is established, wherein the training dataset indicates a comparison result of the quality evaluation values ​​of the first online answer and the offline answer, and the training dataset includes the online question, the first online answer, the offline answer, the safety evaluation value of the first online answer, and the safety evaluation value of the offline answer.

[0017] In one possible implementation, the dataset expansion module is specifically configured to obtain a security risk index value and a comprehensive uncertainty of the online answer based on the evaluation results. The security risk index value indicates the probability that the answer presents a security risk, and the comprehensive uncertainty indicates a combined quantitative value of the answer's usability uncertainty and security uncertainty. A first online answer is determined based on the security risk index value and the comprehensive uncertainty.

[0018] In one possible implementation, the dataset expansion module is specifically used to evaluate the security of online answers using an evaluation model to obtain a security prediction value and security uncertainty of the online answers; wherein the evaluation model includes a security evaluation model, which is used to evaluate the security of the answers, and the security uncertainty indicates the security uncertainty of the answers.

[0019] In one possible implementation, the evaluation model also includes a usability evaluation model, which is used to evaluate the usability of the answer. The dataset expansion module is specifically used to evaluate the online answer using the usability evaluation model to obtain the usability prediction value and usability uncertainty of the online answer.

[0020] In a possible implementation, the data set expansion module is specifically configured to obtain a security risk index value based on the security prediction value and the security uncertainty, and to obtain a comprehensive uncertainty based on the availability uncertainty and the security uncertainty.

[0021] In one possible implementation, the model fine-tuning module is specifically configured to input online questions from a training dataset into a target model to obtain training output answers. The security of the training output answers is evaluated to obtain evaluation results for the training output answers, which include a usability prediction value and a security prediction value. The evaluation results for the training output answers and the training dataset are used to determine a loss function. The target model is then iteratively trained based on the loss function, which includes a usability loss function and a security loss function.

[0022] In one possible implementation, the model fine-tuning module is specifically used to evaluate the security of the training output answer using an evaluation model to obtain a security prediction value of the training output answer, and the evaluation model includes a security evaluation model.

[0023] In a possible implementation, the model fine-tuning module is specifically configured to evaluate the usability of the training output answer using the usability evaluation model to obtain a predicted usability value of the training output answer.

[0024] In one possible implementation, the model fine-tuning module is specifically configured to fine-tune the evaluation model using the training dataset.

[0025] In a third aspect, the present application provides a computing device cluster comprising at least one computing device, each computing device comprising a processor and a memory; the processor of at least one computing device is used to execute instructions stored in the memory of at least one computing device, so that the computing device cluster performs the method described in the first aspect or any possible implementation of the first aspect.

[0026] In a fourth aspect, the present application provides a computer-readable storage medium comprising computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method described in the first aspect or any possible implementation of the first aspect. For example, the computing device cluster may include one or more computing devices.

[0027] In a fifth aspect, the present application provides a computer program product comprising instructions that, when executed by a computing device cluster, cause the computing device cluster to perform the method described in the first aspect or any possible implementation of the first aspect. For example, the computing device cluster may include one or more computing devices.

[0028] The beneficial effects of the second to fifth aspects can be referred to the above introduction to the beneficial effects of the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 This is a technical concept diagram of a data flywheel fine-tuning method provided by an embodiment of the present application;

[0030] Figure 2 This is another technical concept diagram of a data flywheel fine-tuning method provided in an embodiment of the present application;

[0031] Figure 3 This is a schematic diagram of the data flywheel fine-tuning service provided in an embodiment of the present application deployed on a cloud computing platform;

[0032] Figure 4 This is a flow chart of a data flywheel fine-tuning method provided in an embodiment of the present application;

[0033] Figure 5 This is a flow chart of a data flywheel fine-tuning method provided in an embodiment of the present application;

[0034] Figure 6 This is a schematic diagram of input data in the data flywheel fine-tuning method provided in an embodiment of the present application;

[0035] Figure 7 Schematic diagram of a specific implementation of the data flywheel fine-tuning method provided in an embodiment of the present application;

[0036] Figure 8 Schematic diagram of the composition of the data flywheel fine-tuning device provided in an embodiment of the present application;

[0037] Figure 9 A schematic diagram of the structure of a computing device provided in an embodiment of the present application;

[0038] Figure 10 A schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;

[0039] Figure 11 A schematic diagram of the structure of another computing device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION

[0040] The following describes the solutions provided by the embodiments of the present application in conjunction with the accompanying drawings. In the embodiments of the present application, "plurality" refers to two or more. "First," "second," and the like are merely used to distinguish similar objects and are not necessarily used to describe a specific order or number of objects.

[0041] To facilitate understanding of the solutions provided by the embodiments of the present application, the technical terms that may be involved in the embodiments of the present application are first introduced.

[0042] A large language model (LLM) is a computer model capable of processing and generating natural language. LLM is an artificial intelligence technology based on deep learning. It can predict the next word or sentence by learning the statistical patterns and semantic information of language data. LLMs typically consist of multiple neural networks (such as feedforward neural networks and recurrent neural networks). Feedforward neural networks are used to extract features from input text, while recurrent neural networks are used to model text sequences. By continuously training on large amounts of text data, LLMs can gradually improve their understanding and prediction capabilities of natural language.

[0043] Reinforcement Learning from Human Feedback (RLHF) is a technology that aligns large language models with human preferences. A reward model is trained based on human preference data to model human tendencies in responding to the model. A reinforcement learning algorithm is then used to optimize the language model to maximize cumulative rewards, thereby aligning the language model output with human preferences.

[0044] The Data Flywheel refers to the dynamic interaction and iterative optimization of data and models during the training and application of large models, forming a mutually reinforcing closed-loop system. The Data Flywheel includes the following steps: data production, model training, data screening and enhancement, and a feedback loop.

[0045] In order to ensure the security and credibility of the content output by the large language model during its life cycle. Among the existing technical solutions, one solution is to adopt an offline fine-tuning method based on secure reinforcement learning, which optimizes the large model in combination with a pre-collected human feedback dataset. The core of this method is to collect human feedback data on offline data, and use the human feedback data to train the reward model to fine-tune the large model. The offline fine-tuning method based on secure reinforcement learning is adopted here. Through human feedback reward modeling and security constraint optimization, it can effectively solve the security risks of the large model in the fine-tuning process and will not cause real risks. However, when the large language model is applied to the data flywheel and the incremental data generated by the Internet is used to fine-tune the large language model, it may output content that is not guaranteed to be safe to users, which poses a high real risk.

[0046] Another approach is the data flywheel approach, which doesn't explicitly address security. This approach accumulates data from the internet, evaluates it using task metrics, and then blends data from multiple sources. This multi-source data is then used to fine-tune the large language model. This approach allows for continuous growth in the size of the training data, allowing the updated training data to be used to fine-tune the large language model. This approach can acquire more high-quality, novel data, helping to enhance the capabilities of the large model. However, this approach evaluates data solely using task metrics during the data accumulation process, leading to the inclusion of harmful content during the data fusion phase. Subsequently, the large language model is fine-tuned using this contaminated data. If strict security requirements are imposed on the output content during the fine-tuning process, high-quality, novel data cannot be produced, resulting in a waste of manpower and resources. Increasing the novelty of the output content compromises its security, ultimately making it impossible to guarantee the security of the output data from the fine-tuned model.

[0047] The aforementioned methods have explored ways to ensure the security of the output content of large language models. However, these methods lack the accumulation of secure and effective online data, fail to measure security during the fine-tuning phase, and fail to balance the security and novelty of the output content.

[0048] In view of this, the present invention provides a method for fine-tuning a data flywheel. In this method, security metrics and security constraints are introduced throughout the data flywheel. In light of the security and novelty of the output content, uncertainty assessment is introduced to guide the data flywheel in exploring and fine-tuning the output content, balancing security and usability. Ultimately, a large language model with improved security and usability is obtained.

[0049] For example, Figure 1 A technical concept diagram of a data flywheel fine-tuning method is shown. Figure 1 As shown, the architecture of the data flywheel fine-tuning method mainly includes: an online data acquisition section, an online answer evaluation section, a training dataset generation section, a training evaluation section, and a model fine-tuning section. The online data acquisition section mainly obtains the input online questions and the answers output by the target model. Online questions are input instructions or questions used to guide the model to generate specific types of answers. The answers output by the target model include multiple online answers generated in the online exploration mode and offline answers generated in the offline mode. Exemplarily, in this section, the online data acquisition module 110 can be used to obtain online questions and answers.

[0050] The online answer evaluation section primarily performs security and usability assessments on online answers to obtain evaluation results. Using these assessment results, the first online answer is selected, which is the answer that best balances security and usability among all online answers. In some embodiments, the evaluation results may include a security prediction value and a usability prediction value for each online answer. They may also include a security uncertainty and a usability uncertainty for each online answer. The security prediction value is a quantitative assessment indicator of the content's security risk. The usability prediction value is a quantitative assessment indicator of whether the content meets usability requirements. The security uncertainty measures the inaccuracy of the security prediction value for the content. The usability uncertainty measures the inaccuracy of the usability prediction value for the content. Exemplarily, in this section, the online answer evaluation module 120 may evaluate n online answers to obtain the online answer that best balances security and usability. In some embodiments, the n online answers may also be input into an evaluation model to obtain the online answer evaluation results. The evaluation model may include a security assessment model and a usability assessment model. The usability assessment model may utilize a reward model. The reward model (RM) is the reward-guided policy optimization process of the RLHF. The reward model takes {user question, model output} as input and outputs a scalar score, which is used to evaluate the quality of the responses generated by the policy model. The security assessment model can be implemented using a cost model. The cost model (CM) takes {user question, model output} as input and outputs a scalar score, which is used to evaluate the security of the responses generated by the model.

[0051] The training dataset generation module is primarily used to obtain evaluation results for the first online answer and the offline answer, thereby obtaining a user-level evaluation of their security and usability. A training dataset is then constructed based on the user's evaluation. For example, the training dataset generation module 130 can obtain the user's selected answer with the highest quality from the first online answer and the offline answer, while the other answer is designated as the second-best answer. The security of each answer is then evaluated. This results in a training data record. This training data record includes the online question, the highest quality answer, the second-best answer, and the security evaluation value of the highest quality answer and the second-best quality answer. Thus, during the training dataset generation phase of the data flywheel, online data generated in a networked manner is obtained, ensuring its novelty. The online answers are evaluated for security and usability, and online answers that demonstrate both security and usability are selected. Furthermore, these online answers are used as the training dataset, rather than solely using offline data. This ensures the novelty, security, and usability of the output content. Furthermore, the training dataset also includes user preferences and user assessments of the security of the output content.

[0052] The training evaluation component primarily involves inputting online questions from the training dataset into the target model, obtaining the target model's output answers, and then evaluating the usability and security of these output answers. For example, the training evaluation module 140 can select any online questions from the training dataset, input these online questions into the target model, and obtain the target model's output answers. Due to the complex interactions between internal parameters in large language models, these output answers may differ from both previous online and offline answers. The evaluation model is then used to evaluate these output answers, obtaining their predicted security and usability values.

[0053] The model fine-tuning component primarily utilizes the security prediction value and availability prediction value of the output answer to derive a loss function for the output answer. This loss function is then used to fine-tune the target model, taking into account the dual optimization goals of availability and security. For example, the target model can be fine-tuned using the model fine-tuning module 150.

[0054] Based on the above solution framework, by evaluating the security and usability of online content during dataset generation, training the model using a training dataset that balances security and usability, and evaluating the security and usability of the generated content during training, we ensure that both security and usability are taken into account during model fine-tuning. This implements security monitoring and evaluation throughout the data flywheel, ensuring that the model's output content always meets security requirements and avoids the generation of harmful content. Furthermore, to address the need for balancing security and usability, we guide the data flywheel to balance security and usability through assessments of security uncertainty and usability uncertainty, resulting in a model that aligns effectiveness and security.

[0055] In order to improve the accuracy of the evaluation of the security and availability of the output content, the above solution framework can also be used to add a fine-tuning part of the evaluation model. Figure 2 The program framework shown. Figure 2 Another technical concept diagram of a data flywheel fine-tuning method is shown. Figure 2 The architecture of the data flywheel fine-tuning method mainly includes: online data acquisition part, online answer evaluation part, training dataset generation part and training evaluation part, model fine-tuning part and evaluation model fine-tuning part. Figure 2 The online data acquisition part, online answer evaluation part, training dataset generation part, training evaluation part and model fine-tuning part can be found in the above Figure 1 The relevant description in will not be described here.

[0056] exist Figure 2The evaluation model fine-tuning section is primarily used to fine-tune the evaluation model using the training dataset. In some embodiments, a loss function can be established for the evaluation module. For example, the evaluation model fine-tuning module 160 can establish a loss function for the usability evaluation model. The usability prediction values ​​of the superior answer and the suboptimal answer in the training dataset are input into the loss function of the security evaluation model to obtain a loss value. This loss value reflects the difference between the usability prediction results of the usability evaluation model for the answer and the user's preference for the answer. The parameters of the usability evaluation model are optimized by minimizing the loss value. A loss function is also established for the security evaluation model. The security prediction values ​​of the superior answer, the security prediction values ​​of the suboptimal answer, the security evaluation values ​​of the superior answer, and the security evaluation values ​​of the suboptimal answer in the training dataset are input into the loss function of the security evaluation model to obtain a loss value. This loss value reflects the difference between the security prediction results of the security evaluation model for the answer and the user's security evaluation of the answer. The parameters of the security evaluation model are optimized by minimizing the loss value. Finally, the fine-tuned evaluation model can be updated to the evaluation model in the online answer evaluation section. Furthermore, the fine-tuned evaluation model can also be updated to the evaluation model used in the training evaluation section.

[0057] based on Figure 2 The solution framework shown here corrects erroneous predictions by fine-tuning the evaluation model. This fine-tuned evaluation model is then used to collect training data and fine-tune the target model, improving the accuracy of the security and usability assessments of the output content. Ultimately, a model with improved security and usability is achieved.

[0058] In some embodiments, each part of the solution framework described above can be configured on a cloud computing platform. For example, it can be deployed on at least one virtual machine or container instance, so that the cloud computing platform can provide data flywheel fine-tuning services for the model. Of course, each part of the solution framework described above can also be configured on nodes other than the cloud computing platform. For example, it can be deployed in at least one data center, or on at least one server. The specific details can be determined according to actual conditions and are not limited here. Among them, the cloud computing platform can provide pages related to public cloud services for users to remotely access public cloud services. In this embodiment, users can purchase data flywheel fine-tuning services on the cloud computing platform in advance. For ease of understanding, the interaction between users and cloud computing platforms is described below. As Figure 3As shown, the user's interaction with the cloud computing platform primarily involves logging into the cloud computing platform 300 through a client webpage, selecting and purchasing the Data Flywheel Tuning service within the cloud computing platform 300. After purchase, the user can collect training datasets on the cloud computing platform 300, leveraging the Data Flywheel Tuning service's functionality to ensure both security and novelty of output content. The cloud computing platform 300 primarily manages the infrastructure for running the Data Flywheel Tuning service. For example, the infrastructure for running the Data Flywheel Tuning service may include multiple data centers located in different regions, each containing multiple servers. Data centers may provide basic resources for the Data Flywheel Tuning service, such as computing resources and storage resources. Therefore, when purchasing and using the Data Flywheel Tuning service, users primarily pay for the resources used. When using the Data Flywheel Tuning service, users can specify tuning frequencies and other parameters through the configuration interface, application programming interface, or user interaction interface provided by the cloud computing platform 300. The cloud computing platform 300 then performs Data Flywheel Tuning on the large language model according to the user's instructions. In some embodiments, each part of the solution framework described above may also be configured on a local server, and the specific configuration may depend on the actual situation and is not limited here.

[0059] The specific implementation process of the above technical concept is described below.

[0060] See also Figure 4 , Figure 4 A flow chart of a data flywheel fine-tuning method is shown. It is understood that the method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities. Exemplarily, the method can be executed by a data flywheel fine-tuning device, wherein the device can be implemented by software and / or hardware, and can be configured in an electronic device or a server, typically, on a server. Figure 4 As shown, the data flywheel fine-tuning method may include the following steps:

[0061] S401. Obtain online questions and online answers.

[0062] In the embodiment of the present application, the online question can be input by the user or sent by a device, apparatus or service, etc., which is not limited here.

[0063] The target model can generate one offline answer in offline mode or multiple online answers in online mode.

[0064] For example, a probability p can be set. When the question input this time falls under the probability p, the target model generates k online answers in the network mode. In this way, the operating performance of the target model deployed in the Internet environment will not be affected.

[0065] S402: Perform security assessment on the online answer and output the security assessment result.

[0066] In embodiments of the present application, security matching rules and security risk keywords can be established through natural language processing methods. Online answers can be matched using the security matching rules and security risk keywords to obtain security assessment results for the online answers. Similarly, online answers can be matched using usable keywords and usability matching rules to obtain usability assessment results. For example, online answers can be evaluated using the evaluation model described in the online answer evaluation section above to obtain usability and security assessment results for the online answers.

[0067] S403: Determine a first online answer based on the security evaluation result of the online answers.

[0068] In the embodiment of the present application, after obtaining the security assessment results of each online answer, pre-set criteria can be used to screen out the first online answer whose assessment results meet the criteria. For example, when the assessment results are represented by an assessment score, online answers with a security assessment score greater than or equal to a security threshold can be selected, and from these online answers, an online answer with a usability assessment score greater than or equal to the usability threshold can be selected as the first online answer.

[0069] S404: Construct a training data set using the first online answer and the offline answer.

[0070] In an embodiment of the present application, a training data record may include an online question, a first online answer, an offline answer, and predicted usability and security values ​​for the first online answer and the offline answer. Exemplarily, a user's evaluation results of the first online answer and the offline answer are obtained, and a training dataset is established based on the evaluation results. In this example, the first online answer and the offline answer are simultaneously output to the user. The user evaluates the first online answer and the offline answer based on their preferences, selecting the answer they consider to be of higher quality, and the other answer is marked as a suboptimal answer. The user then evaluates the security of the higher quality answer and the suboptimal answer, respectively. Thus, a training data record is obtained. This training data record includes the online question, the higher quality answer, the suboptimal answer, the user-side security evaluation value of the higher quality answer, and the user-side security evaluation value of the suboptimal answer. In an embodiment of the present application, the online answer and the offline answer can also be input into a large language model to evaluate the quality and security of the answers using the large language model.

[0071] S405: Fine-tune the target model using the training dataset.

[0072] In an embodiment of the present application, online questions in a training data set are input into a target model to obtain training output answers. The training output answers are evaluated for usability and security to obtain evaluation results. Based on the evaluation results, a loss function is constructed with the output content achieving usability and security as the goal, and the target model is updated using the loss function. Exemplarily, the training output answers can be evaluated using the method described in step 402. The loss function is set based on the predicted security and usability values ​​of the training output answers and the predicted security and usability values ​​of the better answers in the training data set, and the parameters in the target model are adjusted using the loss function.

[0073] In this way, by screening the training dataset for security and usability during its generation, using this training dataset to train the model, evaluating the security and usability of the generated content during training, and also considering security and usability during model fine-tuning, we achieve security monitoring and assessment throughout the entire data flywheel, ensuring that the model's output content always meets security requirements and avoids the generation of harmful content.

[0074] In some embodiments, see also Figure 5 . Figure 5 FIG. 1 shows a flow chart of a data flywheel fine-tuning method. Figure 5 As shown, in the dataset generation step 501, online questions are input into the target model to obtain offline answers generated in offline mode and multiple online answers generated in online mode. The n online answers are respectively input into the security assessment model and the usability assessment model. Each online answer is evaluated using the security assessment model to obtain a security prediction value and security uncertainty for each online answer. Each online answer is evaluated using the usability assessment model to obtain a usability prediction value and usability uncertainty for each online answer. The security prediction value is a quantitative assessment indicator of the security risk of the content. The security uncertainty is a measure of the inaccuracy of the security prediction value of the content. The usability prediction value is a quantitative assessment indicator of whether the content meets the user's online questions. The usability uncertainty is a measure of the inaccuracy of the usability prediction value of the content.

[0075] In some embodiments, the usability evaluation model can be implemented through a reward model. The reward model (RM) is a reward-guided strategy optimization process of the RLHF. The reward model takes {user questions, model outputs} as input and outputs a scalar score to evaluate the quality of the responses generated by the strategy model. In this embodiment of the application, {online questions, online answers} are used as inputs of the usability evaluation model, and the output is the usability prediction value of the online answers. and uncertainty about the availability of online answers . ‌Availability uncertainty‌ is an indicator that can measure the credibility or reliability of online answers, reflecting the range in which online answers may deviate from the true value in terms of availability. The security assessment model can be implemented through a cost model. The cost model (Reward Model, CM) takes {user questions, model output} as input and outputs a scalar score to evaluate the security of the model-generated responses. In the embodiment of the present application, {online questions, online answers} are used as the input of the security assessment model, and the security prediction value of the online answer is output. and the security uncertainty of online answers Security uncertainty can be used to measure the security level or security reliability of online answers, reflecting the range in which online answers may deviate from the true value in terms of security.

[0076] For example, the online answers are evaluated using the usability evaluation model to obtain the usability prediction value and usability uncertainty of the online answers. Input the reward model separately to get the online answer Availability prediction value and availability uncertainty .

[0077] In some embodiments, a security risk index value can also be obtained based on the security prediction value and the security uncertainty. A comprehensive uncertainty can be obtained based on the availability uncertainty and security uncertainty of the online answer. For example, the security prediction value and safety uncertainty Enter the security risk index value calculation formula Get online answers security risk index value. is the cumulative distribution function of the standard normal distribution. and safety uncertainty Enter the comprehensive uncertainty calculation formula , and get the comprehensive uncertainty of the online answer. Among them, is a weighting factor used to balance availability and security.

[0078] like Figure 5As shown, in data set generation step 501, online answers whose security risk index values ​​are below the risk tolerance threshold are selected. Among the eligible online answers, the online answer with the largest comprehensive uncertainty index value is selected as the first online answer. The first online answer and the offline answer are presented to the user. Based on the user's quality evaluation of the first online answer and the offline answer, the superior and suboptimal answers are determined. The user's security evaluation value for the superior answer and the suboptimal answer are obtained. The (online question, superior answer, suboptimal answer, user-side security evaluation value for the superior answer, user-side security evaluation value for the suboptimal answer) is recorded as a training data item.

[0079] Exemplary, offline answers With the first online answer Presented to the user, the user selects the one with higher quality according to their own preference and records it as , and the other is Then, the user evaluates whether the two results are safe, denoted as and Integrate the above data to get the quintuple ;

[0080] In addition, all collected training data records can be appended to the training dataset, and the initial dataset It is also added to the training dataset, where the initial dataset is a dataset generated offline.

[0081] In the model training step 502, as Figure 5 As shown, online questions from the training dataset are input into the target model to obtain training output answers. The training output answers are evaluated for usability and security using the evaluation model to obtain evaluation results. Based on the evaluation results, a loss function is constructed to ensure that the output content of the target model achieves usability and security. This loss function is then used to update the target model.

[0082] In an embodiment of the present application, the training output answer is evaluated using a usability evaluation model to obtain a usability prediction value of the training output answer. The training output answer is evaluated using a security evaluation model to obtain a security prediction value of the training output answer. Based on the usability prediction value and the security prediction value, an advantage function is obtained. The advantage function can be used to determine a quantitative indicator value of the relative advantage of the training output answer in terms of usability and security relative to a better answer. The advantage function is used to determine a loss function. The loss function is used to fine-tune the target model to achieve dual optimization objectives of security loss and usability loss, ultimately obtaining a target model that takes both security and usability into account.

[0083] For example, a cost evaluation model can be used to evaluate the availability cost of the training output answer, obtaining an availability cost evaluation value. The availability cost evaluation value and the predicted availability value of the training output answer are used to generate a first advantage function. This first advantage function can be used to quantify the relative usability advantage of the training output answer over a superior answer. Based on the first advantage function, an availability loss function is determined. The cost evaluation model can be used to evaluate the security cost of the training output answer, obtaining a security cost evaluation value. The security cost evaluation value and the predicted security value of the training output answer are used to calculate a second advantage function. This second advantage function can be used to quantify the relative security advantage of the training output answer over superior answers in the training dataset. The second advantage function is used to generate a security loss function. The security loss function and the availability loss function are used to generate a total loss function that balances both security and usability. This total loss function is used to fine-tune the target model, thereby obtaining a target model that balances both the usability and security of the output content. The fine-tuned target model can then be updated to the online service. The cost evaluation model can be used to quantify model prediction error, specifically, to obtain the prediction error for the content's usability prediction value and the prediction error for the content's security prediction value.

[0084] For example, online questions are extracted from the training dataset, and the online questions are input into the target model to obtain the output answers. The output answers are evaluated through the reward model and the cost model to obtain the availability prediction value and the security prediction value.

[0085] The Generalized Advantage Estimation (GAE) method is used to obtain the safety advantage function based on the availability prediction value and the availability estimate value. Using the safety prediction value and safety estimate value, we can get the safety advantage function .

[0086] Calculate the availability loss using the same method used in the Proximal Policy Optimization (PPO) algorithm. and security loss .in,

[0087]

[0088] ,

[0089] in, is the expectation of the current policy distribution, is the policy update ratio, that is, the probability ratio of the new policy to the old policy. are the parameters of the policy network. is the clipping threshold, which is used to control the update amplitude. is the clipping function. It's an online problem. is the training dataset, It is the training output answer obtained by inputting the online answer into the target model.

[0090] Introducing Lagrange multipliers , construct the Lagrangian function to realize the total loss function of multiple optimization objectives for availability and security .in,

[0091] In the process of optimizing the target model using the total loss function, the total loss function can be used to , update the target model. Or update the Lagrange multiplier , ,in is the update rate.

[0092] By evaluating the security and usability of the model output during training and establishing a loss function for dual-objective optimization of usability and security, we achieved a balance between security and usability during model fine-tuning, ultimately realizing security monitoring and evaluation of the data flywheel at all stages.

[0093] In some embodiments, an evaluation model training step 503 is also included. Figure 5 As shown, in the evaluation model training step 503, the evaluation model is fine-tuned using the training data set. For example, a loss function is established for the usability evaluation model. The usability prediction values ​​of the superior answer and the suboptimal answer in the training data set are input into the loss function of the first evaluation model to obtain a loss value. This loss value reflects the difference between the usability evaluation model's usability prediction results for the answer and the user's preference for the answer. The parameters of the usability evaluation model are optimized by minimizing the loss value. In this example, a loss function is established for the security evaluation model. The security prediction values ​​of the superior answer, the suboptimal answer, the security evaluation values ​​of the superior answer, and the security evaluation values ​​of the suboptimal answer in the training data set are input into the loss function of the security evaluation model to obtain a loss value. This loss value reflects the difference between the security evaluation model's security prediction results for the answer and the user's security evaluation of the answer. The parameters of the security evaluation model are optimized by minimizing the loss value.

[0094] For example, the Bradley-Terry model can be used to model the evaluation model. The Bradley-Terry model is a classic probability model used to model the preference relationship between paired data. It assigns a positive real number score to each option to represent its relative preference strength and calculates the relative preference probability between two options. For example, for option and , preference Exceed The probability is:

[0095] in and Respectively options and When training the model, the parameters are optimized using the negative log-likelihood loss function:

[0096] The loss function of the usability evaluation model is:

[0097]

[0098] in,

[0099]

[0100] Loss function of the security assessment model:

[0101]

[0102] in, is the predicted probability, is a linear combination of the input values.

[0103] In summary, we use the loss function to fine-tune the security and usability evaluation models, correcting the errors in the evaluation model's predictions. The fine-tuned evaluation model is then used to fine-tune the target model, ultimately resulting in a model with improved security and usability.

[0104] The above is an introduction to the data flywheel fine-tuning method provided in this application. In some embodiments, this method can be combined with blockchain technology. The blockchain records the generation and use of training data, ensuring the authenticity and immutability of the data. Within blockchain smart contracts, automatic updates of the reward model (RM) and cost model (CM) can be implemented. For example, when new data is added to the dataset, the smart contract can trigger model fine-tuning to ensure the model remains optimal.

[0105] Figure 6 Schematic diagram of input data in the data flywheel fine-tuning method in the embodiment of the present application. Figure 6As shown, the online questions or online answers in the data flywheel fine-tuning device in the embodiment of the present application can also be multimodal data, such as audio or video data. The embodiment of the present application may use a lightweight adapter to fine-tune the target model instead of directly training the entire target model. Federated learning technology can also be used to allow multiple clients such as different companies or user groups to jointly train the model without sharing data. In this way, the generalization ability of the model can be improved while protecting data privacy.

[0106] It is understood that the order of execution of the steps in the above embodiments does not necessarily imply a specific order of execution. The order of execution of each process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of this application. In addition, the various embodiments and technical features described above can be combined according to actual circumstances, and the combined solutions are still within the scope of protection of this application.

[0107] Next, based on the method in the above embodiment, the specific implementation method of the data flywheel fine-tuning method provided in the embodiment of the present application is introduced. Figure 7 Schematic diagram of a specific implementation of the data flywheel fine-tuning method provided in the embodiment of the present application. Figure 7 As shown, the online question is: "How do I protect the security of my social media account?" The large language model outputs the following online answer:

[0108] Online answer 1: There is no need for protection at all. Ordinary people’s accounts are not worth protecting anyway.

[0109] Online answer 2: Bind your phone or authentication app and enable two-factor authentication.

[0110] Online answer 3: Do not click on private message links from unknown sources.

[0111] Online answer 4: Rely on the platform, the platform will take care of everything for you.

[0112] Use the evaluation model to evaluate the online answers. Get the evaluation results of the online answers. The evaluation results are as follows:

[0113] Online answer 1: Safety prediction value 0, availability prediction value 0, safety uncertainty 0.6, availability uncertainty 0.8; Online answer 2: Safety prediction value 1, availability prediction value 2, safety uncertainty 0.8, availability uncertainty 0.8; Online answer 3: Safety prediction value 1, availability prediction value 1, safety uncertainty 0.8, availability uncertainty 0.6; Online answer 4: Safety prediction value 0, availability prediction value 0, safety uncertainty 0.5, availability uncertainty 0.5.

[0114] From the two online answers 2 and 3 that meet the security constraints, online answer 2 with the highest usability evaluation score is selected as the first online answer. Online answer 2 and the offline answer are displayed to the user, who then evaluates their quality and security. For the offline answer, a strong password policy is recommended: set the password length to ≥ 12 characters, mix uppercase and lowercase letters, numbers, and symbols, and prohibit password reuse. Based on the user's evaluation results, a training data record is constructed.

[0115] For the online question, the user prefers online answer 2, making it the superior answer. The offline answer is the suboptimal answer. Next, the user scores the security of online answer 2 and the offline answer. A training data record is obtained: (How to prevent my social media account from being hacked, online answer 2, offline answer 1, 1).

[0116] During model training, the online question "How do I prevent my social media account from being hacked" from the training dataset is fed into the large language model, which then outputs the answer. The answer is then fed into the evaluation model to obtain the security and usability predictions for the answer. The security and usability predictions of the output answers are used to determine usability and security loss functions. Multiple optimization objectives are applied to the usability and security loss functions to obtain a loss function that balances usability and security. The resulting loss function is then used to fine-tune the large language model.

[0117] The above is an example of the data flywheel fine-tuning method in practical application of the embodiment of the present application. Based on the method in the above embodiment, a data flywheel fine-tuning device provided by the embodiment of the present application is introduced. Figure 8 A schematic diagram of the composition of a data flywheel fine-tuning device is shown. Figure 8 As shown, the data flywheel fine-tuning device 800 includes a dataset expansion module 810 and a model fine-tuning module 820. The dataset expansion module 810 primarily receives input online questions and generates multiple answers based on the target model in the online exploration mode and the offline answers generated in the offline mode. It then evaluates the security and usability of the multiple online answers to obtain an online answer evaluation result. Based on the evaluation results, it determines an online answer that balances security and usability. The training dataset is expanded based on the online and offline answers, and the expanded dataset is output. The model fine-tuning module 820 uses the expanded dataset to fine-tune the target model.

[0118] In an embodiment of the present application, the target model may be a policy model. The policy model is a large language model to be optimized in RLHF, which takes user questions as input and outputs model responses. The ultimate goal of RLHF is to optimize the policy model so that it outputs high-quality text and aligns with human preferences. Among them, the target model may be any neural network model that can be deployed to an online service application. The embodiment of the present application does not limit the specific implementation method of the online service application function. For example, the target model can be deployed to an online consultation application to respond to online questions raised by users.

[0119] like Figure 8 As shown, the dataset expansion module 810 primarily includes an online answer evaluation module, a first online answer generation module, and a dataset user evaluation module. The online answer evaluation module uses an evaluation model to evaluate the usability and security of multiple online answers. The first online answer generation module, based on the usability and security evaluation results of the multiple online answers, determines a first online answer that balances usability and security and achieves the optimal quantitative usability and security indicators among all online answers. The dataset user evaluation module expands the training dataset based on user evaluations of the offline answers and the first online answer.

[0120] In the embodiments of this application, the evaluation model can be any neural network model capable of evaluating the usability and security of text. The embodiments of this application do not limit the specific implementation of the usability and security evaluation. For example, the evaluation model can be a reward model, a cost model, etc., which can quantify the usability and security of the answers output by the large language model.

[0121] In the embodiments of the present application, the evaluation model includes a security evaluation model and a usability evaluation model. In the online answer evaluation module, the usability evaluation model is used to evaluate online answers, obtaining the usability prediction value and usability uncertainty of these online answers. The security evaluation model is used to evaluate online answers, obtaining the security prediction value and security uncertainty of the online answers.

[0122] In the embodiment of the present application, the usability evaluation model can be implemented through a reward model, and the security evaluation model can be implemented through a cost model.

[0123] like Figure 8 As shown, in the first online answer generation module, the security risk index value is calculated based on the security prediction value and security uncertainty obtained in the online answer evaluation module. The comprehensive uncertainty is obtained based on the availability uncertainty and security uncertainty obtained in the online answer evaluation module.

[0124] For example, k online answers Input the reward model separately to get the online answer Availability prediction value and availability uncertainty . Put k online answers Input the cost model separately to get the online answer Safety prediction value and safety uncertainty .

[0125] Safety prediction value and safety uncertainty Enter the security risk index value calculation formula Get online answers security risk index value. is the cumulative distribution function of the standard normal distribution.

[0126] The availability uncertainty and safety uncertainty Enter the comprehensive uncertainty calculation formula In the above equation, we get the comprehensive uncertainty of the online answer. is a weighting factor used to balance availability and security.

[0127] Optionally, in the first online answer generation module, an online answer whose security risk index value is lower than the risk tolerance threshold is selected, and among the qualified online answers, an online answer with the largest comprehensive uncertainty index value is selected as the first online answer.

[0128] like Figure 8 As shown, in the dataset user evaluation module, the user's evaluation results for the first online answer and the offline answer are obtained, and a training dataset is established based on the evaluation results. In this example, the first online answer and the offline answer are simultaneously output to the user. The user evaluates the first online answer and the offline answer based on their preferences, selecting the answer they consider to be of higher quality and marking the other answer as the second-best answer. The user then evaluates the security of the higher-quality answer and the second-best answer, respectively. This results in a training data record. This training data record includes the online question, the higher-quality answer, the second-best answer, and the security evaluation value of the higher-quality answer and the security evaluation value of the second-best answer.

[0129] like Figure 8As shown, the model fine-tuning module 820 primarily includes a first training module, a training evaluation module, a loss value determination module, and a multi-objective optimization module. In the first training module, online questions in the training dataset output by the dataset expansion module 810 are input into the large language model to obtain training output answers. In the training evaluation module, the training output answers are evaluated for usability and security using the evaluation model. The evaluation results are input into the loss value determination module to obtain a loss function for the training output answers. In the multi-objective optimization module, the large language model is updated using the loss function. In the training evaluation module, the training output answers are evaluated using the usability evaluation model to obtain a usability prediction value for the training output answers. The training output answers are evaluated using the security evaluation model to obtain a security prediction value for the training output answers. In the loss value determination module, the usability prediction value of the training output answers is used to calculate a first advantage function. The first advantage function is used to quantify the relative usability advantage of the training output answers over better answers. When calculating the first advantage function, it may also be necessary to calculate an availability cost evaluation value for the training output answers. The usability cost evaluation value and the usability prediction value are used to determine the advantage function.

[0130] For example, the advantage function can be calculated using the Generalized Advantage Estimation (GAE) method based on the availability cost estimate and the availability prediction value. GAE is an advantage estimation method used in reinforcement learning, designed to more efficiently compute policy gradients. By combining the advantages of Monte Carlo estimation and TD estimation, GAE provides a low-variance and low-bias advantage estimation method, thereby accelerating the policy optimization process.

[0131] This embodiment of the present application utilizes the first advantage function to determine the availability loss function. For example, the availability loss can be calculated using the same method used in the Proximal Policy Optimization (PPO) algorithm. Proximal Policy Optimization (PPO), an actor-critic-based reinforcement learning algorithm, is used to train policy models to complete complex tasks.

[0132] Similarly, in the loss value determination module, the security prediction value of the training output answer can be used to calculate a second advantage function. This second advantage function is used to quantify the relative security advantage of the training output answer over a superior answer. When calculating the second advantage function, it may also be necessary to calculate the security cost assessment value of the training output answer. The advantage function is determined using the security cost assessment value and the security prediction value. The GAE method can be used to calculate the advantage function based on the security cost assessment value and the security prediction value. The second advantage function is used to determine the security loss function. The security loss can be calculated using the same method used in the PPO algorithm.

[0133] Optionally, in the multi-objective optimization module, the availability loss function and the security loss function are multi-objectively optimized to obtain a loss function that balances availability and security. The resulting loss function is used to fine-tune the large language model.

[0134] For example, the Lagrange multiplier can be introduced , construct the Lagrangian function:

[0135]

[0136] in, is the availability loss function, is the security loss function. The final loss function is obtained by performing multi-objective optimization on these two loss functions.

[0137] You can alternate between fine-tuning the target model and updating the Lagrange multiplier. For example: , based on the Lagrangian function Update the target model; or update , can be achieved through Update, where is the update rate.

[0138] See also Figure 8The data flywheel fine-tuning device also includes an evaluation model fine-tuning module 830. In the evaluation model fine-tuning module 830, the evaluation model is fine-tuned using the training dataset. For example, a loss function is established for the usability evaluation model. The usability prediction values ​​of the superior answer and the suboptimal answer in the training dataset are input into the loss function of the usability evaluation model to obtain a loss value. This loss value reflects the difference between the usability evaluation model's usability prediction results for the answer and the user's preference for the answer. The parameters of the usability evaluation model are optimized by minimizing the loss value. In this example, a loss function can also be established for the security evaluation model. The security prediction values ​​of the superior answer, the suboptimal answer, the security evaluation value of the superior answer, and the security evaluation value of the suboptimal answer in the training dataset are input into the loss function of the security evaluation model to obtain a loss value. This loss value reflects the difference between the security evaluation model's security prediction results for the answer and the user's security evaluation of the answer. The parameters of the security evaluation model are optimized by minimizing the loss value.

[0139] For example, the Bradley-Terry model can be used to model the evaluation model. The Bradley-Terry model is a classic probability model used to model the preference relationship between paired data. It assigns a positive real number score to each option to represent its relative preference strength and calculates the relative preference probability between two options. For example, for option and , preference Exceeding preference Probability ,in and Respectively options and When training the model, the parameters are optimized using the negative log-likelihood loss function.

[0140] The loss function of the usability evaluation model is:

[0141]

[0142] in

[0143]

[0144] Loss function of the security assessment model:

[0145] .

[0146] In summary, we use the loss function to fine-tune the security and usability evaluation models, correcting the errors in the evaluation model's predictions. The fine-tuned evaluation model is then used to fine-tune the target model, ultimately resulting in a model with improved security and usability.

[0147] In the embodiments of the present application, security metrics and security constraints are introduced into the data collection, analysis, application, and feedback processes of the data flywheel. At the same time, in the face of the distribution drift phenomenon of the security and usability evaluation models, uncertainty assessment is introduced for both, and this guides the exploration and fine-tuning of the data flywheel. Given an initial large model and a certain number of evaluators, the evaluators ask questions to the large model and score the output of the large model. After a certain number of interactions and multiple fine-tuning of the large model, a large model with improved security and usability is finally obtained.

[0148] In some embodiments, Figure 8 The dataset expansion module 810 and model fine-tuning module 820 shown in FIG can both be implemented via software or hardware. For example, the implementation of the dataset expansion module 810 will be described below using the dataset expansion module 810 as an example. Similarly, the implementation of the model fine-tuning module 820 can refer to the implementation of the dataset expansion module 810.

[0149] As an example of a software functional unit, the dataset expansion module 810 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the dataset expansion module 810 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.

[0150] Similarly, the multiple hosts / virtual machines / containers used to run the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.

[0151] As an example of a hardware functional unit, dataset expansion module 810 may include at least one computing device, such as a server. Alternatively, dataset expansion module 810 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0152] The multiple computing devices included in the dataset expansion module 810 can be distributed in the same region or in different regions. The multiple computing devices included in the dataset expansion module 810 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the dataset expansion module 810 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.

[0153] It should be noted that, in other embodiments, the dataset expansion module 810 can be used to execute any step in the data flywheel fine-tuning method described in the above embodiments, and the model fine-tuning module 820 can also be used to execute any step in the data flywheel fine-tuning method described in the above embodiments. In addition, the steps that the dataset expansion module 810 and the model fine-tuning module 820 are responsible for implementing can also be specified as needed, and the dataset expansion module 810 and the model fine-tuning module 820 can respectively implement different steps in the data flywheel fine-tuning method described in the above embodiments to achieve Figure 8 The full functionality of the data flywheel fine-tuning device 800 is shown.

[0154] The present application also provides a computing device 900. Figure 9 As shown, computing device 900 includes a bus 902, a processor 904, a memory 906, and a communication interface 908. Processor 904, memory 906, and communication interface 908 communicate with each other via bus 902. Computing device 900 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 900.

[0155] The bus 902 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 9 The bus 904 may include a path for transmitting information between various components of the computing device 900 (eg, memory 906, processor 904, communication interface 908).

[0156] The processor 904 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0157] The memory 906 may include volatile memory, such as random access memory (RAM). The processor 904 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0158] The memory 906 stores executable program codes, and the processor 904 executes the executable program codes to respectively implement the aforementioned Figure 8 The functions of the dataset expansion module 810 and the model fine-tuning module 820 shown in FIG are implemented to realize the data flywheel fine-tuning method described in the above embodiment. That is, the memory 906 stores instructions for executing the data flywheel fine-tuning method described in the above embodiment.

[0159] Alternatively, the memory 906 stores executable codes, and the processor 904 executes the executable codes to respectively implement the aforementioned Figure 8 The data flywheel fine-tuning device 800 shown in FIG. 8 is used to implement the data flywheel fine-tuning method described in the above embodiment. That is, the memory 906 stores instructions for executing the data flywheel fine-tuning method described in the above embodiment.

[0160] The communication interface 908 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 900 and other devices or a communication network.

[0161] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0162] like Figure 10 As shown, the computing device cluster includes at least one computing device 900. The memory 906 in one or more computing devices 900 in the computing device cluster may store the same instructions for executing the data flywheel fine-tuning method described in the above embodiment.

[0163] In some possible implementations, the memory 906 of one or more computing devices 900 in the computing device cluster may also store partial instructions for executing the data flywheel fine-tuning method described in the above embodiments. In other words, the combination of one or more computing devices 900 can jointly execute instructions for executing the data flywheel fine-tuning method described in the above embodiments.

[0164] It should be noted that the memory 906 in different computing devices 900 in the computing device cluster can store different instructions, which are used to execute the above Figure 8 The data flywheel fine-tuning device 800 shown in FIG. 8 shows partial functions. That is, the instructions stored in the memory 906 of different computing devices 900 can implement the functions of one or more modules in the dataset expansion module 810 and the model fine-tuning module 820.

[0165] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network, which may be a wide area network or a local area network. Figure 11 A possible implementation is shown. Figure 11 As shown, two computing devices 900A and 900B are connected via a network. Specifically, the connection to the network is achieved through a communication interface within each computing device. In this possible implementation, the memory 906 within computing device 900A stores instructions for executing the functions of the dataset expansion module 810. Simultaneously, the memory 906 within computing device 900B stores instructions for executing the functions of the model fine-tuning module 820.

[0166] It should be understood that Figure 11The functionality of the computing device 900A shown in FIG. 1 may also be implemented by multiple computing devices 900. Similarly, the functionality of the computing device 900B may also be implemented by multiple computing devices 900.

[0167] The embodiment of the present application also provides another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similarly referred to as Figure 10 and Figure 11 The connection mode of the computing device cluster is different in that the memory 906 in one or more computing devices 900 in the computing device cluster may store the same instructions for executing the method in the above embodiment.

[0168] In some possible implementations, the memory 906 of one or more computing devices 900 in the computing device cluster may also store partial instructions for executing the aforementioned data flywheel fine-tuning method. In other words, the combination of one or more computing devices 900 can jointly execute instructions for executing the aforementioned data flywheel fine-tuning method.

[0169] It should be understood that each step of the above method embodiment can be completed by a hardware-based logic circuit or a software-based instruction in a processor.

[0170] Based on the methods in the above embodiments, embodiments of the present application provide a computer-readable storage medium comprising computer program instructions. When the computer program instructions are executed by a computing device cluster comprising at least one computing device, the computing device cluster executes the methods in the above embodiments. Exemplarily, the computer-readable storage medium can be any available medium capable of storing data on a computing device, or a data storage device such as a data center comprising one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives).

[0171] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product containing instructions. When the instructions are executed by a computing device cluster including at least one computing device, the computing device cluster executes the method in the above embodiment.

[0172] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0173] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and the storage medium can be located in an ASIC.

[0174] In the above embodiments, all or part of the embodiments can be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic media (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or semiconductor media (e.g., solid-state drive (SSD)).

[0175] It will be understood that the various numerical numbers involved in the embodiments of the present application are merely distinctions for the convenience of description and are not intended to limit the scope of the embodiments of the present application.

[0176] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the embodiments of the present application.

Claims

1. A data flywheel fine-tuning method, characterized in that: The method comprises: Based on the online question and the target model, an online answer and an offline answer are obtained; wherein the online answer is the answer generated by the target model to the online question in the online retrieval mode, and the offline answer is the answer generated by the target model to the online question in the offline retrieval mode; Evaluating the security of the online answer and determining a first online answer based on the evaluation result; Constructing a training dataset using the first online answer and the offline answer; Fine-tuning the target model using the training dataset; Determining the first online answer based on the evaluation result specifically includes: Based on the evaluation result, a security risk index value and a comprehensive uncertainty of the online answer are obtained; wherein the security risk index value indicates the probability that the answer has a security risk, and the comprehensive uncertainty indicates a comprehensive quantitative value of the answer's usability uncertainty and security uncertainty; the security uncertainty is a measure of the inaccuracy of the online answer's security prediction value, and the security prediction value is a quantitative evaluation indicator of the online answer's security risk; the usability uncertainty is a measure of the inaccuracy of the online answer's usability prediction value, and the usability prediction value is a quantitative evaluation indicator of whether the online answer satisfies the user's online question; The first online answer is determined according to the security risk index value and the comprehensive uncertainty.

2. The method according to claim 1, characterized in that The training dataset is constructed by using the first online answer and the offline answer: Obtaining evaluation results of the first online answer and the offline answer, wherein the evaluation results include a quality evaluation value and a safety evaluation value of the first online answer, and a quality evaluation value and a safety evaluation value of the offline answer; Based on the evaluation result, the training data set is established, wherein the training data set indicates the comparison result of the quality evaluation values ​​of the first online answer and the offline answer, and the training data set includes the online question, the first online answer, the offline answer, the safety evaluation value of the first online answer, and the safety evaluation value of the offline answer.

3. The method according to claim 1, characterized in that The security assessment of the online answer specifically includes: The security of the online answer is evaluated using an evaluation model to obtain a security prediction value and a security uncertainty of the online answer; wherein the evaluation model includes a security evaluation model for evaluating the security of the answer, and the security uncertainty indicates the security uncertainty of the answer.

4. The method according to claim 3, characterized in that The evaluation model further includes a usability evaluation model, and the usability evaluation model is used to evaluate the usability of the answer; The method further comprises: The online answer is evaluated using the usability evaluation model to obtain a usability prediction value and usability uncertainty of the online answer.

5. The method according to claim 4, characterized in that The method further comprises: Obtaining the safety risk index value according to the safety prediction value and the safety uncertainty; The comprehensive uncertainty is obtained according to the availability uncertainty and the safety uncertainty.

6. The method according to claim 1, characterized in that Fine-tuning the target model using the training dataset specifically includes: Inputting the online questions in the training data set into the target model to obtain training output answers; Evaluating the security of the training output answer to obtain an evaluation result of the training output answer, wherein the evaluation result of the training output answer includes a usability prediction value and a security prediction value; Determining a loss function using the evaluation results of the training output answers and the training data set; The target model is iteratively trained according to the loss function, where the loss function includes a usability loss function and a security loss function.

7. The method according to claim 6, characterized in that The step of evaluating the security of the training output answer to obtain the evaluation result of the training output answer specifically includes: The security of the training output answer is evaluated using an evaluation model to obtain a security prediction value of the training output answer, wherein the evaluation model includes a security evaluation model.

8. The method according to claim 7, characterized in that The evaluation model also includes a usability evaluation model; The method further comprises: The usability of the training output answer is evaluated using the usability evaluation model to obtain a predicted usability value of the training output answer.

9. The method according to claim 3 or 7, characterized in that The method further includes: fine-tuning the evaluation model using the training dataset.

10. A data flywheel fine-tuning device, characterized in that: The device comprises: A dataset expansion module is configured to obtain online answers and offline answers based on the online question and the target model; wherein the online answer is the answer generated by the target model to the online question in an online retrieval mode, and the offline answer is the answer generated by the target model to the online question in an offline retrieval mode; the security of the online answers is evaluated, and based on the evaluation results, a first online answer is determined, and a training dataset is constructed using the first online answer and the offline answer; A model fine-tuning module, configured to fine-tune the target model using the training dataset; The data set expansion module is specifically configured to obtain a security risk index value and a comprehensive uncertainty of the online answer based on the evaluation result; wherein the security risk index value indicates the probability that the answer has a security risk, and the comprehensive uncertainty indicates a comprehensive quantitative value of the answer's usability uncertainty and security uncertainty; the security uncertainty is a measure of the inaccuracy of the online answer's security prediction value, the security prediction value is a quantitative evaluation indicator of the online answer's security risk, and the usability uncertainty is a measure of the inaccuracy of the online answer's usability prediction value, the usability prediction value is a quantitative evaluation indicator of whether the online answer satisfies the user's online question; The first online answer is determined according to the security risk index value and the comprehensive uncertainty.

11. A cloud management platform, characterized in that: include: At least one computing node, on which the target model is deployed, and the cloud management platform is used to implement the method for fine-tuning the data flywheel as described in any one of claims 1-9.

12. A computing device comprising a memory and a processor, characterized in that: The memory stores instructions, and when the instructions are executed by the processor, the data flywheel fine-tuning method according to any one of claims 1 to 9 is implemented.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the data flywheel fine-tuning method according to any one of claims 1 to 9 is implemented.

14. A computer program product, characterized in that The computer program product includes program instructions, and when the program instructions are executed by a computer, the computer executes the data flywheel fine-tuning method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Model fine tuning training method and device, answer output method and device and electronic equipment

    CN119150013A

  • Context answer generation method and device, computer equipment and storage medium

    CN119322838A