Data flywheel fine tuning method and device

By introducing a security evaluation and screening mechanism into the data flywheel, the training data set is built using the answers of the target model in offline and networked modes, and the model is fine-tuned, solving the security and credibility problems of the output content of large language models during the life cycle, and achieving efficient security screening and model fine-tuning.

CN120234408AActive Publication Date: 2025-07-01SHENZHEN HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510705943.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-01
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

When large language models are learned in Internet data, they may inherit biased information, misinformation or harmful information, resulting in the output of discriminatory and negative remarks. Due to the "black box" nature of the model, it is difficult to trace and repair, affecting the security and credibility of the output content of the model during its life cycle.

Method used

By introducing safety assessment in the data production, model training, data screening and enhancement and feedback closed loop of the data flywheel, we obtain the answers of the target model in offline and networked modes, conduct safety assessments, filter out answers that meet safety requirements, and use these answers to build a training data set to fine-tune the model.

Benefits of technology

The security screening of the output content of the model is realized to ensure that the output of the model meets security requirements and avoids the generation of harmful content. At the same time, it improves the security and credibility of the output content of the model during its life cycle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234408A_ABST
    Figure CN120234408A_ABST
Patent Text Reader

Abstract

The invention provides a data flywheel fine tuning method and device, and relates to the technical field of artificial intelligence. The data flywheel fine tuning method comprises the following steps: based on an online question and a target model, obtaining an online answer and an offline answer; wherein the online answers are answers generated by the target model for the online questions in the networking retrieval mode, and the offline answers are answers generated by the target model for the online questions in the offline retrieval mode. And evaluating the security of the online answer, and determining a first online answer according to an evaluation result. And constructing a training data set through the first online answer and the offline answer. And performing fine tuning on the target model by using the training data set. According to the method, safety evaluation is introduced into data production, model training, data screening and enhancement and feedback closed loop of the data flywheel, and the content output by the model is guided to meet safety and availability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of artificial intelligence (AI), and in particular, to a method and device for fine-tuning a data flywheel. Background Art

[0002] Large language models learn from Internet data and may inherit biased information, false information, or harmful information in Internet data, resulting in harmful answers such as discriminatory and negative remarks in the output. However, as a "black box" model, the internal parameter interactions of large language models are complex, and the same input may also produce different answers, making it difficult to trace and repair harmful answers, thus affecting the security and credibility of the content output during the life cycle of large language models. Summary of the Invention

[0003] This application provides a method and device for fine-tuning a data flywheel. Safety evaluation is introduced in the data production, model training, data screening and enhancement, and feedback loop of the data flywheel to guide the content output by the model to meet safety and usability requirements.

[0004] In a first aspect, a method for fine-tuning a data flywheel is provided. In this method, the target model is deployed to an online service application. A user inputs a question to the online service application. The target model generates an offline answer to the online question in an offline retrieval mode and generates multiple online answers in an online retrieval mode. The content of the multiple online answers is subjected to safety evaluation. According to the evaluation results, a first online answer is determined from the multiple online answers. A training data set is constructed through the first online answer and the offline answer, and the target model is fine-tuned using the training data set.

[0005] In this way, by obtaining the answers output by the target model in the offline mode and the online mode, online data is collected in the data production link of the data flywheel, rather than just obtaining the answers output in the offline mode. By performing safety evaluation on the content of the online answers, safety screening of the content of the online answers is realized, and incremental data on the Internet is used to fine-tune the model, ensuring that the output content of the model meets safety requirements and avoiding the generation of harmful content.

[0006] In a possible implementation, the method further includes: by obtaining the quality evaluation values and security evaluation values of the first online answer and the offline answer, determining the order of the first online answer and the offline answer in the training dataset according to the quality evaluation values of the first online answer and the offline answer, where the training dataset includes online questions, the first online answer, the offline answer, the security evaluation value of the first online answer, and the security evaluation value of the offline answer. In this way, by using the user to evaluate the quality of the online answer and the offline answer, the user's preference is obtained. Thus, in the training dataset, the user's preference for the answer and the security evaluation of the answer are integrated, and the generated training dataset indicates the user's content preference and security requirements. Using this training data to fine-tune the model can improve the content output by the model to meet the user's preference and the user's security requirements.

[0007] In a possible implementation, the method further includes: evaluating the security of multiple online answers to obtain the security risk index values and comprehensive uncertainty of each online answer, and determining the first online answer according to the security risk index values and comprehensive uncertainty. In this way, by evaluating the security risk of the content of multiple online answers, the probability values of each online answer having security risks are obtained, and by introducing the security uncertainty and availability uncertainty evaluations, the distribution drift phenomenon between security and availability that is likely to occur in the process of evaluating the security of online answers can be solved, guiding the data flywheel to balance the availability and security of the content output by the model during the operation stage.

[0008] In a possible implementation, the method further includes: using an evaluation model to evaluate the security of the online answer to obtain the security prediction value and security uncertainty of the online answer. The evaluation model includes a security evaluation model for evaluating the answer in terms of security, and the security uncertainty indicates the security uncertainty of the answer.

[0009] In a possible implementation, the evaluation model further includes an availability evaluation model for evaluating the availability of the answer. The method further includes: using the availability evaluation model to evaluate the online answer to obtain the availability prediction value and availability uncertainty of the online answer. In this way, by using the security evaluation model and the availability evaluation model to evaluate the online answer respectively, through the security uncertainty and availability uncertainty evaluations in the evaluation results, the deviation of the security prediction result and the availability prediction result that may occur in different stages or different data distributions of the evaluation model can be solved.

[0010] In a possible implementation, the method further includes: obtaining a safety risk index value according to the safety prediction value and the safety uncertainty; obtaining a comprehensive uncertainty according to the availability uncertainty and the safety uncertainty. In this way, the safety risk index value is obtained by using the safety prediction value and the safety uncertainty which represents the degree of doubt about the safety conclusion. Compared with predicting the safety risk only by using the safety prediction value, the safety risk of the online answer can be obtained more accurately.

[0011] In a possible implementation, the method further includes: inputting the online questions in the training data set into the target model to obtain the training output answers. By evaluating the safety of the training output answers, the availability prediction value and the safety prediction value of the training output answers are obtained. The loss function is determined by using the evaluation results of the training output answers and the training data set. The target model is iteratively trained according to the loss function, and the loss function includes the loss function of availability and the loss function of safety. In this way, the loss function of availability and the loss function of safety determined by using the training data set and the safety prediction value and the availability prediction value of the training output answers guide the model output content to improve availability under the safety constraint in the iterative process of the model through multi-objective optimization of availability and safety.

[0012] In a possible implementation, the method further includes: using the evaluation model to evaluate the safety of the training output answers to obtain the safety prediction value of the training output answers, and the evaluation model includes a safety evaluation model.

[0013] In a possible implementation, the evaluation model further includes an availability evaluation model, and the method further includes: using the availability evaluation model to evaluate the availability of the training output answers to obtain the availability prediction value of the training output answers. In this way, in the model fine-tuning stage, the availability evaluation model and the safety evaluation model are used to guide the model to output content beneficial to availability under the condition of safety constraints.

[0014] In a possible implementation, the method further includes: using the training data set to fine-tune the evaluation model. In this way, the evaluation model is also fine-tuned by the training data set, and the results obtained by using the fine-tuned evaluation model to evaluate the content output by the target model are more accurate.

[0015] In a second aspect, a data flywheel fine-tuning device is provided, including a dataset expansion module and a model fine-tuning module. Among them, the dataset expansion module is used to obtain online answers and offline answers based on online questions and a target model. Among them, the online answer is the answer generated by the target model for the online question in the online retrieval mode, and the offline answer is the answer generated by the target model for the online question in the offline retrieval mode. The security of the online answer is evaluated, and according to the evaluation result, the first online answer is determined, and a training dataset is constructed through the first online answer and the offline answer. The model fine-tuning module is used to fine-tune the target model using the training dataset.

[0016] In a possible implementation manner, the dataset expansion module specifically obtains the evaluation results of the first online answer and the offline answer. Among them, the evaluation results include the quality evaluation value and the security evaluation value of the first online answer and the offline answer. According to the evaluation results, a training dataset is established. Among them, the comparison result of the quality evaluation value indicating the first online answer and the offline answer in the training dataset, and the training dataset includes the online question, the first online answer, the offline answer, the security evaluation value of the first online answer, and the security evaluation value of the offline answer.

[0017] In a possible implementation manner, the dataset expansion module is specifically used to obtain the security risk index value and the comprehensive uncertainty of the online answer according to the evaluation result. Among them, the security risk index value indicates the probability that the answer has a security risk, and the comprehensive uncertainty indicates the comprehensive quantification value of the availability uncertainty and the security uncertainty of the answer. According to the security risk index value and the comprehensive uncertainty, the first online answer is determined.

[0018] In a possible implementation manner, the dataset expansion module is specifically used to evaluate the security of the online answer using an evaluation model to obtain the security prediction value and the security uncertainty of the online answer; among them, the evaluation model includes a security evaluation model for evaluating the answer in terms of security, and the security uncertainty indicates the security uncertainty of the answer.

[0019] In a possible implementation manner, the evaluation model further includes an availability evaluation model, and the availability evaluation model is used to evaluate the availability of the answer. The dataset expansion module is specifically used to evaluate the online answer using the availability evaluation model to obtain the availability prediction value and the availability uncertainty of the online answer.

[0020] In a possible implementation manner, the dataset expansion module is specifically used to obtain the security risk index value according to the security prediction value and the security uncertainty. The comprehensive uncertainty is obtained according to the availability uncertainty and the security uncertainty.

[0021] In a possible implementation, the model fine-tuning module is specifically configured to input the online questions in the training dataset into the target model to obtain the training output answers. Evaluate the security of the training output answers to obtain the evaluation results of the training output answers, and the evaluation results of the training output answers include the availability prediction value and the security prediction value. Determine the loss function by using the evaluation results of the training output answers and the training dataset. Iteratively train the target model according to the loss function, and the loss function includes the loss function of availability and the loss function of security.

[0022] In a possible implementation, the model fine-tuning module is specifically configured to use the evaluation model to evaluate the security of the training output answers to obtain the security prediction value of the training output answers, and the evaluation model includes a security evaluation model.

[0023] In a possible implementation, the model fine-tuning module is specifically configured to use the availability evaluation model to evaluate the availability of the training output answers to obtain the availability prediction value of the training output answers.

[0024] In a possible implementation, the model fine-tuning module is specifically configured to fine-tune the evaluation model by using the training dataset.

[0025] In a third aspect, the present application provides a computing device cluster, including at least one computing device, and each computing device includes a processor and a memory; the processor of at least one computing device is configured to execute the instructions stored in the memory of at least one computing device, so that the computing device cluster executes the method described in the first aspect or any possible implementation of the first aspect.

[0026] In a fourth aspect, the present application provides a computer-readable storage medium, including computer program instructions, when the computer program instructions are executed by the computing device cluster, the computing device cluster executes the method described in the first aspect or any possible implementation of the first aspect. Exemplarily, the computing device cluster may include one or more computing devices.

[0027] In a fifth aspect, the present application provides a computer program product including instructions, when the instructions are run by the computing device cluster, the computing device cluster is caused to execute the method described in the first aspect or any possible implementation of the first aspect. Exemplarily, the computing device cluster may include one or more computing devices.

[0028] For the beneficial effects of the second aspect to the fifth aspect, reference may be made to the introduction of the beneficial effects of the first aspect above, and details are not described herein again. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is a schematic diagram of the technical concept of a data flywheel fine-tuning method provided by an embodiment of the present application;

[0030] Figure 2 It is another schematic diagram of the technical concept of a data flywheel fine-tuning method provided by an embodiment of the present application;

[0031] Figure 3 It is a schematic diagram of the deployment of the data flywheel fine-tuning service provided by an embodiment of the present application on a cloud computing platform;

[0032] Figure 4 It is a schematic flowchart of a data flywheel fine-tuning method provided by an embodiment of the present application;

[0033] Figure 5 It is a schematic flowchart of a data flywheel fine-tuning method provided by an embodiment of the present application;

[0034] Figure 6 It is a schematic diagram of input data in the data flywheel fine-tuning method provided by an embodiment of the present application;

[0035] Figure 7 It is a schematic diagram of the specific implementation manner of the data flywheel fine-tuning method provided by an embodiment of the present application;

[0036] Figure 8 It is a schematic diagram of the composition of the data flywheel fine-tuning device provided by an embodiment of the present application;

[0037] Figure 9 It is a schematic diagram of the structure of a computing device provided by an embodiment of the present application;

[0038] Figure 10 It is a schematic diagram of the structure of a computing device cluster provided by an embodiment of the present application;

[0039] Figure 11 It is a schematic diagram of the structure of another computing device cluster provided by an embodiment of the present application. Detailed implementation manner

[0040] Next, the solutions provided by the embodiments of the present application will be described in conjunction with the accompanying drawings. Among them, in the embodiments of the present application, "a plurality of" means two or more. "First", "second", etc. are only used to distinguish similar objects and do not have to be used to describe a specific order or the number of objects.

[0041] To facilitate understanding of the solutions provided by the embodiments of the present application, the technical terms that may be involved in the embodiments of the present application will be introduced first.

[0042] A large language model (LLM) is a computer model capable of processing and generating natural language. LLM is an artificial intelligence technology based on deep learning that can predict the next word or sentence by learning the statistical patterns and semantic information of language data. LLM usually consists of multiple neural networks (such as feedforward neural networks and recurrent neural networks). Among them, the feedforward neural network is used for feature extraction of the input text, and the recurrent neural network is used for modeling text sequences. By continuously training on a large amount of text data, LLM can gradually improve its understanding and prediction ability of natural language.

[0043] Reinforcement Learning from Human Feedback (RLHF) is a technology that matches large language models with human preferences. A reward model is trained based on human preference data to model human tendencies towards model responses, and the language model is optimized through reinforcement learning algorithms to maximize the cumulative reward, thereby aligning the output of the language model with human preferences.

[0044] Data Flywheel refers to the dynamic interaction and iterative optimization between data and models during the training and application of large models, forming a mutually reinforcing closed-loop system. The data flywheel includes the following links: data production, model training, data screening and enhancement, and feedback closed-loop.

[0045] To ensure the security and credibility of the output content of large language models throughout their life cycle. In existing technical solutions, one approach is to use an offline fine-tuning method based on safe reinforcement learning, which optimizes the large model by combining a pre-collected human feedback dataset. The core of this method is to collect human feedback data on offline data, train a reward model using the human feedback data, and then fine-tune the large model. Here, the offline fine-tuning method based on safe reinforcement learning can effectively address the security risks during the fine-tuning process of large models through human feedback reward modeling and safety constraint optimization, without generating real-world risks. However, when applying this large language model to the data flywheel and using the Internet to generate incremental data for fine-tuning the large language model, it may output content that does not guarantee safety to users, presenting a relatively high real-world risk.

[0046] Another solution is the data flywheel method that does not explicitly focus on security. This method accumulates data on the Internet. After evaluating the data through task metrics, it mixes multi-source data. It uses the multi-source data to fine-tune the large language model. This method can continuously increase the scale of the training data and use the updated training data to fine-tune the large language model. This method can obtain more high-quality and more novel data, which helps to improve the capabilities of the large model. However, in the process of data accumulation, this method only evaluates the data through task metrics, resulting in the mixing of harmful content data in the data fusion stage. Subsequently, the large language model is fine-tuned using the data doped with harmful content. If strict requirements are imposed on the security of the output content during the fine-tuning process, high-quality novel data cannot be output, consuming manpower and material resources. If the novelty of the output content is to be improved, the security of the output content cannot be guaranteed. Ultimately, the security of the data output by the fine-tuned model cannot be guaranteed.

[0047] The above methods have explored ways to ensure the security of the output content of the large language model. However, none of these methods have accumulated safe and effective online data, nor have they measured security during the fine-tuning stage, and they cannot balance the security and novelty of the output content.

[0048] In view of this, the embodiments of this application provide a data flywheel fine-tuning method. In the embodiments of this application, security measurement and security constraints are introduced in all links of the data flywheel. Facing the security and novelty of the output content, by introducing uncertainty evaluation, and using this to guide the data flywheel to explore and fine-tune in terms of balancing the security and usability of the output content, a large language model with improved security and usability is finally obtained.

[0049] Exemplarily, Figure 1 shows a schematic diagram of the technical concept of a data flywheel fine-tuning method. As Figure 1 shown, the architecture for the data flywheel fine-tuning method mainly includes: an online data acquisition part, an online answer evaluation part, a training dataset generation part, a training evaluation part, and a model fine-tuning part. Among them, the online data acquisition part is mainly used to acquire the input online questions and the answers output by the target model. The online questions are the input instructions or questions used to guide the model to generate specific types of answers. The answers output by the target model include multiple online answers generated in the online exploration mode and offline answers generated in the offline mode. Exemplarily, in this part, the online questions and answers can be acquired through the online data acquisition module 110.

[0050] The online answer evaluation part mainly conducts security evaluation and usability evaluation on the online answers to obtain the evaluation results. Using the evaluation results, the first online answer is selected, which is the answer that best balances security and usability among all online answers. In some embodiments, the evaluation results may include the security prediction value and usability prediction value of each online answer. At the same time, it may also include the security uncertainty and usability uncertainty of each online answer. The security prediction value is a quantitative evaluation index for the security risk of the content. The usability prediction value is a quantitative evaluation index for whether the content meets usability. The security uncertainty is a measure of the inaccuracy of the security prediction value of the content. The usability uncertainty is a measure of the inaccuracy of the usability prediction value of the content. Exemplarily, in this part, the online answer evaluation module 120 can evaluate n online answers to obtain the online answer that best balances security and usability. In some embodiments, the n online answers can also be input into the evaluation model to obtain the evaluation results of the online answers. The evaluation model may include a security evaluation model and a usability evaluation model. The usability evaluation model can use a reward model. The reward model (Reward Model, RM) is the policy optimization process of RLHF under the guidance of rewards. The reward model takes {user question, model output} as input and outputs a scalar score to evaluate the quality of the replies generated by the policy model. The security evaluation model can be implemented through a cost model. The cost model (Reward Model, CM) takes {user question, model output} as input and outputs a scalar score to evaluate the security of the replies generated by the model.

[0051] The training dataset generation part is mainly used to obtain the evaluation results of the first online answer and the offline answer, so as to obtain the evaluation results of the security and usability of the first online answer and the offline answer from the user's perspective, and form a training dataset based on the user's evaluation. Exemplarily, the training dataset generation module 130 can obtain the answer with better quality selected by the user from the first online answer and the offline answer, and the other answer is recorded as the sub-optimal answer. Then, the security of the answer with better quality and the sub-optimal answer are evaluated respectively. Thus, a training data record is obtained. This training data record includes the online question, the better answer, the sub-optimal answer, the security evaluation value of the better answer, and the security evaluation value of the sub-optimal answer. In this way, in the production link of the training dataset in the data flywheel, by obtaining the online data generated in the networked mode, these online data are novel. The security and usability of the online answers are evaluated, and the online answers with security and usability and novelty are selected. Using this online answer as the training dataset instead of just using offline data ensures the novelty, security, and usability of the output content. At the same time, the training dataset also includes the user's preferences and the evaluation results of the user on the security of the output content.

[0052] The training and evaluation part mainly inputs the online questions in the training dataset into the target model to obtain the output answers of the target model, and evaluates the usability and security of the output answers. Exemplarily, the training and evaluation module 140 can randomly select an online question from the training dataset and input the online question into the target model to obtain the answer output by the target model. Due to the complex interaction of internal parameters in the large language model, the output answer may be different from both the previous online answers and offline answers. The evaluation model is used to evaluate the output answer to obtain the security prediction value and usability prediction value of the output answer.

[0053] The model fine-tuning part mainly uses the security prediction value and usability prediction value of the output answer to obtain the loss function of the output answer, so as to take into account the dual optimization goals of usability and security, and uses the loss function to fine-tune the target model. Exemplarily, the target model can be fine-tuned by the model fine-tuning module 150.

[0054] Based on the above scheme framework, by evaluating the security and usability of the online-generated content in the dataset generation link, training the model with a training dataset that takes into account both security and usability, and evaluating the security and usability of the generated content during training, the security and usability are also taken into account during the model fine-tuning process. The security monitoring and evaluation of the data flywheel in all links are realized, ensuring that the output content of the model always meets the security requirements and avoiding the generation of harmful content. In addition, in response to the need to balance security and usability, through the evaluation of security uncertainty and usability uncertainty, the data flywheel is guided to balance security and usability, and a model aligned with effectiveness and security is obtained.

[0055] Under the above scheme framework, in order to improve the evaluation accuracy of the security and usability of the output content, a fine-tuning part of the evaluation model can also be added to the above scheme framework. At this time, the above scheme framework can be transformed into Figure 2 the scheme framework shown. Figure 2 Fig. shows another schematic diagram of the technical concept of a data flywheel fine-tuning method. In Figure 2 it, the architecture for the data flywheel fine-tuning method mainly includes: an online data acquisition part, an online answer evaluation part, a training dataset generation part, a training and evaluation part, a model fine-tuning part, and an evaluation model fine-tuning part. Among them, Figure 2 the online data acquisition part, the online answer evaluation part, the training dataset generation part, the training and evaluation part, and the model fine-tuning part in it can refer to the relevant descriptions in the foregoing Figure 1 and will not be described here.

[0056] In Figure 2Among them, the fine-tuning part of the evaluation model is mainly used to obtain the fine-tuning of the evaluation model using the training dataset. In some embodiments, a loss function of the evaluation module can be established. Exemplarily, a loss function can be established for the usability evaluation model through the evaluation model fine-tuning module 160, and the usability prediction values of the better answers and the sub-optimal answers in the training dataset are input into the loss function of the security evaluation model to obtain a loss value. This loss value reflects the difference between the usability prediction result of the usability evaluation model for the answers and the user's preference for the answers. The parameters of the usability evaluation model are optimized by minimizing the loss value. A loss function is established for the security evaluation model, and the security prediction values of the better answers, the security prediction values of the sub-optimal answers, the security evaluation values of the better answers, and the security evaluation values of the sub-optimal answers in the training dataset are input into the loss function of the security evaluation model to obtain a loss value. This loss value reflects the difference between the security prediction result of the security evaluation model for the answers and the user's security evaluation of the answers. The parameters of the security evaluation model are optimized by minimizing the loss value. Finally, the fine-tuned evaluation model can be updated to the evaluation model in the online answer evaluation part. In addition, the fine-tuned evaluation model can also be updated to the evaluation model used in the training evaluation part.

[0057] Based on Figure 2 The scheme framework shown, through the fine-tuning of the evaluation model, corrects the wrong results predicted by the evaluation model. Using the fine-tuned evaluation model for the training data collection link and the link for fine-tuning the target model, improves the evaluation accuracy of the security and usability of the output content, and finally obtains a model with improved security and usability.

[0058] In some embodiments, each part in the above-described scheme framework can be configured on a cloud computing platform. For example, it can be deployed on at least one virtual machine or container instance, so that the cloud computing platform can provide data flywheel fine-tuning services for the model. Of course, each part in the above-described scheme framework can also be configured on nodes other than the cloud computing platform. For example, it can be deployed in at least one data center or on at least one server, which can be determined according to the actual situation and is not limited here. Among them, the cloud computing platform can provide a page related to the public cloud service for users to remotely access the public cloud service. In this embodiment, the user can purchase data flywheel fine-tuning services on the cloud computing platform in advance. For ease of understanding, the interaction form between the user and the cloud computing platform is described below. As Figure 3As shown in the figure, the interaction between the user and the cloud computing platform mainly includes: the user logs in to the cloud computing platform 300 through the client web page, selects and purchases the data flywheel fine-tuning service in the cloud computing platform 300. After the purchase, the user can collect the training dataset on the cloud computing platform 300 based on the function provided by the data flywheel fine-tuning service, which can ensure both the security and novelty of the output content. Among them, the cloud computing platform 300 is mainly used to manage the infrastructure for running the data flywheel fine-tuning service. Exemplarily, the infrastructure for running the data flywheel fine-tuning service may include multiple data centers set in different regions, and each data center includes multiple servers. The data center can provide basic resources for the data flywheel fine-tuning service, such as computing resources, storage resources, etc. Therefore, when the user purchases and uses the data flywheel fine-tuning service, the user mainly pays for the resources used. When the user uses the data flywheel fine-tuning service, the user can issue the fine-tuning frequency and number of times, etc. through the configuration interface, application programming interface or interface interacting with the user provided by the cloud computing platform 300. After that, the cloud computing platform 300 can fine-tune the large language model according to the instructions issued by the user. In some embodiments, each part in the above-described solution framework can also be configured on the local server, which can be determined according to the actual situation and is not limited here.

[0059] The specific implementation process of the above technical concept is described below.

[0060] Please refer to Figure 4 , Figure 4 , which shows a schematic flowchart of a data flywheel fine-tuning method. It can be understood that this method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities. Exemplarily, this method can be executed by a data flywheel fine-tuning device, where the device can be implemented by software and / or hardware, and can be but not limited to being configured in an electronic device or a server. Typically, it can be configured on a server. As Figure 4 shown, the data flywheel fine-tuning method may include the following steps:

[0061] S401. Obtain online questions and online answers.

[0062] In the embodiment of the present application, the online questions can be input by the user or sent by a device, apparatus, or service, etc., which is not limited here.

[0063] The target model can generate an offline answer in the offline mode or generate multiple online answers in the online mode.

[0064] Exemplarily, a probability p can be set. When the question input this time falls within the probability p, the target model generates k online answers in the online mode . In this way, the running performance of the target model deployed in the Internet environment can be not affected.

[0065] S402. Perform a security assessment on the online answer and output the security assessment result.

[0066] In the embodiments of the present application, a security matching rule and security risk keywords can be established through natural language processing methods. The online answer is matched through the security matching rule and security risk keywords to obtain the security assessment result of the online answer. Similarly, the online answer can also be matched through available keywords and availability matching rules to obtain the availability assessment result. For example, the online answer can be evaluated through the evaluation model described in the aforementioned online answer evaluation section to obtain the availability and security assessment results of the online answer.

[0067] S403. Determine the first online answer according to the security assessment result of the online answer.

[0068] In the embodiments of the present application, after obtaining the security assessment results of each online answer, the online answers that meet the first online answer can be screened out through a preset standard. For example, when the assessment result is represented by an assessment score, the online answers with a security assessment score greater than or equal to the security threshold can be selected, and the online answers with an availability assessment score greater than or equal to the availability threshold are selected from these online answers as the first online answer.

[0069] S404. Construct a training data set through the first online answer and the offline answer.

[0070] In the embodiments of the present application, a training data record may include an online question, the first online answer, the offline answer, and the availability prediction values and security prediction values of the first online answer and the offline answer. Exemplarily, obtain the evaluation results of the user on the first online answer and the offline answer, and establish a training data set according to the evaluation results. In this example, the first online answer and the offline answer are output to the user at the same time. The user evaluates the first online answer and the offline answer according to their own preferences, selects the answer that the user thinks has better quality, and the other answer is recorded as the sub-optimal answer. The user evaluates the security of the better-quality answer and the sub-optimal answer respectively. Thus, a training data record is obtained. This training data record includes the online question, the better-quality answer, the sub-optimal answer, the user-side security evaluation value of the better-quality answer, and the user-side security evaluation value of the sub-optimal answer. In the embodiments of the present application, the online answer and the offline answer can also be input into a large language model, and the large language model evaluates the quality and security of the answer.

[0071] S405. Fine-tune the target model using the training data set.

[0072] In the embodiments of the present application, the online questions in the training dataset are input into the target model to obtain the training output answers. The training output answers are evaluated for usability and security to obtain the evaluation results. According to the evaluation results, a loss function is constructed with the goal of the output content meeting usability and security, and the target model is updated using this loss function. Exemplarily, the training output answers can be evaluated by the method described in step 402. By the security and usability prediction values of the training output answers and the security and usability prediction values of the better answers in the training dataset, the loss function is set, and the parameters in the target model are adjusted through the loss function.

[0073] In this way, by screening the dataset for security and usability in the training dataset generation link, training the model using the training dataset that takes into account both security and usability, evaluating the generated content for security and usability during training, and also taking into account security and usability during the model fine-tuning process. The full-link security monitoring and evaluation of the data flywheel are realized, ensuring that the output content of the model always meets the security requirements and avoiding the generation of harmful content.

[0074] In some embodiments, reference may also be made to Figure 5 . Figure 5 shows a schematic flow diagram of a data flywheel fine-tuning method. As Figure 5 shown, in the dataset generation link 501, the online questions are input into the target model to obtain the offline answers generated in the offline mode and multiple online answers generated in the online mode. The n online answers are respectively input into the security evaluation model and the usability evaluation model. Each online answer is evaluated by the security evaluation model to obtain the security prediction value and the security uncertainty of each online answer. Each online answer is evaluated by the usability evaluation model to obtain the usability prediction value and the usability uncertainty of each online answer. The security prediction value is a quantitative evaluation index for the security risk of the content. The security uncertainty is a measure of the inaccuracy of the security prediction value of the content. The usability prediction value is a quantitative evaluation index for whether the content meets the user's online questions. The usability uncertainty is a measure of the inaccuracy of the usability prediction value of the content.

[0075] In some embodiments, the usability evaluation model can be implemented by a reward model. The reward model (RewardModel, RM) is a policy optimization process under the guidance of rewards in RLHF. The reward model takes {user question, model output} as input and outputs a scalar score to evaluate the quality of the replies generated by the policy model. In the embodiments of the present application, {online question, online answer} is used as the input of the usability evaluation model, and the usability prediction value of the online answer and the usability uncertainty of the online answer The usability uncertainty is an indicator that can measure the credibility or reliability of online answers, reflecting the range in which online answers may deviate from the true value in terms of usability. The security evaluation model can be implemented through a cost model. The cost model (Reward Model, CM) takes {user question, model output} as input and outputs a scalar score for evaluating the security of the replies generated by the model. In the embodiments of the present application, {online question, online answer} is used as the input of the security evaluation model, and the security prediction value of the online answer is output and the security uncertainty of the online answer The security uncertainty is an indicator that can measure the security level or security reliability of online answers, reflecting the range in which online answers may deviate from the true value in terms of security.

[0076] Exemplarily, the usability evaluation model is used to evaluate the online answer, and the usability prediction value and usability uncertainty of the online answer are obtained. The k online answers are respectively input into the reward model, and the usability prediction value of the online answer and the usability uncertainty are obtained.

[0077] In some embodiments, the security risk indicator value can also be obtained according to the security prediction value and security uncertainty. The comprehensive uncertainty is obtained according to the usability uncertainty and security uncertainty of the online answer. Exemplarily, the security prediction value and the security uncertainty are input into the security risk indicator value calculation formula to obtain the security risk indicator value of the online answer . is the cumulative distribution function of the standard normal distribution. The usability uncertainty and the security uncertainty are input into the comprehensive uncertainty calculation formula to obtain the comprehensive uncertainty of the online answer. Among them, is the weighting coefficient for balancing usability and security.

[0078] For example Figure 5As shown, in the dataset generation step 501, online answers with security risk indicator values lower than the risk tolerance threshold are selected, and among the eligible online answers, the online answer with the largest comprehensive uncertainty indicator value is selected as the first online answer. The first online answer and the offline answer are presented to the user, and based on the user's quality evaluation of the first online answer and the offline answer, the better answer and the second-best answer among the first online answer and the offline answer are determined. The user's security evaluation value for the better answer and the user's security evaluation value for the second-best answer are obtained. (Online question, better answer, second-best answer, user-side security evaluation value of the better answer, user-side security evaluation value of the second-best answer) is used as a training data record.

[0079] Exemplarily, the offline answer and the first online answer are presented to the user, and the user selects the one with higher quality according to their own preference and records it as , and the other is . Then, the user respectively evaluates whether the two results are secure, denoted as and . The above data is integrated to obtain a five-tuple ;

[0080] In addition, all the collected training data records can be appended to the training dataset, and the initial dataset is also added to the training dataset, where the initial dataset is a dataset generated offline.

[0081] In the model training step 502, as Figure 5 shown, the online questions in the training dataset are input into the target model to obtain the training output answers. The evaluation model is used to evaluate the usability and security of the training output answers to obtain the evaluation results. According to the evaluation results, a loss function of the target model whose output content meets the usability and security is constructed. The target model is updated using this loss function.

[0082] In the embodiments of the present application, the usability evaluation model is used to evaluate the training output answers to obtain the usability prediction value of the training output answers. The security evaluation model is used to evaluate the training output answers to obtain the security prediction value of the training output answers. According to the usability prediction value and the security prediction value, an advantage function is obtained. The advantage function can be used to determine the quantitative index value of the relative advantage of the training output answers over the better answers in terms of usability and security. The loss function is determined using the advantage function. The target model is fine-tuned with a dual optimization objective for the security loss and the usability loss through the loss function, and finally a target model that takes into account both security and usability is obtained.

[0083] Exemplarily, a cost evaluation model can be used to evaluate the usability cost of the training output answer to obtain a usability cost evaluation value. Using the usability cost evaluation value and the usability prediction value of the training output answer, a first advantage function is obtained. The first advantage function can be used to obtain a quantitative index value of the relative advantage of the training output answer over a better answer in terms of usability. According to the first advantage function, a usability loss function is determined. A cost evaluation model is used to evaluate the security cost of the training output answer to obtain a security cost evaluation value. Using the security cost evaluation value and the security prediction value of the training output answer, a second advantage function is calculated. The second advantage function can be used to obtain a quantitative index value of the relative advantage of the training output answer over a better answer in the training dataset in terms of security. Using the second advantage function, a security loss function is obtained. Using the security loss function and the usability loss function, a total loss function that takes into account both security and usability is obtained. The target model is fine-tuned using the total loss function, so as to obtain a target model that takes into account the usability and security of the output content. The fine-tuned target model can be updated to the online service. The cost evaluation model here can be used to quantify the model prediction error. Specifically, it can obtain the prediction error of the usability prediction value of the content and the prediction error of the security prediction value of the content.

[0084] For example, online questions are extracted from the training dataset, and the online questions are input into the target model to obtain an output answer. The output answer is evaluated by a reward model and a cost model to obtain a usability prediction value and a security prediction value.

[0085] Using the Generalized Advantage Estimation (GAE) method, a safety advantage function is obtained based on the usability prediction value and the usability prediction estimate. Using the security prediction value and the security prediction estimate, a security advantage function is obtained. 。

[0086] The usability loss is calculated using the way of calculating the loss in the Proximal Policy Optimization (PPO) algorithm. And the security loss 。Among them,

[0087]

[0088] ,

[0089] Among them, is the expectation of the current policy distribution, is the policy update ratio, that is, the probability ratio of the new policy to the old policy. is the parameter of the policy network. is the clipping threshold, used to control the update amplitude. is the clipping function. It's an online problem. is the training dataset, It is the training output answer obtained by inputting the online answer into the target model.

[0090] Introducing Lagrange multipliers , construct the Lagrangian function to realize the total loss function of multiple optimization objectives for availability and security .in,

[0091] In the process of optimizing the target model using the total loss function, the total loss function can be used to , update the target model. Or update the Lagrange multiplier , ,in is the update rate.

[0092] By evaluating the security and availability of the model output during training and establishing a loss function for dual-objective optimization of availability and security, we have achieved a balance between security and availability during model fine-tuning, and ultimately implemented security monitoring and evaluation of the data flywheel at all stages.

[0093] In some embodiments, an evaluation model training step 503 is also included. Figure 5 As shown, in the evaluation model training link 503, the evaluation model is fine-tuned using the training data set. For example, a loss function is established for the usability evaluation model, and the usability prediction value of the better answer and the usability prediction value of the suboptimal answer in the training data set are input into the loss function of the first evaluation model to obtain a loss value. The loss value reflects the difference between the usability prediction result of the usability evaluation model for the answer and the user's preference for the answer. The parameters of the usability evaluation model are optimized by minimizing the loss value. In this example, a loss function is established for the security evaluation model, and the security prediction value of the better answer, the security prediction value of the suboptimal answer, the security evaluation value of the better answer and the security evaluation value of the suboptimal answer in the training data set are input into the loss function of the security evaluation model to obtain a loss value. The loss value reflects the difference between the security prediction result of the security evaluation model for the answer and the user's security evaluation of the answer. The parameters of the security evaluation model are optimized by minimizing the loss value.

[0094] Exemplarily, the Bradley-Terry model can be used to model the evaluation model. The Bradley-Terry model is a classical probability model used to model the preference relationship between paired data. It assigns a positive real number score to each option to represent its relative preference strength and calculates the relative preference probability between two options. For example, for options and , the probability that the preference exceeds is:

[0095] where and are the scores of options and respectively. When training this model, the parameters are optimized through the negative log-likelihood loss function:

[0096] The loss function of the usability evaluation model is:

[0097]

[0098] where,

[0099]

[0100] The loss function of the security evaluation model:

[0101]

[0102] where, is the predicted probability, is a linear combination of the input values.

[0103] In summary, the loss function is used to fine-tune the security and usability evaluation models to correct the error degree of the evaluation model prediction. The target model is fine-tuned using the fine-tuned evaluation model, and finally a model with improved security and usability is obtained.

[0104] The above is the related introduction to the data flywheel fine-tuning method provided in this application. In some embodiments, this method can be combined with blockchain technology. The generation and use process of training data is recorded through the blockchain to ensure the authenticity and immutability of the data. In the smart contract of the blockchain, an automatic update mechanism for the reward model (RM) and the cost model (CM) can be implemented. For example, when new data is added to the dataset, the smart contract can trigger the fine-tuning process of the model to ensure that the model is always in the optimal state.

[0105] Figure 6 is the schematic diagram of the input data in the data flywheel fine-tuning method in the embodiments of this application. As Figure 6As shown, the online questions or online answers in the data flywheel fine-tuning device in the embodiments of the present application can also be multimodal data, such as audio or video data. In the embodiments of the present application, a lightweight adapter may be used to fine-tune the target model instead of directly training the entire target model. Federated learning technology can also be used to allow multiple clients, such as different enterprises or user groups, to jointly train the model without sharing data. In this way, while protecting data privacy, the generalization ability of the model can be improved.

[0106] It can be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. In addition, the various embodiments described above and the technical features in each embodiment can be combined according to actual situations, and the combined solutions are still within the protection scope of the present application.

[0107] Next, based on the method in the above embodiments, the specific implementation manners of the data flywheel fine-tuning method provided in the embodiments of the present application will be introduced. Figure 7 It is a schematic diagram of the specific implementation manner of the data flywheel fine-tuning method provided in the embodiments of the present application. As Figure 7 shown, the online question is: "How to protect the security of my social media account". The large language model outputs the following online answers:

[0108] Online answer 1: There is no need to protect at all. Anyway, the accounts of ordinary people are not worth protecting.

[0109] Online answer 2: Bind the mobile phone or authenticate the APP, and start two-factor authentication.

[0110] Online answer 3: Do not click on private message links from unknown sources.

[0111] Online answer 4: Rely on the platform, and the platform will handle everything for you.

[0112] Use the evaluation model to evaluate the online answers. Obtain the evaluation results of the online answers. The evaluation results are as follows:

[0113] Online answer 1: Security prediction value 0, usability prediction value 0, security uncertainty 0.6, usability uncertainty 0.8; Online answer 2: Security prediction value 1, usability prediction value 2, security uncertainty 0.8, usability uncertainty 0.8; Online answer 3: Security prediction value 1, usability prediction value 1, security uncertainty 0.8, usability uncertainty 0.6; Online answer 4: Security prediction value 0, usability prediction value 0, security uncertainty 0.5, usability uncertainty 0.5.

[0114] Among the online answers 2 and 3 that meet the security constraints, the online answer 2 with the highest availability evaluation score is determined as the first online answer. The online answer 2 and the offline answer are displayed to the user, and the user conducts quality evaluation and security evaluation on the online answer 2 and the offline answer. Among them, the offline answer: It is recommended to implement a strong password policy, set the password length ≥ 12 bits, mix uppercase and lowercase letters, numbers, and symbols, and prohibit reusing passwords. According to the user's evaluation results, a training data record is constructed.

[0115] For the online question, the user has a greater preference for the online answer 2. The online answer 2 is the better answer, and the offline answer is the second-best answer. Then, the user gives a security score to the online answer 2 and a security score to the offline answer. A training data record (How to prevent my social media account from being hacked, online answer 2, offline answer, 1, 1) is obtained.

[0116] In the model training session, the online question "How to prevent my social media account from being hacked" in the training dataset is input into the large language model, and an answer is output. The answer is input into the evaluation model to obtain the security prediction value and availability prediction value of the answer. Using the security prediction value and availability prediction value of the output answer, the availability loss function and security loss function are determined. The availability loss function and security loss function are subjected to multi-optimization objectives to obtain the loss function that balances availability and security. The large language model is fine-tuned using the final loss function.

[0117] The above is an example of the data flywheel fine-tuning method in the embodiments of the present application in practical applications. Based on the method in the above embodiments, a data flywheel fine-tuning device provided in the embodiments of the present application is introduced. Figure 8 The composition schematic diagram of a data flywheel fine-tuning device is shown. As Figure 8 shown, the data flywheel fine-tuning device 800 includes a dataset expansion module 810 and a model fine-tuning module 820. The dataset expansion module 810 mainly obtains the input online questions, multiple answers generated in the online exploration mode of the target model, and the offline answers generated in the offline mode. The security and availability of the multiple online answers are evaluated to obtain the evaluation results of the online answers. According to the evaluation results, the online answers that take into account both security and availability are determined. The training dataset is expanded based on the online answers and offline answers, and the expanded dataset is output. The model fine-tuning module 820 fine-tunes the target model using the expanded dataset.

[0118] In the embodiments of the present application, the target model may be a policy model. The policy model is the large language model to be optimized in RLHF. Taking the user's question as the input, it outputs the model's response. The ultimate goal of RLHF is to optimize the policy model so that it outputs high-quality text and aligns with human preferences. Among them, the target model can be any neural network model that can be deployed to an online service application. The embodiments of the present application do not limit the specific implementation manner of the online service application function. Exemplarily, the target model can be deployed to an online consultation application to answer online questions raised by users.

[0119] As Figure 8 shown, the dataset expansion module 810 mainly includes: an online answer evaluation module, a first online answer generation module, and a dataset user evaluation module. In the online answer evaluation module, an evaluation model is used to evaluate the usability and security of multiple online answers. In the first online answer generation module, based on the evaluation results of the usability and security of multiple online answers, a first online answer is determined. This online answer balances usability and security and is the best among all online answers in terms of the quantified index values of usability and security. The dataset user evaluation module expands the training dataset based on the evaluation results of users on offline answers and the first online answer.

[0120] In the embodiments of the present application, the evaluation model can be any neural network model that can evaluate the usability and security of text. The embodiments of the present application do not limit the specific implementation manner of the usability and security evaluation. Exemplarily, the evaluation model can be a reward model, a cost model, etc., which can quantify the usability and security of the answers output by the large language model.

[0121] In the embodiments of the present application, the evaluation model includes a security evaluation model and a usability evaluation model. In the online answer evaluation module, the usability evaluation model is used to evaluate the online answers to obtain the usability prediction values and usability uncertainties of these online answers. The security evaluation model is used to evaluate the online answers to obtain the security prediction values and security uncertainties of the online answers.

[0122] In the embodiments of the present application, the usability evaluation model can be implemented through a reward model. The security evaluation model can be implemented through a cost model.

[0123] As Figure 8 shown, in the first online answer generation module, according to the security prediction values and security uncertainties obtained in the online answer evaluation module, a security risk index value is calculated. According to the usability uncertainties and security uncertainties obtained in the online answer evaluation module, a comprehensive uncertainty is obtained.

[0124] Exemplarily, k online answers Input them into the reward model respectively to obtain the online answers of the availability prediction value and the availability uncertainty . Input the k online answers into the cost model respectively to obtain the online answers of the security prediction value and the security uncertainty .

[0125] Input the security prediction value and the security uncertainty into the security risk index value calculation formula to obtain the online answers of the security risk index value. is the cumulative distribution function of the standard normal distribution.

[0126] Input the availability uncertainty and the security uncertainty into the comprehensive uncertainty calculation formula to obtain the comprehensive uncertainty of the online answers. Among them, is the weighting coefficient used to balance availability and security.

[0127] Optionally, in the first online answer generation module, select the online answers with the security risk index value lower than the risk tolerance threshold, and among the qualified online answers, select the online answer with the largest comprehensive uncertainty index value as the first online answer.

[0128] As Figure 8 shown, in the dataset user evaluation module, obtain the evaluation results of the user on the first online answer and the offline answer, and establish a training dataset according to the evaluation results. In this example, the first online answer and the offline answer are output to the user at the same time. The user evaluates the first online answer and the offline answer according to their own preferences, selects the answer with better quality considered by the user, and the other answer is recorded as the sub-optimal answer. The user evaluates the security of the answer with better quality and the sub-optimal answer respectively. Thus, a training data record is obtained. This training data record includes the online question, the better answer, the sub-optimal answer, the security evaluation value of the better answer, and the security evaluation value of the sub-optimal answer.

[0129] As Figure 8As shown in the figure, the model fine-tuning module 820 mainly includes: a first training module, a training evaluation module, a loss value determination module, and a multi-objective optimization module. In the first training module, the online questions in the training dataset output by the dataset expansion module 810 are input into the large language model to obtain training output answers. In the training evaluation module, the evaluation model is used to evaluate the usability and security of the training output answers, and the obtained evaluation results are input into the loss value determination module to obtain the loss function of the training output answers. In the multi-objective optimization module, the large language model is updated using the loss function. In the training evaluation module, the usability evaluation model is used to evaluate the training output answers to obtain the usability prediction value of the training output answers. The security evaluation model is used to evaluate the training output answers to obtain the security prediction value of the training output answers. In the loss value determination module, the first advantage function can be calculated using the usability prediction value of the training output answers. The first advantage function is used to obtain the quantitative index value of the relative advantage of the training output answers over the better answers in terms of usability. When calculating the first advantage function, it may also be necessary to calculate the usability cost evaluation value of the training output answers. The advantage function is determined using the usability cost evaluation value and the usability prediction value.

[0130] Exemplarily, the Generalized Advantage Estimation (GAE) method can be used to calculate the advantage function based on the usability cost evaluation value and the usability prediction value. GAE is an advantage estimation method used in reinforcement learning, aiming to calculate the policy gradient more efficiently. By combining the advantages of Monte Carlo estimation and TD estimation, GAE provides a low-variance and low-bias advantage estimation method, thus accelerating the policy optimization process.

[0131] In the embodiment of this application, the first advantage function is used to determine the usability loss function. Exemplarily, the way of calculating the loss in the PPO algorithm can be used to calculate the usability loss. The Proximal Policy Optimization (PPO) algorithm, an Actor-Critic based reinforcement learning algorithm, is used to train the policy model to complete complex tasks.

[0132] Similarly, in the loss value determination module, the safety prediction value of the training output answer can be utilized to calculate the second advantage function, which is used to obtain a quantitative metric value representing the relative advantage of the training output answer over a better answer in terms of safety. When calculating the second advantage function, it may also be necessary to calculate the safety cost evaluation value of the training output answer. Based on the safety cost evaluation value and the safety prediction value, the advantage function is determined. The Generalized Advantage Estimation (GAE) method can be used to calculate the advantage function according to the safety cost evaluation value and the safety prediction value. Using the second advantage function, the safety loss function is determined. The way of calculating the loss in the Proximal Policy Optimization (PPO) algorithm can be used to calculate the safety loss.

[0133] Optionally, in the multi-objective optimization module, the usability loss function and the safety loss function are optimized for multi-objective to obtain a loss function that balances usability and safety. The large language model is fine-tuned using the final loss function.

[0134] Exemplarily, a Lagrange multiplier can be introduced , and a Lagrangian function is constructed:

[0135]

[0136] where is the usability loss function, and is the safety loss function. By performing multi-objective optimization on these two loss functions, the final loss function is obtained.

[0137] The fine-tuning of the target model and the update of the Lagrange multiplier can be carried out alternately. For example: fixing , the target model is updated based on the Lagrangian function ; or updating , which can be updated through , where is the update rate.

[0138] Please refer to Figure 8, the data flywheel fine-tuning device further includes an evaluation model fine-tuning module 830. In the evaluation model fine-tuning module 830, the evaluation model is fine-tuned using a training data set. For example, a loss function is established for the usability evaluation model, and the usability prediction values of the better answers and the sub-optimal answers in the training data set are input into the loss function of the usability evaluation model to obtain a loss value. This loss value reflects the difference between the usability prediction result of the usability evaluation model for the answers and the user's preference for the answers. The parameters of the usability evaluation model are optimized by minimizing the loss value. In this example, a loss function can also be established for the security evaluation model, and the security prediction values of the better answers, the security prediction values of the sub-optimal answers, the security evaluation values of the better answers, and the security evaluation values of the sub-optimal answers in the training data set are input into the loss function of the security evaluation model to obtain a loss value. This loss value reflects the difference between the security prediction result of the security evaluation model for the answers and the user's security evaluation of the answers. The parameters of the security evaluation model are optimized by minimizing the loss value.

[0139] Exemplarily, the Bradley-Terry model can be used to model the evaluation model. The Bradley-Terry model is a classical probability model for modeling the preference relationship between paired data. It assigns a positive real number score to each option to represent its relative preference strength and calculates the relative preference probability between two options. For example, for options and , the probability that preference exceeds preference is , where and are the scores of options and respectively. When training this model, the parameters are optimized through the negative log-likelihood loss function.

[0140] The loss function of the usability evaluation model is:

[0141]

[0142] where

[0143]

[0144] The loss function of the security evaluation model:

[0145] .

[0146] In summary, the loss function is used to fine-tune the security and usability evaluation model to correct the error degree predicted by the evaluation model. The fine-tuned evaluation model is used to fine-tune the target model, and finally a model with improved security and usability is obtained.

[0147] In the embodiments of the present application, security metrics and security constraints are introduced into the data collection, analysis, application, and feedback processes of the data flywheel. At the same time, in the face of the distribution drift phenomenon of the security and usability evaluation models, uncertainty evaluation is introduced for both, and this is used to guide the exploration and fine-tuning of the data flywheel. Given an initial large model and a certain number of evaluators, the evaluators ask questions to the large model and score the output of the large model. After a certain number of interactions and multiple fine-tuning of the large model, a large model with improved security and usability is finally obtained.

[0148] In some embodiments, Figure 8 Both the dataset expansion module 810 and the model fine-tuning module 820 shown in can be implemented by software or can be implemented by hardware. Exemplarily, next, taking the dataset expansion module 810 as an example, the implementation manner of the dataset expansion module 810 is introduced. Similarly, the implementation manner of the model fine-tuning module 820 can refer to the implementation manner of the dataset expansion module 810.

[0149] As an example of a software functional unit, the dataset expansion module 810 can include code running on a computing instance. Among them, the computing instance can include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance can be one or more. For example, the dataset expansion module 810 can include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running this code can be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers for running this code can be distributed in the same availability zone (AZ) or in different AZs, and each AZ includes one data center or multiple geographically close data centers. Among them, usually one region can include multiple AZs.

[0150] Similarly, the multiple hosts / virtual machines / containers for running this code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, usually one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is realized through the communication gateway.

[0151] As an example of a hardware functional unit, the dataset augmentation module 810 may include at least one computing device, such as a server, etc. Alternatively, the dataset augmentation module 810 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0152] The multiple computing devices included in the dataset augmentation module 810 may be distributed in the same region or in different regions. The multiple computing devices included in the dataset augmentation module 810 may be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the dataset augmentation module 810 may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, the multiple computing devices may be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0153] It should be noted that in other embodiments, the dataset augmentation module 810 may be used to execute any step in the data flywheel fine-tuning method described in the above embodiments, and the model fine-tuning module 820 may also be used to execute any step in the data flywheel fine-tuning method described in the above embodiments. In addition, the steps to be implemented by the dataset augmentation module 810 and the model fine-tuning module 820 may also be specified as needed, and different steps in the data flywheel fine-tuning method described in the above embodiments are respectively implemented through the dataset augmentation module 810 and the model fine-tuning module 820 Figure 8 to implement all the functions of the data flywheel fine-tuning device 800 shown.

[0154] This application also provides a computing device 900. As Figure 9 shown, the computing device 900 includes: a bus 902, a processor 904, a memory 906, and a communication interface 908. The processor 904, the memory 906, and the communication interface 908 communicate with each other through the bus 902. The computing device 900 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 900.

[0155] The bus 902 can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 9 only one line is used in Figure 9 , but it does not mean that there is only one bus or one type of bus. The bus 904 can include a path for transmitting information between various components of the computing device 900 (for example, the memory 906, the processor 904, and the communication interface 908).

[0156] The processor 904 can include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0157] The memory 906 can include volatile memory, such as random access memory (RAM). The processor 904 can also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0158] The executable program code is stored in the memory 906, and the processor 904 executes the executable program code to respectively implement the functions of the data set expansion module 810 and the model fine-tuning module 820 shown in the foregoing Figure 8 , so as to implement the data flywheel fine-tuning method described in the above embodiments. That is, the memory 906 stores instructions for executing the data flywheel fine-tuning method described in the above embodiments.

[0159] Alternatively, the executable code is stored in the memory 906, and the processor 904 executes the executable code to respectively implement the functions of the data flywheel fine-tuning device 800 shown in the foregoing Figure 8 , so as to implement the data flywheel fine-tuning method described in the above embodiments. That is, the memory 906 stores instructions for executing the data flywheel fine-tuning method described in the above embodiments.

[0160] The communication interface 908 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 900 and other devices or communication networks.

[0161] Embodiments of this application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0162] As Figure 10 shown, the computing device cluster includes at least one computing device 900. Instructions for executing the data flywheel fine-tuning method described in the foregoing embodiments can be stored in the memories 906 of one or more of the computing devices 900 in the computing device cluster.

[0163] In some possible implementation manners, partial instructions for executing the data flywheel fine-tuning method described in the foregoing embodiments can also be stored separately in the memories 906 of one or more of the computing devices 900 in the computing device cluster. In other words, a combination of one or more computing devices 900 can jointly execute the instructions for executing the data flywheel fine-tuning method described in the foregoing embodiments.

[0164] It should be noted that the memories 906 in different computing devices 900 in the computing device cluster can store different instructions, which are respectively used to execute partial functions of the data flywheel fine-tuning device 800 described above. That is, the instructions stored in the memories 906 of different computing devices 900 can implement the functions of one or more of the dataset expansion module 810 and the model fine-tuning module 820. Figure 8 shown

[0165] In some possible implementation manners, one or more computing devices in the computing device cluster can be connected through a network. Among them, the network can be a wide area network or a local area network, etc. Figure 11 shows a possible implementation manner. As Figure 11 shown, two computing devices 900A and 900B are connected through a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation manner, instructions for executing the function of the dataset expansion module 810 are stored in the memory 906 of the computing device 900A. At the same time, instructions for executing the function of the model fine-tuning module 820 are stored in the memory 906 of the computing device 900B.

[0166] It should be understood that Figure 11The functions of the computing device 900A shown can also be completed by multiple computing devices 900. Similarly, the functions of the computing device 900B can also be completed by multiple computing devices 900.

[0167] The embodiments of the present application also provide another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similarly referred to Figure 10 and Figure 11 the connection mode of the computing device cluster. The difference is that the same instructions for executing the method in the above embodiments can be stored in the memory 906 of one or more computing devices 900 in the computing device cluster.

[0168] In some possible implementation manners, partial instructions for executing the foregoing data flywheel fine-tuning method can also be stored separately in the memory 906 of one or more computing devices 900 in the computing device cluster. In other words, a combination of one or more computing devices 900 can jointly execute the instructions for executing the foregoing data flywheel fine-tuning method.

[0169] It should be understood that each step of the foregoing method embodiment can be completed by a logic circuit in the form of hardware in the processor or an instruction in the form of software.

[0170] Based on the method in the above embodiments, the embodiments of the present application provide a computer-readable storage medium, including computer program instructions. When the computer program instructions are executed by a computing device cluster including at least one computing device, the computing device cluster is caused to execute the method in the above embodiments. Exemplarily, the computer-readable storage medium can be any available medium that the computing device can store or a data storage device such as a data center including one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive), etc.

[0171] Based on the method in the above embodiments, the embodiments of the present application provide a computer program product including instructions. When the instructions are executed by a computing device cluster including at least one computing device, the computing device cluster is caused to execute the method in the above embodiments.

[0172] It can be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0173] The method steps in the embodiments of the present application may be implemented in a hardware manner or by a processor executing software instructions. The software instructions may be composed of corresponding software modules, and the software modules may be stored in a random access memory (RAM), flash memory, read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), registers, hard disks, removable hard disks, CD-ROMs, or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium may also be a component of the processor. The processor and the storage medium may be located in the ASIC.

[0174] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.

[0175] It can be understood that the various numerical numbers involved in the embodiments of the present application are only for the convenience of description and are not used to limit the scope of the embodiments of the present application.

[0176] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present application.

Claims

1. A data flywheel fine-tuning method, characterized in that, The method includes: Based on the online question and the target model, obtain an online answer and an offline answer; wherein, the online answer is the answer generated by the target model for the online question in the networked retrieval mode, and the offline answer is the answer generated by the target model for the online question in the offline retrieval mode; Evaluate the security of the online answer, and determine the first online answer according to the evaluation result; Construct a training data set through the first online answer and the offline answer; Use the training data set to fine-tune the target model.

2. The method according to claim 1, wherein The constructing the training data set through the first online answer and the offline answer: Obtain the evaluation results of the first online answer and the offline answer, wherein the evaluation results include the quality evaluation value and the security evaluation value of the first online answer, and the quality evaluation value and the security evaluation value of the offline answer; According to the evaluation results, establish the training data set, wherein the training data set indicates the comparison result of the quality evaluation values of the first online answer and the offline answer, and the training data set includes the online question, the first online answer, the offline answer, the security evaluation value of the first online answer, and the security evaluation value of the offline answer.

3. The method according to claim 1, characterized in that The determining the first online answer according to the evaluation result specifically includes: According to the evaluation result, obtain the security risk index value and the comprehensive uncertainty of the online answer; wherein, the security risk index value indicates the probability that the answer has a security risk, and the comprehensive uncertainty indicates the comprehensive quantification value of the availability uncertainty and the security uncertainty of the answer; Determine the first online answer according to the security risk index value and the comprehensive uncertainty.

4. The method according to claim 1, wherein The evaluating the security of the online answer specifically includes: Use an evaluation model to evaluate the security of the online answer, and obtain the security prediction value and the security uncertainty of the online answer; wherein, the evaluation model includes a security evaluation model for evaluating the security of the answer, and the security uncertainty indicates the security uncertainty of the answer.

5. The method according to claim 4, wherein The evaluation model further includes an availability evaluation model for evaluating the availability of the answer; The method further includes: Use the availability evaluation model to evaluate the online answer, and obtain the availability prediction value and the availability uncertainty of the online answer.

6. The method according to claim 5, wherein The method further includes: Obtain the security risk index value according to the security prediction value and the security uncertainty; Obtain the comprehensive uncertainty according to the availability uncertainty and the security uncertainty.

7. The method according to claim 1, wherein The using the training data set to fine-tune the target model specifically includes: Input the online question in the training data set into the target model to obtain a training output answer; Evaluate the security of the training output answer to obtain the evaluation result of the training output answer, and the evaluation result of the training output answer includes the availability prediction value and the security prediction value; Use the evaluation result of the training output answer and the training data set to determine the loss function; Iteratively train the target model according to the loss function, where the loss function includes a loss function for availability and a loss function for security.

8. The method according to claim 7, wherein The evaluation of the security of the training output answer to obtain the evaluation result of the training output answer specifically includes: Use an evaluation model to evaluate the security of the training output answer to obtain a security prediction value of the training output answer, where the evaluation model includes a security evaluation model.

9. The method according to claim 8, wherein The evaluation model further includes an availability evaluation model; The method further includes: Use the availability evaluation model to evaluate the availability of the training output answer to obtain an availability prediction value of the training output answer.

10. The method according to claim 4 or 8, characterized in that The method further includes: fine-tuning the evaluation model using the training data set.

11. A data flywheel fine-tuning device, characterized in that, The device includes: A data set expansion module for obtaining an online answer and an offline answer based on an online question and a target model; where the online answer is the answer generated by the target model for the online question in the online retrieval mode, and the offline answer is the answer generated by the target model for the online question in the offline retrieval mode, evaluate the security of the online answer, determine a first online answer according to the evaluation result, and construct a training data set through the first online answer and the offline answer; A model fine-tuning module for fine-tuning the target model using the training data set.

12. A cloud management platform, characterized in that, Includes: At least one computing node with the target model deployed thereon, and the cloud management platform is used to implement the data flywheel fine-tuning method according to any one of claims 1-10.

13. A computing device, comprising a memory and a processor, characterized in that, Instructions are stored in the memory, and when the instructions are executed by a processor, the data flywheel fine-tuning method according to any one of claims 1-10 is implemented.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the data flywheel fine-tuning method according to any one of claims 1-10 is implemented.

15. A computer program product, characterized in that, The computer program product includes program instructions, and when the program instructions are executed by a computer, the computer is caused to execute the data flywheel fine-tuning method according to any one of claims 1-10.

Citation Information

Patent Citations

  • Question and answer method and device, electronic equipment and storage medium

    CN116719918A

  • Model training processing method and device, information processing method and device, equipment and medium

    CN117290490A

  • Model fine tuning training method and device, answer output method and device and electronic equipment

    CN119150013A

  • Context answer generation method and device, computer equipment and storage medium

    CN119322838A

  • Computer-readable recording medium storing evaluation program, evaluation method, and evaluation apparatus

    US20250103960A1