Model training method and device, electronic equipment and storage medium

By performing multi-stage fine-tuning training of emotional, professional and daily dialogue training data on natural language models, the problem of insufficient understanding and expression ability of the model in complex emotional companion scenarios is solved, and natural, professional and safe conversational interaction is achieved.

CN120068983AInactive Publication Date: 2025-05-30HANGZHOU HUAXI INTELLIGENT TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510534726.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing natural language models cannot deeply understand the language expression habits of specific groups in complex emotional companion scenarios, cannot accurately grasp the deep emotional needs of the conversation, and lack flexibility and adaptability.

Method used

By performing the first instruction supervision fine-tuning training based on emotional training data on the natural language model, the first intermediate model is obtained; then performing the second instruction supervision fine-tuning training based on professional training data to obtain the second intermediate model; finally, direct preference optimization training is carried out based on daily dialogue training data to obtain the target model.

Benefits of technology

It realizes the natural, professional and safe nature of the natural language model in conversation interaction, and improves emotional expression ability, flexibility and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068983A_ABST
    Figure CN120068983A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method and device, electronic equipment and a storage medium, and the method comprises the steps: carrying out the first instruction supervision fine tuning training of a natural language model based on preset emotion training data, and obtaining a first intermediate model; performing second instruction supervision fine tuning training on the first intermediate model based on preset professional training data to obtain a second intermediate model; and based on preset daily dialogue training data, carrying out direct preference optimization training on the second intermediate model to obtain a target model. According to the invention, natural, professional and safe session interaction of the natural language model can be realized, and the emotional expression ability of the natural language model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular, to a model training method, apparatus, electronic device, and storage medium. Background Art

[0002] With the development of computer technology, natural language models have been increasingly widely used in the field of dialogue interaction, such as psychological counseling, medical Q&A, and emotional companionship. In the prior art, natural language models usually generate responses based on rule templates, or use a BERT (Bidirectional Encoder Representations from Transformers) - like model to perform semantic matching on conversations to generate corresponding responses, or generate responses in a corresponding style by guiding the natural language model based on a prompt template.

[0003] However, in some relatively complex emotional companionship scenarios, such as those for the elderly or blind, the natural language models in the prior art cannot deeply understand the complex scenarios and the language expression habits of specific groups, cannot accurately grasp the deep - level emotional needs of conversations, and lack flexibility and adaptability. Summary of the Invention

[0004] The present invention provides a model training method, apparatus, electronic device, and storage medium to achieve natural, professional, and secure conversation interaction of natural language models and improve the emotional expression ability of natural language models.

[0005] In a first aspect, an embodiment of the present invention provides a model training method, which includes:

[0006] Performing first - order supervised fine - tuning training on a natural language model based on pre - set emotional training data to obtain a first intermediate model;

[0007] Performing second - order supervised fine - tuning training on the first intermediate model based on pre - set professional training data to obtain a second intermediate model;

[0008] Performing direct preference optimization training on the second intermediate model based on pre - set daily conversation training data to obtain a target model.

[0009] In a second aspect, an embodiment of the present invention further provides a model training apparatus, which includes:

[0010] A first - order supervised fine - tuning training module, configured to perform first - order supervised fine - tuning training on a natural language model based on pre - set emotional training data to obtain a first intermediate model;

[0011] The second instruction supervised fine-tuning training module is used to perform second instruction supervised fine-tuning training on the first intermediate model based on pre-set professional training data to obtain a second intermediate model;

[0012] The direct preference optimization training module is used to perform direct preference optimization training on the second intermediate model based on pre-set daily conversation training data to obtain a target model.

[0013] In a third aspect, an embodiment of the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the model training method according to any one of the embodiments of the present invention.

[0014] In a fourth aspect, an embodiment of the present invention further provides a storage medium storing computer-executable instructions. When the computer-executable instructions are executed by a computer processor, they are used to execute the model training method according to any one of the embodiments of the present invention.

[0015] The technical solution of the embodiment of the present invention obtains a first intermediate model by performing first instruction supervised fine-tuning training on a natural language model based on sentiment training data, obtains a second intermediate model by performing second instruction supervised fine-tuning training on the first intermediate model based on professional training data, and performs direct preference optimization training on the second intermediate model based on daily conversation training data to obtain a final target model for session interaction. It solves the problem that the natural language model in the prior art cannot deeply understand the language expression habits of complex scenarios and specific groups during session interaction, cannot accurately grasp the deep emotional needs of the session, and lacks flexibility and adaptability. The technical solution of the present invention realizes natural, professional, and safe session interaction of the natural language model, and improves the emotional expression ability, flexibility, and adaptability of the natural language model.

[0016] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0018] Figure 1 It is a flowchart of a model training method provided in Embodiment 1 of the present invention;

[0019] Figure 2 It is a flowchart of a model training method provided in the second embodiment of the present invention;

[0020] Figure 3 It is a schematic structural diagram of a model training device provided in the third embodiment of the present invention;

[0021] Figure 4 It is a schematic structural diagram of an electronic device provided in the fourth embodiment of the present invention. Detailed implementation manners

[0022] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0023] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. In the embodiments of the present application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned, and they should be regarded as exemplary. The purpose is only to illustrate the feasibility in the implementation of the technical solution of the present application, but it does not mean that the applicant has already or necessarily used this solution.

[0024] In the technical solution of the present application, the acquisition, transmission, storage, use, processing, etc. of data all comply with the relevant regulations of national laws and regulations.

[0025] Embodiment 1

[0026] Figure 1FIG. 0 is a flowchart of a model training method provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of training a natural language model to improve the professionalism and emotional expression ability of the natural language model. This method can be executed by a model training device, which can be implemented in the form of hardware and / or software, and the model training device can be configured in a server.

[0027] As Figure 1 shown, the method includes:

[0028] S110. Based on pre-set emotion training data, perform first instruction supervised fine-tuning training on the natural language model to obtain a first intermediate model.

[0029] Among them, the emotion training data is used to guide the natural language model to perform emotion perception, and improve the emotional quotient and anthropomorphic style of the natural language model for dialogue responses. The emotion training data can be generated based on emotion companionship conversations, psychological counseling conversations, etc.

[0030] The natural language model is based on Natural Language Processing (NLP) technology, aiming to enable a computer to understand, process, and generate human natural language. In this embodiment, the type and quantity of the natural language model can be one or multiple. When the type and quantity of the natural language model are multiple, the operations of S110-S130 can be respectively performed on each natural language model to obtain a target model after fine-tuning and training of each natural language model. Further, the target models obtained after fine-tuning and training of each natural language model can be horizontally compared to select the optimal target model.

[0031] The "first" and "second" in the first instruction supervised fine-tuning training and the second instruction supervised fine-tuning training hereinafter are only used to distinguish different instruction supervised fine-tuning trainings, and do not represent order, etc. The instruction supervised fine-tuning training is a training method of performing supervised fine-tuning optimization on the natural language model by giving instruction data. The first intermediate model is an intermediate model obtained after performing the first instruction supervised fine-tuning training on the natural language model. Similarly, the "first" and "second" in the first intermediate model and the second intermediate model hereinafter are only used to distinguish intermediate models in different training stages, and are not used to represent order, sequence, etc.

[0032] Further, the first instruction supervised fine-tuning training can be instruction supervised fine-tuning training based on the low-rank adaptation technology, that is, instruction supervised fine-tuning training based on the LoRA (Low-Rank Adaptation) technology. In the instruction supervised fine-tuning training based on the LoRA technology, most of the parameters of the natural language model are frozen, and only a small number of trainable low-rank matrix parameters are adjusted to adapt to new tasks or data. In this way, effective fine-tuning of the natural language model can be achieved without significantly increasing the computational cost and storage requirements.

[0033] Specifically, in this embodiment, when performing instruction supervised fine-tuning training on the natural language model based on the LoRA technology, the rank value of the low-rank matrix can be set to 32, and the weight scaling factor can be configured to 16 to control the parameter scale while ensuring the model's expressive ability. This embodiment adopts a Dropout mechanism of 0.1 to prevent overfitting. In the treatment of the bias term, this embodiment chooses not to update any bias parameters and specifies the task type as an autoregressive language modeling task. In terms of the configuration of the optimizer and training parameters, this embodiment adopts the AdamW (Adaptive Moment Estimation with Weight Decay) optimizer, where the first moment of the momentum coefficient is set to 0.9 and the second moment is set to 0.999. Considering the stability of training, this embodiment adopts a base learning rate of 1e-5. In terms of batch processing, this embodiment sets a base batch size of 4 and expands the effective batch size through 16 steps of gradient accumulation. The entire training process is set to 13 epochs, and a gradient clipping threshold of 1.0 is set to prevent gradient explosion. This embodiment adopts a conservative strategy in weight decay and sets its coefficient to 0. In the learning rate scheduling strategy, this embodiment implements a combination of linear warm-up and cosine decay. Specifically, in the initial stage of training, a warm-up ratio of 0.03 is adopted, and the learning rate is steadily increased through the LinearLR (Linear Learning Rate Scheduler) scheduler, and then the learning rate is periodically decayed through the CosineAnnealingLR (Cosine Annealing Learning Rate) to ensure that the model can stably converge to a better solution.

[0034] In this embodiment, when fine-tuning and training the natural language model, a three-stage progressive training scheme is adopted. In the first stage, the sentiment training data is used as the training data, and the natural language model is trained for sentiment ability using the instruction supervised fine-tuning training based on the LoRA technology, so that the natural language model can possess humanized sentiment understanding and expression abilities and achieve anthropomorphic high-emotional intelligence alignment.

[0035] S120. Based on pre-set professional training data, perform second instruction supervised fine-tuning training on the first intermediate model to obtain a second intermediate model.

[0036] Among them, the professional training data is used to ensure the professionalism and security of the model output results, and the professional training data can be generated based on the professional knowledge of a specific field.

[0037] Furthermore, the second instruction supervised fine-tuning training in this embodiment can be instruction supervised fine-tuning training based on the Quantized Low-Rank Adaptation (QLoRA) technique, that is, instruction supervised fine-tuning training based on the QLoRA (Quantized Low-Rank Adaptation) technique.

[0038] In this embodiment, by adopting the instruction supervised fine-tuning training based on the QLoRA technique and combining quantization and adapters, it is possible to reduce the video memory occupancy while ensuring the model performance.

[0039] In this embodiment, in the first stage of model training, after training the natural language model to obtain the first intermediate model, use the professional training data as the training data, and adopt the instruction supervised fine-tuning training based on the QLoRA technique to perform knowledge infusion and safety alignment training on the first intermediate model to obtain the second intermediate model. Through the model training in the second stage of this embodiment, the safety of the output of the second intermediate model and the accuracy of professional knowledge are ensured.

[0040] S130. Based on pre-set daily conversation training data, perform direct preference optimization training on the second intermediate model to obtain a target model.

[0041] Among them, the daily conversation training data is used to enhance the inference ability and anthropomorphic expression effect of the model, and the daily conversation training data can be generated based on daily conversations and general knowledge in various fields.

[0042] The direct preference optimization training is also the DPO (Direct Preference Optimization) training. The DPO training focuses on aligning with human preference behaviors, while enhancing the general inference ability of the model to ensure that the target model has an output consistent with human preferences.

[0043] Specifically, when performing DPO training in this embodiment, a sigmoid-type DPO loss function can be adopted, and the temperature coefficient is set to 0.1 to precisely control the KL (Kullback-Leibler divergence) divergence constraint strength, while not enabling the label smoothing mechanism to maintain the purity of training. To improve the training efficiency, this embodiment adopts a 4-bit quantization training strategy, uses the float16 data type during the calculation process, and combines double quantization technology and nf4 quantization type, which can significantly reduce the video memory occupancy. In terms of the LoRA configuration of the actor model, the present invention increases the rank value of the low-rank matrix to 64 to enhance the model's adjustment ability, while maintaining a weight scaling factor of 16 and a Dropout ratio of 0.1. Considering the particularity of DPO training, the present invention adopts a more cautious training strategy, reduces the base learning rate to 4e-6, sets the single-card batch size to 1, and ensures the training stability through 16 steps of gradient accumulation. Only 1 round of training is performed in the entire DPO stage, and the same 0.03 learning rate warm-up ratio as in the first stage is maintained, and the combined scheduling strategy of linear warm-up and cosine decay is continued.

[0044] In this embodiment, after the first and second stages of model training, the model has been able to balance emotional understanding, expression ability, professionalism, and knowledge. In the third stage, the daily dialogue training data is used as the training data, and direct preference optimization training is adopted to obtain the target model, which can enable the target model to focus on aligning with human preference behaviors and enhance the general reasoning ability of the target model.

[0045] The technical solution of this embodiment adopts instruction supervised fine-tuning training based on the LoRA technology in the first stage, focusing on enabling the natural language model to learn humanized emotional understanding and expression ability and perform anthropomorphic high-emotional intelligence alignment. The second stage focuses on introducing professional training data for instruction supervised fine-tuning training based on the QLoRA technology to ensure the safety of the model output and the accuracy of professional knowledge. The third stage further enhances the model's reasoning ability and anthropomorphic expression effect through the DPO reinforcement learning method. Through this three-stage progressive model training, the balanced development of the natural language model in multiple dimensions such as professionalism, emotional expression, and safety is achieved. At the same time, through the training strategy settings and parameter configurations in each stage, the efficient utilization of training resources is realized while ensuring the model performance, and the stability of the training process and the reliability of the results are ensured.

[0046] Further, after obtaining the target model, it further includes:

[0047] S1. Determine the dialogue test data in at least two scenarios, where the dialogue test data includes single-round dialogue test data and multi-round dialogue test data;

[0048] S2. Input the conversation test data into the target model to obtain reply data that matches the conversation test data;

[0049] S3. Determine the evaluation result of the target model according to the reply data.

[0050] Among them, the conversation test data is used to test the conversation effect of the target model, and the scenario of the conversation test data can be flexibly set according to the applicable scenario of the target model. For example, when the target model is an emotional companionship model for the elderly group, the conversation test data can cover various emotional need scenarios of the elderly group such as living alone, being widowed, and intergenerational conflicts. Another example is that when the target model is an emotional companionship model for the disabled group, the conversation test data can cover various emotional need scenarios such as rehabilitation, medical treatment, and mental health.

[0051] Furthermore, to test the ability of the target model to handle extreme situations, conversation test data in high-risk scenarios can be set; to test the professional ability of the target model, conversation test data in professional knowledge application scenarios can be set.

[0052] Single-round conversation test data refers to a conversation form in which only one information interaction occurs between the two parties of the conversation. Multi-round conversation test data refers to a situation where the two parties of the conversation conduct multiple information interactions and complete a complete conversation process through multiple rounds of questions and answers and exchanges. In this embodiment, the ratio between single-round conversation test data and multi-round conversation test data can be preset, for example, it can be 4:1.

[0053] When the conversation test data is single-round conversation test data, the reply data is single-round reply; when the conversation test data is multi-round conversation test data, the reply data is the set of all replies during multi-round conversations.

[0054] In this embodiment, the conversation test data is input into the target model, and the target model is evaluated according to the reply data output by the target model. Since the conversation test data includes not only single-round conversation test data but also multi-round conversation test data, when evaluating the target model in this embodiment, not only the quality of single-round replies is concerned, but also the coherence and emotional progression of replies during multi-round conversations are evaluated to achieve the overall evaluation of multi-round conversations.

[0055] The technical solution of this embodiment can deeply and comprehensively evaluate the advantages and disadvantages of the target model in actual applications through multi-dimensional test scenarios, thereby providing an accurate optimization direction for the iteration of the target model.

[0056] Furthermore, S3 can further include:

[0057] S31. Determine the first evaluation results of the reply data in at least two dimensions, and determine the second evaluation results of the reply data in at least two dimensions through other natural language models;

[0058] S32. Determine the evaluation results of the target model according to the first evaluation results and the second evaluation results.

[0059] Among them, at least two dimensions represent evaluation dimensions, for example, it may include emotional resonance ability, wisdom support ability, dialogue art level, personality charm shaping, companionship temperature creation, and character vitality display, etc. Different evaluation criteria can be set for different evaluation dimensions in advance. For example, the emotional resonance dimension focuses on whether to use a tone close to the user and whether it can timely capture the user's emotional needs, etc.

[0060] The "first" and "second" in the first evaluation results and the second evaluation results are only used to distinguish the two evaluation results, and do not represent order or sequence, etc. The evaluation results can be in the form of scores (such as 1-5 points), or in the form of words (such as to be improved, qualified, good, and excellent, etc.). The first evaluation results can be manual evaluation results. For example, evaluators can evaluate the reply data of the target model in each evaluation dimension through the user interface. The second evaluation results are the results obtained by other natural language models evaluating the reply data of the target model.

[0061] When determining the evaluation results of the target model according to the first evaluation results and the second evaluation results, on the one hand, different weights can be set for different evaluation dimensions, and on the other hand, different weights can be set for the first manual evaluation results and the second evaluation results of other natural language models.

[0062] Specifically, for the evaluation results of the same reply data in different evaluation dimensions by the same evaluation subject (manual or other natural language models), weighted summation is performed according to the weights of different evaluation dimensions to obtain the evaluation results of this evaluation subject for this reply data. According to the weights of each evaluation subject, weighted summation is performed on the evaluation results of different evaluation subjects for this reply data to obtain the evaluation results of this reply data. Taking the average value or the mode, etc., of the evaluation results of each reply data to obtain the evaluation results of the target model. Different weights can also be set for different dialogue test data. For example, higher weights can be set for multi-round dialogue test data, or higher weights can be set for dialogue test data in high-risk scenarios. According to the weights of the dialogue test data corresponding to different reply data, weighted summation is performed on the evaluation results of each reply data, and finally the evaluation results of the target model are obtained.

[0063] The technical solution of this embodiment combines human-machine collaborative evaluation and multi-dimensional multi-scenario evaluation, which can ensure the professionalism, objectivity, and accuracy of the model evaluation results. At the same time, the scenarios and evaluation dimensions of the dialogue test data can be expanded, improving the scalability of the model evaluation.

[0064] Furthermore, after determining the evaluation result of the target model, the iterative optimization results of the model can be compared longitudinally according to the evaluation results of the target model at different iteration stages; the dialogue capabilities of the model can also be compared horizontally according to the evaluation results of the target model and the evaluation results of natural language models in other scenarios or fields. This not only comprehensively evaluates the emotional companionship ability and professionalism of the target model but also effectively guides the optimization direction of the target model. Through multi-scenario testing and multi-dimensional evaluation, the advantages and disadvantages of the target model in actual applications can be deeply discovered, providing an accurate improvement direction for the iteration of the target model.

[0065] The technical solution of the embodiment of the present invention performs first instruction supervised fine-tuning training on a natural language model based on emotion training data to obtain a first intermediate model, performs second instruction supervised fine-tuning training on the first intermediate model based on professional training data to obtain a second intermediate model, and performs direct preference optimization training on the second intermediate model based on daily dialogue training data to obtain a final target model for session interaction. This solves the problem that the natural language model in the prior art cannot deeply understand the language expression habits of complex scenarios and specific groups during session interaction, cannot accurately grasp the deep emotional needs of the conversation, and lacks flexibility and adaptability. The technical solution of the present invention realizes natural, professional, and safe session interaction of the natural language model, improving the emotional expression ability, flexibility, and adaptability of the natural language model.

[0066] Embodiment 2

[0067] Figure 2 The figure is a flowchart of a model training method provided by the second embodiment of the present invention. Based on the above embodiment, the present embodiment further specifies the setting process of emotion training data, professional training data, and daily dialogue training data.

[0068] As Figure 2 shown, the method includes:

[0069] S210. Generate an initial emotion dialogue according to the initial emotion data through a multi-agent dialogue method.

[0070] Among them, the multi-agent dialogue method refers to a situation where multiple dialogue agents with different functions, roles, or characteristics participate in the dialogue process in a dialogue system or interaction environment to provide richer and more efficient dialogue services or achieve more complex tasks.

[0071] According to the different application scenarios of the natural language model, the initial sentiment data can be sentiment companionship cases, psychological counseling cases, etc. The initial sentiment dialogue is a multi-turn dialogue generated based on the initial sentiment data and can coherently and completely represent the process of sentiment companionship or psychological counseling. For example, each group of initial sentiment dialogues can include 5-10 turns.

[0072] In this embodiment, through the multi-agent dialogue method, different roles are assigned to different agents according to the initial sentiment data, such as psychological counselors, patients, etc. Key information is extracted from the initial sentiment data, and based on the extracted key information and roles, the agents conduct conversation interactions. During the conversation process, the agents also perform logical reasoning and plot advancement, and finally generate a dialogue among the multi-agents.

[0073] The technical solution of this embodiment converts the static initial sentiment data into dynamic and interactive initial sentiment dialogues through multi-agent conversations, enabling the initial sentiment dialogues to have a professional emotional healing effect while ensuring coherence and integrity.

[0074] S220. Generate sentiment training data according to the initial sentiment dialogue through prompt engineering.

[0075] Prompt engineering refers to the technology and method of guiding the natural language model to generate output content that meets specific requirements and expectations by designing and optimizing the prompts input to the natural language model.

[0076] In this embodiment, through prompt engineering, the initial sentiment dialogue is input into the natural language model to guide the natural language model to generate sentiment training data, achieve the transfer of a high-emotional-intelligence anthropomorphic style, integrate preset role information (such as lively, steady, authoritative, etc.) into the sentiment training data, and mix in self-cognition fine-tuning data to enhance the diversity and cognitive ability of the sentiment training data.

[0077] S230. Determine the safety alignment data and medical Q&A data.

[0078] Among them, the safety alignment data refers to a data set specially constructed to ensure the safety, reliability, and compliance with ethical and other requirements of the natural language model. It is a collection composed of a series of data samples with specific annotations or features. The safety alignment data aims to help the natural language model learn and understand how to follow safety guidelines, ethical norms, and legal requirements in various tasks and scenarios, so that the output and behavior of the natural model are consistent with human values and safety needs.

[0079] Medical Q&A data is a collection composed of medical-related questions and their corresponding answers. Medical Q&A data usually covers various medical fields and disease-related knowledge, aiming to help natural language models learn and understand the patterns of medical questions and the logic of answers, so as to be able to output accurate and reliable medical information and advice.

[0080] S240. Through prompt engineering, generate professional training data based on the safe alignment data and medical Q&A data.

[0081] In this embodiment, also through prompt engineering, the safe alignment data and medical Q&A data are input into the natural language model to guide the natural language model to generate professional training data. This setting realizes the anthropomorphic style alignment of the professional training data, ensuring that when performing the second-stage model training based on the professional training data, while infusing knowledge and ensuring safe alignment, it can also guarantee the high EQ and anthropomorphism of the dialogue style.

[0082] S250. Through the data flywheel method, generate positive samples and negative samples of daily conversations based on the initial daily conversation data.

[0083] Among them, the data flywheel method is a cyclic model that uses data to drive business growth. Through the continuous accumulation, analysis, and application of data, it promotes the continuous optimization and improvement of all aspects of the business, and then forms a virtuous cycle, enabling the business to continuously accelerate like a flywheel.

[0084] The initial daily conversation data can be data collected from multiple sources. The initial daily conversation data can cover a rich variety of topics and scenarios to ensure the diversity of the daily conversation training data.

[0085] Positive samples of daily conversations refer to daily conversation samples whose reply effects and interaction depths meet the requirements, and negative samples of daily conversations refer to daily conversation samples whose reply effects and interaction depths do not meet the requirements.

[0086] In this embodiment, through the data flywheel method, positive samples and negative samples of daily conversations are generated based on the initial daily conversation data. Specifically, evaluate the reply quality of the initial daily conversation data, and determine the initial daily conversation data as positive samples or negative samples of daily conversations according to the reply quality; generate new daily conversation data based on the initial daily conversation data; evaluate the reply quality of the new daily conversation data, and determine the new daily conversation data as positive samples or negative samples of daily conversations according to the reply quality; continuously repeat the above process until the quantities of positive samples and negative samples of daily conversations meet the preset quantity requirements.

[0087] The technical solution of this embodiment can generate positive and negative samples for daily conversations through the data flywheel method, which can improve the quality and quantity of daily conversation samples, providing strong data support for training a more intelligent and efficient target model subsequently.

[0088] S260. Through prompt engineering, generate training data for positive samples of daily conversations based on positive samples of daily conversations.

[0089] In this embodiment, also through prompt engineering, input positive samples of daily conversations into the natural language model to guide the natural language model to generate training data for positive samples of daily conversations. Such a setting realizes the enhancement of the anthropomorphic personality of the training data for positive samples of daily conversations, so that when the model training in the third stage is carried out subsequently, the reasoning ability and anthropomorphic expression effect of the model can be improved.

[0090] S270. Use the training data for positive samples of daily conversations, negative samples of daily conversations, pre-determined positive samples of general knowledge, and negative samples of general knowledge as training data for daily conversations.

[0091] In this embodiment, mix and proportion the training data for positive samples of daily conversations, negative samples of daily conversations, pre-determined positive samples of general knowledge, and negative samples of general knowledge. For example, the total number of training data for positive samples of daily conversations and negative samples of daily conversations can be 1000, and the ratio of the training data for positive samples of daily conversations to negative samples of daily conversations can be 4:1; the total number of positive samples of general knowledge and negative samples of general knowledge can be 10000, and the ratio is also set to 4:1.

[0092] The technical solution of this embodiment, through the mixed proportion of the training data for positive samples of daily conversations with enhanced anthropomorphic personality, negative samples of daily conversations, and positive and negative samples of general knowledge, can improve the reasoning ability and anthropomorphic expression effect of the natural language model in the model training stage.

[0093] S280. Based on the pre-set emotional training data, perform the first instruction supervised fine-tuning training on the natural language model to obtain the first intermediate model.

[0094] S290. Based on the pre-set professional training data, perform the second instruction supervised fine-tuning training on the first intermediate model to obtain the second intermediate model.

[0095] S2100. Based on the pre-set training data for daily conversations, perform direct preference optimization training on the second intermediate model to obtain the target model.

[0096] The model training processes in the above three stages have been described in the above embodiments, and will not be elaborated herein.

[0097] The technical solution of this embodiment adopts a phased training data generation method, realizes the quality control and feature enhancement of training data through the adaptive mechanism of dynamic style transfer and the dual guarantee system of security and medical professionalism. Through the anthropomorphic style transfer in the training data generation stage and the subsequent multi-stage model training, the emotional expression ability of the natural language model is significantly improved, enabling the natural language model to show stronger emotional resonance ability and a more natural high-emotional intelligence anthropomorphic dialogue style. Through the multi-stage progressive model training strategy, combined with the instruction supervision fine-tuning training technology and the dynamic adjustment mechanism of direct preference alignment, the ability performance of the model in various aspects such as naturalness, security, and professionalism is successfully balanced, ensuring both the security of the natural language model and the improvement of the professional level. At the same time, a scientific model evaluation system is established, providing a reliable basis for model optimization through the quantification method of multiple evaluation dimensions, the innovative mechanism of human-machine collaborative evaluation, and the systematic scheme of cross-model comparison.

[0098] Embodiment III

[0099] Figure 3 The following is a schematic structural diagram of a model training device provided by Embodiment III of the present invention. As Figure 3 shown, the device includes:

[0100] The first instruction supervision fine-tuning training module 310 is configured to perform first instruction supervision fine-tuning training on the natural language model based on the pre-set emotional training data to obtain a first intermediate model;

[0101] The second instruction supervision fine-tuning training module 320 is configured to perform second instruction supervision fine-tuning training on the first intermediate model based on the pre-set professional training data to obtain a second intermediate model;

[0102] The direct preference optimization training module 330 is configured to perform direct preference optimization training on the second intermediate model based on the pre-set daily conversation training data to obtain a target model.

[0103] In the technical solution of the embodiment of the present invention, through the first instruction supervised fine-tuning training of the natural language model based on the sentiment training data, a first intermediate model is obtained. Then, the second instruction supervised fine-tuning training of the first intermediate model based on the professional training data is carried out to obtain a second intermediate model. Finally, the direct preference optimization training of the second intermediate model based on the daily conversation training data is carried out to obtain the final target model for conversation interaction. This solves the problem that in the prior art, the natural language model cannot deeply understand the complex scenarios and the language expression habits of specific groups during conversation interaction, cannot accurately grasp the deep emotional needs of the conversation, and lacks flexibility and adaptability. The technical solution of the present invention realizes natural, professional, and safe conversation interaction of the natural language model, and improves the emotional expression ability, flexibility, and adaptability of the natural language model.

[0104] Based on the above embodiment, optionally, the device further includes:

[0105] An initial emotion dialogue generation module, configured to generate an initial emotion dialogue according to the initial emotion data through a multi-agent dialogue method;

[0106] A sentiment training data generation module, configured to generate sentiment training data according to the initial emotion dialogue through prompt engineering.

[0107] Based on the above embodiment, optionally, the device further includes:

[0108] A professional data determination module, configured to determine safety alignment data and medical Q&A data;

[0109] A professional training data generation module, configured to generate professional training data according to the safety alignment data and medical Q&A data through prompt engineering.

[0110] Based on the above embodiment, optionally, the device further includes:

[0111] A data flywheel module, configured to generate daily conversation positive samples and daily conversation negative samples according to the initial daily conversation data through the data flywheel method;

[0112] A daily conversation positive sample training data generation module, configured to generate daily conversation positive sample training data according to the daily conversation positive samples through prompt engineering;

[0113] A daily conversation training data generation module, configured to use the daily conversation positive sample training data, the daily conversation negative samples, the pre-determined general knowledge positive samples, and the general knowledge negative samples as the daily conversation training data.

[0114] Based on the above embodiments, optionally, the first instruction supervised fine-tuning training is instruction supervised fine-tuning training based on the low-rank adaptation technique, and the second instruction supervised fine-tuning training is instruction supervised fine-tuning training based on the quantized low-rank adaptation technique.

[0115] Based on the above embodiments, optionally, the device further includes:

[0116] A dialogue test data determination module, configured to determine dialogue test data in at least two scenarios, where the dialogue test data includes single-round dialogue test data and multi-round dialogue test data;

[0117] A reply data determination module, configured to input the dialogue test data into a target model to obtain reply data matching the dialogue test data;

[0118] An evaluation result determination module, configured to determine an evaluation result of the target model according to the reply data.

[0119] Based on the above embodiments, optionally, the evaluation result determination module includes:

[0120] A multi-dimensional evaluation result determination unit, configured to determine a first evaluation result of the reply data in at least two dimensions, and determine a second evaluation result of the reply data in at least two dimensions through other natural language models;

[0121] The evaluation result of the target model determines the evaluation result of the target model according to the first evaluation result and the second evaluation result.

[0122] The model training device provided by the embodiments of the present invention can execute the model training method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0123] Embodiment 4

[0124] Figure 4 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0125] As Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0126] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0127] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the model training method.

[0128] In some embodiments, the model training method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the model training method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the model training method by any other appropriate means (e.g., by means of firmware).

[0129] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0130] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0131] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0132] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user can be in any form (including acoustic input, voice input, or tactile input).

[0133] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0134] A computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0135] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0136] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A model training method, characterized in that: include: Based on the preset sentiment training data, the natural language model is fine-tuned by first instruction supervision to obtain a first intermediate model; Based on the preset professional training data, the first intermediate model is fine-tuned by performing second instruction supervision training to obtain a second intermediate model; Based on the preset daily conversation training data, the second intermediate model is directly trained for preference optimization to obtain a target model.

2. The model training method according to claim 1, characterized in that: The setting process of the emotion training data includes: Generate initial emotional dialogue based on initial emotional data through multi-agent dialogue method; Through the prompt word engineering, emotional training data is generated based on the initial emotional dialogue.

3. The model training method according to claim 1, characterized in that: The process of setting the professional training data includes: Identify secure alignment data and medical question and answer data; Through prompt word engineering, professional training data is generated based on security alignment data and medical question and answer data.

4. The model training method according to claim 1, characterized in that: The process of setting the daily conversation training data includes: Through the data flywheel method, daily conversation positive samples and daily conversation negative samples are generated based on the initial daily conversation data; Through the prompt word engineering, we generate positive sample training data of daily conversations based on positive samples of daily conversations; Daily conversation positive sample training data, daily conversation negative samples, predetermined general knowledge positive samples, and general knowledge negative samples are used as daily conversation training data.

5. The model training method according to any one of claims 1 to 4, characterized in that: The first instruction supervised fine-tuning training is instruction supervised fine-tuning training based on low-rank adaptive technology, and the second instruction supervised fine-tuning training is instruction supervised fine-tuning training based on quantized low-rank adaptive technology.

6. The model training method according to claim 1, characterized in that: After obtaining the target model, it also includes: Determine dialogue test data under at least two scenarios, wherein the dialogue test data includes single-round dialogue test data and multi-round dialogue test data; Inputting the dialogue test data into the target model to obtain response data matching the dialogue test data; An evaluation result of the target model is determined according to the response data.

7. The model training method according to claim 6, characterized in that: Determining the evaluation result of the target model according to the response data includes: Determine a first evaluation result of the reply data in at least two dimensions, and determine a second evaluation result of the reply data in at least two dimensions by using other natural language models; An evaluation result of the target model is determined according to the first evaluation result and the second evaluation result.

8. A model training device, characterized in that: include: A first instruction supervised fine-tuning training module is used to perform first instruction supervised fine-tuning training on the natural language model based on preset sentiment training data to obtain a first intermediate model; A second instruction supervised fine-tuning training module is used to perform second instruction supervised fine-tuning training on the first intermediate model based on pre-set professional training data to obtain a second intermediate model; The direct preference optimization training module is used to perform direct preference optimization training on the second intermediate model based on preset daily conversation training data to obtain a target model.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the model training method as described in any one of claims 1-7 is implemented.

10. A storage medium storing computer executable instructions, characterized in that: The computer executable instructions, when executed by a computer processor, are used to execute the model training method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Answer feedback method and device applied to large language model

    CN117708292A

  • Large model-based medical field model evaluation method and system

    CN118152768A

  • Large language model security alignment training method and device, electronic equipment and medium

    CN118966299A

  • Intelligent dialogue method, device, computer equipment, storage medium and product

    CN119739817A

  • Psychological counseling simulation method and system based on fine-tuning large language model

    CN119884297A

Cited By

  • Robot motion planning method and system based on local sub-target guided learning

    CN121492037A