Model training method, model-based interaction method and corresponding device and equipment

By evaluating the model's results in the training sub-stages and preset threshold conditions, the training mode is dynamically selected, which solves the problem of resource waste in inference chain training of large language models and improves the model's complex task processing capabilities and training efficiency.

CN120804715APending Publication Date: 2025-10-17北京中关村科金技术有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511012379.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

After the introduction of inference chain training, the existing large language models have increased resource overhead and time consumption, resulting in low training efficiency. In particular, the response is lengthy and the generalization ability is weakened when facing simple tasks.

Method used

By evaluating the results of the model in the current training sub-stage and the preset evaluation threshold conditions, the training mode is dynamically selected, including direct inference and inference chain training, to optimize the forward reasoning method during the model training process.

Benefits of technology

It improves the model's processing and reasoning capabilities for complex tasks, while alleviating overfitting and redundant training for simple tasks, improving training efficiency and effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804715A_ABST
    Figure CN120804715A_ABST
Patent Text Reader

Abstract

The invention provides a model training method, a model-based interaction method and corresponding devices and equipment, and belongs to the technical field of artificial intelligence. The method comprises the steps of evaluating a to-be-processed model to obtain an evaluation result of the to-be-processed model in a current training sub-stage; according to the evaluation result and a preset evaluation threshold condition corresponding to the current training sub-stage, a training mode of the to-be-processed model in the next training sub-stage is determined, and the training mode is used for indicating a processing mode of forward reasoning of the to-be-processed model in the training process; and training the to-be-processed model based on the determined training mode in the next training sub-stage. According to the embodiment of the invention, the training effect and the training efficiency of the model can be effectively considered.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, in particular to a model training method, a model-based interaction method, a model training device, a model-based interaction device, an electronic device, a computer readable storage medium, and a computer program product. BACKGROUND

[0002] In recent years, large language models have achieved remarkable results in task-oriented dialogues and text generation. As the requirements of reasoning tasks (such as legal question answering, mathematical problem solving, and complex dialogue generation) for model logic chains are increasing, the Chain-of-Thought (CoT) training method is widely used: that is, the model not only outputs the answer, but also needs to show the detailed reasoning process. By introducing the training method of the reasoning chain, the training effect of the model is improved. However, after introducing the reasoning chain, the resource consumption and time consumption required for model training will also increase accordingly. SUMMARY

[0003] The present application provides a model training method, a model-based interaction method, a model training device, a model-based interaction device, an electronic device, a computer readable storage medium, and a computer program product, which can select a reasonable training mode according to the evaluation result of the model, so as to balance the training effect and training efficiency of the model.

[0004] In a first aspect, the present application provides a model training method, which comprises: evaluating a to-be-processed model to obtain an evaluation result of the to-be-processed model in a current training sub-stage; determining a training mode of the to-be-processed model in a next training sub-stage according to the evaluation result and a preset evaluation threshold condition corresponding to the current training sub-stage, the training mode being used to indicate a processing mode of the to-be-processed model in a forward reasoning process; and training the to-be-processed model based on the determined training mode in the next training sub-stage.

[0005] In a second aspect, the present application provides a model-based interaction method, which comprises: receiving user dialogue data sent by a user terminal; processing the user dialogue data based on a preset dialogue model to obtain agent reply data corresponding to the user dialogue data; and sending the agent reply data to the user terminal; wherein the preset dialogue model is obtained based on the model training method of any one of the embodiments of the present application.

[0006] In a third aspect, the present application provides a model training apparatus, comprising: an evaluation module configured to evaluate a to-be-processed model to obtain an evaluation result of the to-be-processed model in a current training sub-stage; a determination module configured to determine a training mode of the to-be-processed model in a next training sub-stage according to the evaluation result and a preset evaluation threshold condition corresponding to the current training sub-stage, the training mode being used to indicate a processing mode of the to-be-processed model in a forward inference in a training process; and a training module configured to train the to-be-processed model based on the determined training mode in the next training sub-stage.

[0007] In a fourth aspect, the present application provides a model-based interaction apparatus, comprising: a receiving module configured to receive user conversation data sent by a user terminal; a processing module configured to process the user conversation data based on a preset dialogue model to obtain agent reply data corresponding to the user conversation data; and a sending module configured to send the agent reply data to the user terminal; wherein the preset dialogue model is obtained based on the model training method of any one of the embodiments of the present application.

[0008] In a fifth aspect, the present application provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, and the one or more computer programs are executed by the at least one processor to enable the at least one processor to perform the model training method or the model-based interaction method described above.

[0009] In a sixth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor / processing core, implements the model training method or the model-based interaction method described above.

[0010] In a seventh aspect, the present application provides a computer program product comprising computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the model training method or the model-based interaction method described above.

[0011] The embodiments provided in the application evaluate the to-be-processed model to obtain an evaluation result of the to-be-processed model in a current training sub-stage; determine a training mode of the to-be-processed model in a next training sub-stage according to the evaluation result and a preset evaluation threshold condition corresponding to the current training sub-stage, the training mode being used to indicate a processing manner of the to-be-processed model in forward inference in a training process; and train the to-be-processed model based on the determined training mode in the next training sub-stage.

[0012] In the model training process, the training mode used to train the to-be-processed model in the next training sub-stage can be determined according to the evaluation result of the to-be-processed model in the current training sub-stage and the preset evaluation threshold condition corresponding to the current training sub-stage. The processing manner of the model in forward inference in the training process can be determined through the training mode. Therefore, on the one hand, the processing capability and inference capability of the model for complex tasks can be effectively improved, and on the other hand, overfitting and redundant training for simple tasks can be effectively alleviated. In addition, in full consideration of the case that the training focus of the model may be different in different training stages, the preset evaluation threshold condition corresponding to each training sub-stage is set, so that a more reasonable and accurate training mode can be determined for the to-be-processed model based on the preset evaluation threshold condition, which helps to improve the model training effect and training efficiency. In summary, the model training method of the embodiments of the present disclosure can effectively balance the training effect and training efficiency of the model.

[0013] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent to those skilled in the art through the following description. BRIEF DESCRIPTION OF DRAWINGS

[0014] The accompanying drawings are intended to provide a further understanding of the present application and constitute a part of the specification, together with the embodiments of the present application, to explain the present application and do not constitute a limitation of the present application. The above and other features and advantages will become more apparent to those skilled in the art through a detailed description of the specific example embodiments, with reference to the accompanying drawings, in which:

[0015] Figure 1 The application scenario diagram of the model training method or the model-based interaction method and the corresponding device provided by the embodiments of the present disclosure.

[0016] Figure 2 The flowchart of the model training method provided by the embodiments of the present application.

[0017] Figure 3 The flowchart of the model-based interaction method provided by the embodiments of the present application.

[0018] Figure 4A block diagram of a model training apparatus provided for an embodiment of the present application.

[0019] Figure 5 A block diagram of a model-based interaction apparatus provided for an embodiment of the present application.

[0020] Figure 6 A block diagram of an electronic device provided for an embodiment of the present application. DETAILED DESCRIPTION

[0021] To enable persons skilled in the art to better understand the technical solutions of the present disclosure, exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to help understanding, and should be considered as merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, for the sake of clarity and conciseness, the description below omits the description of well-known functions and structures.

[0022] In the case of no conflict, each embodiment of the present disclosure and each feature in the embodiments can be combined with each other.

[0023] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0024] The terms used herein are only used to describe specific embodiments, and are not intended to limit the present disclosure. As used herein, the singular forms "a" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the terms "comprise" and / or "consist of," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. The terms "connected" or "coupled" and / or similar terms are not limited to a physical or mechanical connection, but can include an electrical connection, whether direct or indirect.

[0025] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted in an overly literal or overly formal sense unless expressly so defined herein.

[0026] In the technical solutions of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information comply with relevant laws and regulations and do not violate public order and good customs. The use of user data in the technical solutions complies with relevant national laws and regulations (for example, the Information Security Technology Personal Information Security Specification). For example, appropriate measures are taken for personal information access control, and restrictions are imposed on the display of personal information. The use of personal information does not exceed the direct or reasonably related range, and the use of personal information eliminates explicit identity pointers and avoids pinpointing specific individuals.

[0027] With the increasing demand for logical chains of models for reasoning tasks (for example, legal question answering, mathematical problem solving, and complex conversation generation), the "thinking chain" is widely used in models: that is, the model not only outputs the answer, but also needs to show the detailed reasoning process.

[0028] However, in some application scenarios (such as the multi-turn dialogue generation task in the outbound call system), a large number of interactions only involve simple question and answer and template confirmation sentences, and do not require complex reasoning. If the thinking chain training is forcibly added for all input samples, the training efficiency will be reduced, the resource consumption will be increased, and the model will respond in a long-winded manner and have weakened generalization ability when facing simple tasks, thereby affecting the user experience.

[0029] Therefore, the present disclosure provides a model training method, a model-based interaction method, a model training device, a model-based interaction device, an electronic device, a computer readable storage medium, and a computer program product.

[0030] According to the model training method provided in the embodiments of the present disclosure, in the model training process, the training mode for training the to-be-processed model in the next training sub-stage can be determined according to the evaluation result of the to-be-processed model in the current training sub-stage and the preset evaluation threshold condition corresponding to the current training sub-stage. The training mode can determine the processing mode of the model in the forward reasoning in the training process. Therefore, on the one hand, the processing capability and reasoning capability of the model for complex tasks can be effectively improved, and on the other hand, the overfitting and redundant training for simple tasks can be effectively alleviated. In addition, in full consideration of the case that the model training focus may be different in different training stages, the preset evaluation threshold condition corresponding to each training sub-stage is set, so that a more reasonable and accurate training mode can be determined for the to-be-processed model based on the preset evaluation threshold condition, which helps to improve the model training effect and training efficiency. In summary, the model training method provided in the embodiments of the present disclosure can effectively balance the training effect and training efficiency of the model.

[0031] The model training method according to the embodiments of the present application can be executed by an electronic device such as a terminal device or a server. The terminal device can be a vehicle-mounted device, a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The method can be implemented by a processor invoking computer readable program instructions stored in a memory. Alternatively, the method can be executed by a server.

[0032] Figure 1 An application scenario of the model training method or the model-based interaction method and the corresponding device provided by the embodiments of the present application is shown in the following figure.

[0033] As shown in the figure, the application scenario of the embodiments of the present application can include a terminal device 101, a network 103 and a server 102. The network 103 is a medium for providing a communication link between the terminal device 101 and the server 102. The network 103 can include various connection types, such as wired, wireless communication links or optical fiber cables, etc. Figure 1

[0034] A user can use the terminal device 101 to interact with the server 102 through the network 103 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0035] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablet computers, laptop computers and desktop computers, etc.

[0036] The server 102 can be a server providing various services, such as a background management server supporting the website browsed by the user using the terminal device 101 (only as an example). The background management server can analyze and process the received user request data, etc., and feed back the processing result (such as a webpage, information or data, etc. obtained or generated according to the user request) to the terminal device.

[0037] ​It should be noted that the model training method or the model-based interaction method and the corresponding device provided by the embodiments of the present disclosure can be executed by the server 102. Accordingly, the model training method or the model-based interaction method and the corresponding device provided by the embodiments of the present disclosure can be arranged in the server 102. The model training method or the model-based interaction method and the corresponding device provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 102 and capable of communicating with the terminal device 101 and / or the server 102. Accordingly, the model training method or the model-based interaction method and the corresponding device provided by the embodiments of the present disclosure can also be arranged in a server or a server cluster different from the server 102 and capable of communicating with the terminal device 101 and / or the server 102.

[0038] It should be understood that Figure 1 The number of terminal devices, networks and servers in the system is only illustrative. Any number of terminal devices, networks and servers can be provided according to the needs of implementation.

[0039] In a first aspect, the embodiments of the present application provide a model training method.

[0040] Figure 2 A flowchart of a model training method provided by the embodiments of the present application is shown in FIG. 2. Referring to FIG. 2, Figure 2 The model training method can include the following steps.

[0041] In step S201, the model to be processed is evaluated to obtain an evaluation result of the model to be processed in the current training sub-stage.

[0042] In step S202, according to the evaluation result and a preset evaluation threshold condition corresponding to the current training sub-stage, a training mode of the model to be processed in the next training sub-stage is determined, and the training mode is used to indicate a processing manner of forward inference of the model to be processed in the training process.

[0043] In step S203, the model to be processed is trained based on the determined training mode in the next training sub-stage.

[0044] Therefore, in the embodiments of the present disclosure, the training process of the model to be processed is divided into multiple training sub-stages, and preset evaluation threshold conditions corresponding to each training sub-stage are set. For any one of the training sub-stages, the model to be processed in the current training sub-stage is evaluated to obtain an evaluation result of the model to be processed in the current training sub-stage, and the evaluation result is compared with the preset evaluation threshold condition corresponding to the current training sub-stage to determine whether the evaluation result meets the preset evaluation threshold condition corresponding to the current training sub-stage, and then the training mode used to train the model to be processed in the next training sub-stage is determined according to the comparison result.

[0045] Further, the training mode is used to indicate the processing manner of the to-be-processed model in the training process for forward reasoning, for example, it can indicate whether to directly perform forward reasoning and output a processing result or to start a reasoning chain, perform forward reasoning through the reasoning chain, and finally output the processing result. The reasoning chain can represent a systematic logical reasoning step from initial information (i.e., input sample data) to a final conclusion (i.e., an output processing result). The core of the reasoning chain is to connect scattered information into a chain with a cause-and-effect or logical relationship through a series of associated reasoning steps, and finally reach the target conclusion.

[0046] It should be noted that, compared with the training mode without starting the reasoning chain, the training mode with starting the reasoning chain requires relatively longer processing time, relatively more resource consumption of processing resources, storage resources, and the like, needs more resource overhead, and the processing result is theoretically more accurate due to a series of reasoning steps. However, for some simple input sample data, complex reasoning based on the reasoning chain is not needed. If the forward reasoning based on the reasoning chain is started in the entire model training process without distinction, the training time will be too long, the training efficiency will be low, a large amount of resources will be consumed, in addition, the model will respond slowly when facing simple tasks, the generalization ability will be weakened, and the allocation of training resources will be unreasonable. Therefore, in the embodiments of the present disclosure, a suitable training mode can be determined according to needs, so as to balance the training efficiency and the training effect of the model, and achieve reasonable allocation of training resources.

[0047] The model training method of the embodiments of the present disclosure will be described below.

[0048] In some optional embodiments, the to-be-processed model is a model to be trained, which can be a large language model (LLM), an image generation model, a multi-modal large model, and the like, and the embodiments of the present disclosure do not limit the to-be-processed model.

[0049] In some optional embodiments, the training process of the to-be-processed model can be divided into multiple training sub-phases, so that each training sub-phase can be evaluated to determine the training mode of the subsequent training sub-phase.

[0050] In some optional embodiments, each training sub-phase includes one training epoch. One training epoch corresponds to a process of completely traversing the entire training data set once in the model training process.

[0051] For example, if the training set includes 1000 training sample data, and the batch size is set to 100, then 1 epoch includes 10 iterations, that is, 1000 training sample data are processed through 10 batches, and 1 training round is completed. Moreover, a single epoch is usually insufficient for the model to fully learn the regularity in the data, so multiple iterations (multiple epochs) are needed to allow the model to continuously adjust its parameters (such as weights and biases), thereby gradually reducing the prediction error and improving the processing performance of the model.

[0052] In some optional embodiments, each training sub-stage corresponds to a preset time period, where the preset time period corresponds to a time period, and the value of the preset time period can be determined according to experience, statistical data, simulation data, etc., and the embodiments of the present disclosure do not limit this.

[0053] It should be noted that the above is only an example of a training sub-stage, and the embodiments of the present disclosure do not limit this.

[0054] In some optional embodiments, the training mode includes a first training mode and a second training mode, the first training mode is a processing mode of directly outputting an inference result, and the second training mode is a processing mode of outputting an inference result based on an inference chain.

[0055] In other words, the first training mode is a training mode without starting an inference chain, and the model to be processed directly outputs a training result when training based on the first training mode, and does not perform inference based on an inference chain in this process. The second training mode is a training mode with starting an inference chain, and the model to be processed starts an inference chain during the forward inference process and obtains and outputs the final processing result through a series of logical derivation steps. Compared with the first training mode, the second training mode requires longer data processing time, consumes more resources, and may have relatively higher accuracy of the processing result.

[0056] In some optional embodiments, the model to be processed is evaluated to obtain an evaluation result of the model to be processed in the current training sub-stage, including: in response to satisfying a preset starting condition, training the model to be processed based on preset training sample data and using the first training mode to obtain output result data corresponding to the training sample data; and determining the evaluation result of the model to be processed in the current training sub-stage according to the training sample data and the output result data. The preset starting condition is a condition for starting model evaluation. If the preset starting condition is met, the evaluation of the model to be processed is started. If the preset starting condition is not met, the evaluation of the model to be processed is not started.

[0057] In some optional embodiments, the preset starting condition comprises at least one of the following: reaching a preset evaluation period, completing a training round.

[0058] For example, if the time period (t1-t2) between the current time t1 and the last time t2 of evaluating the to-be-processed model is greater than or equal to the preset evaluation period T, it is determined that the preset starting condition is met, and the to-be-processed model can be evaluated.

[0059] For example, if one training round of the to-be-processed model is completed, it is determined that the preset starting condition is met, and the to-be-processed model can be evaluated.

[0060] In some optional embodiments, the evaluation result can reflect the current processing capability or processing performance of the to-be-processed model.

[0061] In some optional embodiments, the evaluation result comprises at least one of perplexity, accuracy, precision, recall, and result error; and the preset evaluation threshold condition comprises at least one of the following: the perplexity is less than or equal to a preset perplexity threshold, the accuracy is greater than or equal to a preset accuracy threshold, the precision is greater than or equal to a preset precision threshold, the recall is greater than or equal to a preset recall threshold, and the result error is less than or equal to a preset result error threshold.

[0062] For example, perplexity (PPL) is an important indicator for measuring the prediction capability of a language model, and is mainly used to evaluate the "prediction fluency" or "uncertainty" of the model for a given text sequence. In simple terms, the lower the perplexity, the more accurate and "unpuzzling" the prediction of the model for the text. Correspondingly, the preset evaluation threshold condition comprises that the perplexity is less than or equal to a preset perplexity threshold, wherein the number of the preset perplexity thresholds is multiple, and each preset perplexity threshold corresponds to a training sub-stage. For example, for the i-th training sub-stage, after evaluating the current to-be-processed model, the perplexity PPL(i) is obtained, and the perplexity PPL(i) is compared with the preset perplexity threshold PPL_Thr(i) corresponding to the i-th training sub-stage. If PPL(i)≤PPL_Thr(i), it indicates that the evaluation result does not meet the preset evaluation threshold condition, and if PPL(i)>PPL_Thr(i), it indicates that the evaluation result meets the preset evaluation threshold condition.

[0063] For example, the perplexity can be calculated by the following formula: PPL=exp{(-1 / T)×∑ (t=1,2, … ,T) [logP(y t| y <t , x)]}, wherein y t represents the t-th token in the script data, y<t represents the first to the t-1th word units in the dialogue data, x represents the context of the input dialogue data, and T represents the total length of the dialogue data.

[0064] Exemplarily, the accuracy rate can represent the proportion of the number of samples predicted correctly by the model to the total number of samples, and is used to measure the overall "correctness" of the model. Correspondingly, the preset evaluation threshold condition includes that the accuracy rate is greater than or equal to a preset accuracy threshold, wherein the number of preset accuracy thresholds is multiple, and each preset accuracy threshold corresponds to a training sub-stage.

[0065] Exemplarily, the precision rate is mainly for samples predicted as positive, and measures the proportion of samples that are actually positive. Correspondingly, the preset evaluation threshold condition includes that the precision rate is greater than or equal to a preset precision threshold, wherein the number of preset precision thresholds is multiple, and each preset precision threshold corresponds to a training sub-stage.

[0066] Exemplarily, the recall rate is for samples that are actually positive, and measures the proportion of samples that are correctly predicted as positive by the model. Correspondingly, the preset evaluation threshold condition includes that the recall rate is greater than or equal to a preset recall threshold, wherein the number of preset recall thresholds is multiple, and each preset recall threshold corresponds to a training sub-stage.

[0067] Exemplarily, the result error can represent the difference between the predicted value of the model and the true value, and is used to measure the "degree of deviation of the prediction from the truth". Correspondingly, the preset evaluation threshold condition includes that the result error is greater than or equal to a preset result error threshold, wherein the number of preset result error thresholds is multiple, and each preset result error threshold corresponds to a training sub-stage.

[0068] In some optional embodiments, the training proportion of the second training mode can be gradually increased to improve the training effect of the model, so that the model can focus more on processing complex tasks and improve the processing capability of the model for complex tasks.

[0069] In some optional embodiments, each training sub-stage includes a training round; Correspondingly, the model training method can further include: determining a preset evaluation threshold corresponding to each training round according to a preset drop adjustment parameter, to obtain the preset evaluation threshold corresponding to each training sub-stage; wherein the drop adjustment parameter is used to represent the degree of drop of the value of the preset evaluation threshold with the increase of the training round.

[0070] For example, the preset evaluation threshold can gradually decrease with the increase of the training epoch, and the preset evaluation threshold can be represented as: θ(k) = θ0*exp(-γk), where k represents the training epoch, k is an integer greater than or equal to 1, θ0 represents an initial evaluation threshold, θ(k) represents the preset evaluation threshold corresponding to the kth training epoch, and γ represents a decrease adjustment parameter, which can be 0.01, etc.

[0071] Further, for the kth training epoch, if PPL(k)≤θ PPL (k) is calculated, it is determined that the training mode of the to-be-processed model in the k+1th training epoch is the first training mode; if PPL(k) > θ PPL (k) is calculated, it is determined that the training mode of the to-be-processed model in the k+1th training epoch is the second training mode. Wherein, θ PPL (k) is the preset perplexity threshold corresponding to the kth training epoch.

[0072] In some optional embodiments, the relationship between the value of the preset evaluation threshold and the training sub-stage can also be other forms. For example, the value of the preset evaluation threshold gradually decreases with the increase of the training epoch. For example, the value of the preset evaluation threshold first increases and then decreases with the increase of the training epoch.

[0073] It should be noted that the above evaluation results and preset evaluation threshold conditions are only examples, and the embodiments of the present disclosure do not limit this.

[0074] In some optional embodiments, considering that an evaluation result can only evaluate the model from one dimension and cannot comprehensively reflect the processing ability and processing performance of the model, the evaluation of the model can be realized by multiple evaluation results. For example, perplexity and accuracy can be used to evaluate the to-be-processed model together. As for which evaluation results are selected to evaluate the to-be-processed model together, it can be determined according to experience, statistical data, simulation data, model functions, and actual application scenarios, and the embodiments of the present disclosure do not limit this.

[0075] In some optional embodiments, according to the evaluation result and the preset evaluation threshold condition corresponding to the current training sub-stage, the training mode of the to-be-processed model in the next training sub-stage is determined, including: in the case that the evaluation result does not satisfy the preset evaluation threshold condition, determining that the training mode of the to-be-processed model in the next training sub-stage is the first training mode; in the case that the evaluation result satisfies the preset evaluation threshold condition, determining that the training mode of the to-be-processed model in the next training sub-stage is the second training mode.

[0076] In other words, if the evaluation result does not satisfy the preset evaluation threshold condition, it indicates that the current processing capability or processing performance of the to-be-processed model is good, and therefore, a simpler first training mode can be used in the next training sub-stage to reduce resource overhead, shorten the training time, and improve the training efficiency; otherwise, if the evaluation result does not satisfy the preset evaluation threshold condition, it indicates that the current processing capability or processing performance of the to-be-processed model is good, and therefore, a more complex second training mode can be used in the next training sub-stage to improve the model training effect.

[0077] Further, to further improve the model training effect while minimizing resource overhead in the training process, a reasonable inference chain type can be selected according to requirements.

[0078] In some optional embodiments, the inference chain includes multiple types, and in the second training mode, one type can be selected from the multiple types for model training.

[0079] In some optional embodiments, the types of the inference chain include an explicit inference chain and an implicit inference chain. In the explicit inference chain, each logical step can be explicitly indicated, and the path from the premise (corresponding to the input sample data) to the conclusion (corresponding to the output processing result) is clear. In this processing mode, the inference chain has strong interpretability and is intuitive, and the logical chain of the inference chain can be conveniently verified. The implicit inference chain refers to an inference mode in which the logical steps are not completely explicitly indicated. In other words, this inference chain processing mode is to let the model perform forward inference calculation in its internal latent space. This inference mode mainly considers that language is a tool for human communication, not a carrier of thinking itself. The human brain can maintain a long and coherent inference chain in the latent space with high efficiency, and does not need to continuously translate the inference steps back to the language form during the inference process. Based on this, the model can also perform forward logical inference in a specific internal space to achieve model training based on the inference chain.

[0080] In some optional embodiments, for the implicit inference chain, some pre-set prompt information can be added to appropriately guide the inference direction of the model and accelerate the inference process of the model, to achieve implicit inference chain based on a prompt template. The prompt template is a pre-set template that can guide the model to generate an output that meets the expectations in the inference stage. The essence is to "constrain the inference path of the model through the input format". In addition, for the implicit inference chain, some intermediate supervision information can also be added, for example, the model is "taught" the correct inference steps through intermediate labels. This way can also guide the inference direction of the model and accelerate the inference process of the model, to achieve implicit inference based on intermediate supervision.

[0081] Exemplarily, the types of reasoning chains include at least one of the following: explicit reasoning chains, fully implicit reasoning chains, implicit reasoning chains based on prompt templates, and implicit reasoning chains based on intermediate supervision. A fully implicit reasoning chain is a completely implicit reasoning chain.

[0082] It should be noted that different types of inference chains may require different training resources and play different roles. Therefore, based on the evaluation results of the model to be processed, a suitable inference chain can be selected from multiple types of inference chains as the training mode for the model to be processed in the next training sub-stage, and the model to be processed can be trained based on the selected training mode in the next training sub-stage.

[0083] In some optional embodiments, after determining that the training mode of the model to be processed in the next training sub-stage is the second training mode, the model training method may further include: selecting a target reasoning chain type from multiple reasoning chain types as the reasoning chain type adopted by the model to be processed in the next training sub-stage based on the difference between the evaluation value and the corresponding preset evaluation threshold.

[0084] For example, determine the degree to which multiple inference chain types improve the training effect. If the difference between the evaluation value and the corresponding preset evaluation threshold is large, then select the inference chain type that improves the training effect more as the training mode for the next training sub-stage. If the difference between the evaluation value and the corresponding preset evaluation threshold is small, then select the inference chain type that improves the training effect less as the training mode for the next training sub-stage.

[0085] In some optional embodiments, the model to be processed is trained based on the determined training mode in the next training sub-stage, including: determining a loss function corresponding to the training mode; determining the loss value of the model to be processed with respect to the loss function based on the training mode; and adjusting at least part of the model parameters of the model to be processed according to the loss value to obtain the adjusted model to be processed.

[0086] From this, we can see that for different training modes, we can select the loss function that matches the training mode to adjust the parameters of the model to be processed.

[0087] For example, the loss function L of the model to be processed can be represented as:

[0088]

[0089] Among them, L1 represents the first loss sub-function, which is applicable to the first training mode, L2 is the second loss sub-function, and L1+α×L2 is the loss function applicable to the second training mode. α is the loss balance coefficient, which can balance the first loss sub-function and the second loss sub-function and adjust the role ratio of the first loss sub-function and the second loss sub-function in the model parameter adjustment process.

[0090] It should be noted that the loss balance coefficient can be a fixed value, or a value dynamically matched with the evaluation value. Illustratively, the loss balance coefficient can be in direct proportion to the evaluation value, and the greater the evaluation value, the greater the loss balance coefficient, so that the proportion of the second loss sub-function in the second training mode is greater. Illustratively, the loss balance coefficient can also be in inverse proportion to the evaluation value, and the greater the evaluation value, the smaller the loss balance coefficient, so that the proportion of the second loss sub-function in the second training mode is smaller. For example, the loss balance coefficient can be the reciprocal of the perplexity.

[0091] In some optional embodiments, an iteration condition related to the to-be-processed model can also be pre-set, and if the iteration condition is met, the model is continued to be trained, if the iteration condition is not met, the model is stopped to be trained, and the to-be-processed model at this time is taken as a trained model. The iteration condition can be, for example, that the number of iterations is less than or equal to a pre-set iteration threshold, and the present disclosure embodiments do not limit this.

[0092] In summary, in the present disclosure embodiments, in the model training process, the training mode to be used for training the to-be-processed model in the next training sub-stage can be determined according to the evaluation result of the to-be-processed model in the current training sub-stage and the pre-set evaluation threshold condition corresponding to the current training sub-stage. The training mode can determine the processing manner of the model in the training process. Therefore, on the one hand, the processing capability and inference capability of the model for complex tasks can be effectively improved, and on the other hand, the overfitting and redundant training for simple tasks can be effectively alleviated. In addition, in full consideration of the case that the training focus of the model can be different in different training stages, the pre-set evaluation threshold condition corresponding to each training sub-stage is set, so that a more reasonable and accurate training mode can be determined for the to-be-processed model based on the pre-set evaluation threshold condition, which helps to improve the model training effect and training efficiency. In summary, the model training method of the present disclosure embodiments can effectively balance the training effect and training efficiency of the model.

[0093] In a second aspect, the present disclosure provides a model-based interaction method.

[0094] Figure 3 A flowchart of a model-based interaction method provided by the present disclosure embodiments is shown in FIG. 3. Referring to FIG. 3, Figure 3 The model-based interaction method can include the following steps.

[0095] In step S301, user conversation data sent by a user terminal is received.

[0096] In step S302, the user conversation data is processed based on a pre-set dialogue model to obtain agent reply data corresponding to the user conversation data.

[0097] Step S303, the agent reply data is sent to the user terminal.

[0098] The preset dialogue model is obtained based on the model training method of any embodiment of the present disclosure.

[0099] For example, in a certain interaction process, the user dialogue data sent by the user terminal includes: I have doubts about the payment method of this decoration scheme. After the agent terminal receives the user dialogue data, the preset dialogue model is used to analyze the user dialogue data, and the inference chain includes: the user may worry that the payment method is not flexible enough → whether there is a scheme for payment by stages → whether to support loans or preferential treatment → the user should be comforted first and guided to provide more specific problems. Based on this, the agent reply data corresponding to the aforementioned user dialogue data is: Hello, the payment method supports payment by stages, and it can also be combined with the loan scheme. Which scheme do you want to introduce in detail?

[0100] For example, in a certain interaction process, the user dialogue data sent by the user terminal includes: Please introduce your basic business. After the agent terminal receives the user dialogue data, the preset dialogue model is used to analyze the user dialogue data, and it is determined that the problem corresponding to the user dialogue data is relatively simple, so the inference chain does not need to be introduced, and the corresponding agent reply data is directly given: Hello, our basic business includes voice service, video service and data traffic service.

[0101] It should be noted that whether the preset dialogue model starts the inference chain can be pre-set by the manager, or can be judged by the preset dialogue model according to the user dialogue data whether to start, and the present disclosure does not limit this. For example, if the preset dialogue model determines that the user's intention is complex according to the user dialogue data, the inference chain can be started. For example, if the preset dialogue model determines that the user's intention meets a subgraph in the preset knowledge graph according to the user dialogue data, the inference chain can not be started, but the corresponding subgraph is directly used to give the corresponding agent reply data.

[0102] In the present embodiment, since the preset dialogue model is trained based on the model training method of the present disclosure, it has strong processing capability for complex tasks and can quickly respond to simple tasks, so it can quickly and accurately give the matching agent reply data according to the user dialogue data, improve the interaction effect, and thus help to improve the processing effect and efficiency of the dialogue interaction task.

[0103] It can be understood that the above-mentioned various method embodiments mentioned in the present application can be combined with each other to form combined embodiments without deviating from the principle logic. Limited by the length of the article, the present application will not be described again. Those skilled in the art can understand that in the above-mentioned method of the specific embodiment, the specific execution order of each step should be determined according to its function and possible internal logic.

[0104] Figure 4 A block diagram of a model training device provided by an embodiment of the present application.

[0105] With reference to Figure 4 The embodiment of the present application provides a model training device, which can include the following modules.

[0106] The evaluation module 401 is configured to evaluate the to-be-processed model to obtain an evaluation result of the to-be-processed model in the current training sub-stage.

[0107] The determination module 402 is configured to determine a training mode of the to-be-processed model in the next training sub-stage according to the evaluation result and a preset evaluation threshold condition corresponding to the current training sub-stage, the training mode being used to indicate a processing manner of forward inference of the to-be-processed model in the training process.

[0108] The training module 403 is configured to train the to-be-processed model based on the determined training mode in the next training sub-stage.

[0109] In the embodiment of the present application, during the model training process, the training mode of the to-be-processed model in the next training sub-stage can be determined according to the evaluation result of the to-be-processed model in the current training sub-stage and the preset evaluation threshold condition corresponding to the current training sub-stage. The processing manner of forward inference of the model in the training process can be determined through the training mode. Therefore, on the one hand, the processing capability and inference capability of the model for complex tasks can be effectively improved, and on the other hand, the overfitting and redundant training for simple tasks can be effectively alleviated. In addition, in full consideration of the case that the training focus of the model in different training stages can be different, the preset evaluation threshold condition corresponding to each training sub-stage is set, so that a more reasonable and accurate training mode of the to-be-processed model can be determined based on the preset evaluation threshold condition, which helps to improve the model training effect and training efficiency. In summary, the model training method of the embodiment of the present application can effectively balance the training effect and training efficiency of the model.

[0110] Figure 5 A block diagram of a model-based interaction device provided by an embodiment of the present application.

[0111] With reference to Figure 5The embodiment of the application provides a model-based interaction device, which can include the following modules.

[0112] The receiving module 501 is configured to receive user conversation data sent by a user terminal.

[0113] The processing module 502 is configured to process the user conversation data based on a preset dialogue model to obtain agent reply data corresponding to the user conversation data.

[0114] The sending module 503 is configured to send the agent reply data to the user terminal.

[0115] The preset dialogue model is obtained based on the model training method in any of the embodiments of the application.

[0116] In the embodiments of the application, since the preset dialogue model is trained based on the model training method in the embodiments of the application, it has strong processing capability for complex tasks and can quickly respond to simple tasks, so it can quickly and accurately give matching agent reply data for the user conversation data, improve the interaction effect, and thus help to improve the processing effect and efficiency of the dialogue interaction task.

[0117] Figure 6 A block diagram of an electronic device is provided in the embodiments of the application.

[0118] Reference Figure 6 The embodiments of the application provide an electronic device, which includes at least one processor 601, at least one memory 602, and one or more I / O interfaces 603. The processor 601, the memory 602, and the I / O interface 603 are connected to each other through a bus. The memory 602 stores one or more computer programs executable by the at least one processor 601, and the one or more computer programs are executed by the at least one processor 601 to enable the at least one processor 601 to execute any of the model training methods or model-based interaction methods described in the embodiments of the application.

[0119] The above-mentioned modules can be all or partially implemented by software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to the above-mentioned modules.

[0120] The embodiments of the present disclosure further provide a computer readable storage medium, having stored thereon a computer program, wherein the computer program, when executed by a processor, implements any one of the model training method or the model-based interaction method disclosed by the embodiments of the present disclosure. The computer readable storage medium can be a volatile or nonvolatile computer readable storage medium.

[0121] The embodiments of the present disclosure further provide a computer program product, comprising computer readable code, or a non-volatile computer readable storage medium carrying the computer readable code, when the computer readable code is run in a processor of an electronic device, the processor in the electronic device executes any one of the model training method or the model-based interaction method disclosed by the embodiments of the present disclosure.

[0122] Those of ordinary skill in the art will understand that all or some of the steps in the above-disclosed methods, the functions of the modules / units in the systems and devices can be implemented as software, firmware, hardware, or a combination thereof. In hardware implementation, the division between the modules / units referred to in the above description does not necessarily correspond to the division of physical components; for example, one physical component can have multiple functions, or one function or step can be performed by several physical components in cooperation. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on computer readable storage media, which can include computer storage media (or non-transitory media) and communication media (or transitory media).

[0123] As is well known to those of ordinary skill in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable program instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, random access memory (RAM), read only memory (ROM), erasable programmable read only memory (EPROM), static random access memory (SRAM), flash memory or other memory technology, portable compact disc read only memory (CD-ROM), digital versatile disc (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by a computer. Furthermore, it is well known to those of ordinary skill in the art that communication media typically includes computer readable program instructions, data structures, program modules or other data in modulated data signals such as carrier waves or other transport mechanisms, and can include any information delivery media.

[0124] The computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0125] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computing / processing device, partly on the user's computing / processing device, as a stand-alone software package, partly on the user's computing / processing device and partly on a remote computing / processing device or entirely on the remote computing / processing device or server. In the latter scenario, the remote computing / processing device can be connected to the user's computing / processing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computing / processing device, for example, through the Internet using an Internet Service Provider. In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0126] The computer program product described herein can be embodied in a tangible computer readable storage medium, or embodied in a software product, such as a software development kit (SDK), and the like.

[0127] The computer program product described herein can be embodied in a tangible computer readable storage medium, or embodied in a software product, such as a software development kit (SDK), and the like.

[0128] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can be a computer- readable storage medium having no data storage cycles that change state. The instructions can be executed by one or more processors of a computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process such that the instructions which execute via the one or more processors of the computer or other programmable data processing devices create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0129] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0130] The flow and block diagrams in the drawings show architectural, functional, and operational representations of possible implementations of systems, methods, and computer program products according to the present disclosure. In this regard, each block in the flow and block diagrams can represent a module, a segment, or a portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may

[0131] Example embodiments have been disclosed and although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that features, characteristics or elements described with reference to one embodiment can be used in combination with features, characteristics or elements described with reference to other embodiments unless otherwise explicitly stated. Accordingly, it will be understood that various changes in form and details can be made without departing from the scope of the disclosure as set forth in the appended claims.

Claims

1. A model training method, characterized in that: The method comprises: Evaluate the model to be processed to obtain an evaluation result of the model to be processed in the current training sub-stage; Determining a training mode for the model to be processed in a next training sub-stage according to the evaluation result and a preset evaluation threshold condition corresponding to the current training sub-stage, wherein the training mode is used to indicate a processing method for forward reasoning of the model to be processed during the training process; In the next training sub-phase, the model to be processed is trained based on the determined training mode.

2. The method according to claim 1, characterized in that The training mode includes a first training mode and a second training mode, the first training mode is a processing mode for direct reasoning output results, and the second training mode is a processing mode for reasoning output results based on reasoning chain; The type of the reasoning chain includes at least one of the following: an explicit reasoning chain, a fully implicit reasoning chain, an implicit reasoning chain based on a prompt template, and an implicit reasoning chain based on intermediate supervision.

3. The method according to claim 2, characterized in that The evaluating the model to be processed to obtain an evaluation result of the model to be processed in the current training sub-phase includes: In response to a preset start condition being met, the model to be processed is trained based on preset training sample data and in a first training mode to obtain output result data corresponding to the training sample data; An evaluation result of the model to be processed in the current training sub-phase is determined based on the training sample data and the output result data.

4. The method according to claim 2, characterized in that The evaluation result includes at least one evaluation value, and the preset evaluation threshold condition includes an expected value relationship between each evaluation value and the corresponding preset evaluation threshold; The step of determining a training mode for the model to be processed in a next training sub-stage according to the evaluation result and a preset evaluation threshold condition corresponding to the current training sub-stage includes: If the evaluation result does not meet the preset evaluation threshold condition, determining that the training mode of the to-be-processed model in the next training sub-phase is the first training mode; When the evaluation result satisfies the preset evaluation threshold condition, it is determined that the training mode of the model to be processed in the next training sub-phase is the second training mode.

5. The method according to claim 4, characterized in that After determining that the training mode of the to-be-processed model in the next training sub-phase is the second training mode, the method further includes: According to the difference between the evaluation value and the corresponding preset evaluation threshold, a target reasoning chain type is selected from the multiple reasoning chain types as the reasoning chain type adopted by the to-be-processed model in the next training sub-stage.

6. The method according to claim 4 or 5, characterized in that The evaluation result includes at least one of perplexity, accuracy, precision, recall and result error; The preset evaluation threshold conditions include at least one of the following: the perplexity is less than or equal to the preset perplexity threshold, the accuracy is greater than or equal to the preset accuracy threshold, the precision is greater than or equal to the preset precision threshold, the recall is greater than or equal to the preset recall threshold, and the result error is less than or equal to the preset result error threshold.

7. The method according to claim 4 or 5, characterized in that Each training sub-phase consists of one training round; The method further comprises: Determine the preset evaluation threshold corresponding to each training round according to the preset descent adjustment parameter, and obtain the preset evaluation threshold corresponding to each training sub-stage; The decrease adjustment parameter is used to characterize the degree to which the value of the preset evaluation threshold decreases as the number of training rounds increases.

8. The method according to claim 1, characterized in that The step of training the model to be processed based on the determined training mode in the next training sub-stage includes: Determining a loss function corresponding to the training mode; Determining a loss value of the model to be processed with respect to the loss function based on the training mode; At least part of the model parameters of the model to be processed are adjusted according to the loss value to obtain an adjusted model to be processed.

9. A model-based interaction method, characterized in that: The method comprises: Receiving user conversation data sent by the user terminal; Processing the user conversation data based on a preset speech model to obtain agent response data corresponding to the user conversation data; Sending the agent reply data to the user terminal; Wherein, the preset speech model is obtained based on the model training method described in any one of claims 1-8.

10. A model training device, characterized in that: The device comprises: An evaluation module is used to evaluate the model to be processed and obtain an evaluation result of the model to be processed in the current training sub-stage; a determination module, configured to determine a training mode for the model to be processed in the next training sub-stage based on the evaluation result and a preset evaluation threshold condition corresponding to the current training sub-stage, wherein the training mode is used to indicate a processing method for forward reasoning of the model to be processed during the training process; A training module is used to train the model to be processed based on the determined training mode in the next training sub-phase.

11. A model-based interactive device, characterized in that: The device comprises: A receiving module, configured to receive user conversation data sent by a user terminal; A processing module, configured to process the user conversation data based on a preset speech model to obtain agent response data corresponding to the user conversation data; A sending module, configured to send the agent reply data to the user terminal; Wherein, the preset speech model is obtained based on the model training method described in any one of claims 1-8.

12. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs that can be executed by the at least one processor, and the one or more computer programs are executed by the at least one processor to enable the at least one processor to execute the model training method described in any one of claims 1 to 8, or the model-based interaction method described in claim 9.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When executed by a processor, the computer program implements the model training method according to any one of claims 1 to 8, or the model-based interaction method according to claim 9.

14. A computer program product, characterized in that It includes a computer-readable code, or a non-volatile computer-readable storage medium carrying a computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the model training method as described in any one of claims 1 to 8, or the model-based interaction method as described in claim 9.