Intelligent AI scene simulation training method, system and device and medium

Through intelligent AI scenario simulation training method, multi-round dialogue and real-time evaluation technology are used to solve the problem of lack of authenticity and interactivity in traditional dialogue training methods, and a more efficient and high-quality training experience is achieved.

CN120067273AActive Publication Date: 2025-05-30SICHUAN RONGCHENG LEIMING TECH CO LTD

Patent Information

Application Number
CN202510504160.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-30
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

Traditional dialogue training methods lack authenticity and interactivity, making it difficult to restore the complexity and uncertainty in actual dialogue, resulting in poor training results and user experience.

Method used

Through intelligent AI scenario simulation training methods, the target training data is retrieved to generate the target training library, and multiple rounds of dialogue training are responded to the target user, and the user dialogue information is obtained and evaluated in real time, quality data is generated and training data is adjusted to improve the training effect.

Benefits of technology

Provide a more intuitive and effective training experience, real-time feedback and evaluation, improve the efficiency and quality of training, and improve the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067273A_ABST
    Figure CN120067273A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an intelligent AI scene simulation training method, system and device and a medium, and the method comprises the steps: calling target training data, and generating a target training library; in response to the target user, training is carried out based on target training data in the target training library, the training comprises multiple rounds of dialogues, and for each round of dialogue, the round of dialogue information of the target user is acquired; based on the round training data and the round dialogue information, generating quality data corresponding to the round dialogue information; generating a plurality of candidate training data based on the round dialogue information; based on the difficulty data and the quality data corresponding to the plurality of candidate training data, determining round training data of the next round of dialogue; and sending the round training data and the quality data of the next round of conversation to the user terminal. According to the invention, more visual and effective training experience can be provided, training results can be effectively monitored, statistically analyzed, fed back and evaluated in real time, and the training efficiency and quality are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of artificial intelligence technology, and particularly relates to an intelligent AI scenario simulation training method, system, device and medium. Background Art

[0002] In traditional dialogue training methods, training mainly relies on artificial simulation of dialogue scenarios or simple recordings, etc., lacking authenticity and interactivity, and it is difficult to restore the complexity and uncertainty in actual conversations, resulting in poor training effects and user experience.

[0003] Therefore, it is desired to propose an intelligent AI scenario simulation training method, system, device and medium, which can provide a more intuitive and effective training experience by simulating real dialogue scenarios. It can effectively monitor and statistically analyze the training results, provide real-time feedback and evaluation, improve the training efficiency and quality, and enhance the user experience. Summary of the Invention

[0004] One or more embodiments of this specification provide an intelligent AI scenario simulation training method. The method includes: retrieving target training data to generate a target training library; in response to a target user training based on the target training data in the target training library, the training includes multiple rounds of conversations, and for each round of conversation: obtaining the round conversation information of the target user, the round conversation information corresponding to round training data, and the round training data being the training data used in this round of conversation; generating quality data corresponding to the round conversation information based on the round training data and the round conversation information; generating multiple candidate training data based on the round conversation information; determining the round training data for the next round of conversation based on the difficulty data corresponding to the multiple candidate training data and the quality data; and sending the round training data for the next round of conversation and the quality data to a user terminal.

[0005] One or more embodiments of this specification provide an intelligent AI scenario simulation training system. The system includes a generation module and a training module; the training module is configured to retrieve target training data and generate a target training library; the scoring module is configured to respond to a target user's training based on the target training data in the target training library, the training including multiple rounds of conversations. For each round of conversation: obtain the round conversation information of the target user, the round conversation information corresponding to round training data, and the round training data being the training data used in this round of conversation; generate quality data corresponding to the round conversation information based on the round training data and the round conversation information; generate multiple candidate training data based on the round conversation information; determine the round training data for the next round of conversation based on the difficulty data corresponding to the multiple candidate training data and the quality data; and send the round training data for the next round of conversation and the quality data to the user terminal.

[0006] One or more embodiments of this specification provide an intelligent AI scenario simulation training device, which includes at least one processor and at least one memory; the at least one memory is used to store computer instructions; the at least one processor is used to execute at least some of the computer instructions to implement the intelligent AI scenario simulation training method described in the above embodiments.

[0007] One or more embodiments of this specification provide a computer-readable storage medium, which stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes the intelligent AI scenario simulation training method described in the above embodiments.

[0008] The beneficial effects of this specification include but are not limited to: (1) providing a more intuitive and effective training experience; (2) being able to effectively monitor and statistically analyze the training results, provide real-time feedback and evaluation, and improve the efficiency and quality of training; (3) adopting a machine learning model to more carefully evaluate the influence of multi-dimensional data on sub-quality data and improve the accuracy of determining sub-quality data. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] This specification will be further described by way of exemplary embodiments, which will be described in detail through the drawings. These embodiments are not restrictive. In these embodiments, the same numbers represent the same structures, where: Figure 1 is an exemplary schematic diagram of an intelligent AI scenario simulation training system according to some embodiments of this specification; Figure 2 is an exemplary flowchart of an intelligent AI scenario simulation training method according to some embodiments of this specification; Figure 3It is an exemplary diagram for determining quality data shown in some embodiments of this specification; Figure 4 It is an exemplary diagram of a quality assessment model shown in some embodiments of this specification.

[0010] Explanation of reference numerals: 110: Generation module; 120: Training module; 311-1, …, 311-n: Multiple dialogue sub-data; 312-1, …, 312-n: Multiple dialogue sub-features; 313: Target application scenario A; 314-1, …, 314-n: Multiple simulated dialogue information; 320-1, …, 320-n: Multiple sub-quality data; 330: Quality data; 411: Training sub-data; 412: Target application scenario B; 413: Simulated dialogue information B; 420: Quality assessment model; 421: Information generation layer; 422: Quality assessment layer; 431: Standard dialogue sub-data; 432: Standard dialogue sub-features; 441: Dialogue sub-data B; 442: Dialogue sub-features B; 450: Sub-quality data B. Detailed implementation manners

[0011] To more clearly illustrate the technical solutions of the embodiments of this specification, the accompanying drawings required for the description of the embodiments will be briefly introduced below. The accompanying drawings do not represent all implementation manners.

[0012] It should be understood that the "system", "device", "unit" and / or "module" used herein is a method for distinguishing different components, elements, parts, portions or assemblies at different levels. If other words can achieve the same purpose, the said words can be replaced by other expressions.

[0013] Unless the context clearly indicates an exceptional situation, words such as "a", "an", "one" and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list, and the method or device may also include other steps or elements.

[0014] In the embodiments of this specification, when the operations are described step by step, unless otherwise specified, the order of the steps can be adjusted, the steps can be omitted, and other steps can also be included during the operation process.

[0015] Figure 1 It is an exemplary diagram of an intelligent AI scenario simulation training system shown in some embodiments of this specification.

[0016] In some embodiments, the intelligent AI scenario simulation training system 100 includes a generation module 110 and a training module 120.

[0017] In some embodiments, the generation module may be configured to retrieve target training data and generate a target training library.

[0018] In some embodiments, the training module may be configured to, in response to a target user training based on the target training data in the target training library, obtain the round dialogue information of the target user; generate quality data corresponding to the round dialogue information based on the round training data and the round dialogue information; generate a plurality of candidate training data based on the round dialogue information; determine the round training data for the next round of dialogue based on the difficulty data and quality data corresponding to the plurality of candidate training data; and send the round training data and quality data for the next round of dialogue to the user terminal.

[0019] In some embodiments, for a single dialogue sub - data among a plurality of dialogue sub - data, the training module is further configured to: generate sub - quality data corresponding to the dialogue sub - data based on the dialogue sub - data, the dialogue sub - features corresponding to the dialogue sub - data, the target application scenario, and the simulated dialogue information; and determine the quality data based on the plurality of sub - quality data corresponding to the plurality of dialogue sub - data.

[0020] In some embodiments, the training module is further configured to: generate sub - quality data corresponding to the dialogue sub - data through a quality evaluation model based on the dialogue sub - data, the dialogue sub - features corresponding to the dialogue sub - data, the target application scenario, and the simulated dialogue information.

[0021] The user terminal refers to the terminal device used by the target user, for example, a mobile phone, a tablet computer, etc. In some embodiments, the intelligent AI scenario simulation training system is communicatively connected to the user terminal.

[0022] In some embodiments, the intelligent AI scenario simulation training system may further include a processor and a storage device. The processor is configured to process data from at least one module of the intelligent AI scenario simulation training system 100 or an external data source. The processor includes a central processing unit (CPU), an application - specific integrated circuit (ASIC), a controller, a microcontroller unit, a microprocessor, etc., or any combination thereof.

[0023] The storage device is configured to store data, instructions, and / or information related to the intelligent AI scenario simulation training system 100. For example, the storage device stores target training data and the target training library, etc. In some embodiments, the storage device may be integrated into the processor or may be provided in the cloud, etc.

[0024] In some embodiments, the generation module and the training module may be implemented by the processor.

[0025] For more descriptions of the above content, see Figures 2 - 4 and its related descriptions.

[0026] The intelligent AI scenario simulation training system can generate a target training library through a generation module, and can conduct targeted training for different users. Through the training module, it can train the target users and provide real-time feedback and evaluation of the training results, improving the efficiency and quality of training.

[0027] It should be understood that Figure 1 The intelligent AI scenario simulation training system and its modules shown can be implemented in various ways. It should be noted that the above description of the intelligent AI scenario simulation training system and its modules is only for convenience of description and does not limit this specification to the scope of the examples given. It can be understood that for those skilled in the art, after understanding the principle of the system, they may, without departing from this principle, make any combination of the various modules, or form a subsystem and connect it with other modules. In some embodiments, Figure 1 the generation module 110 and the training module 120 disclosed in [reference] can be different modules in a system, or a single module can implement the functions of the above two or more modules. For example, each module can share a storage module, or each module can have its own storage module respectively. Such variations are all within the scope of protection of this specification.

[0028] Figure 2 is an exemplary flowchart of an intelligent AI scenario simulation training method according to some embodiments of this specification. In some embodiments, process 200 is executed by a processor. As Figure 2 shown, process 200 includes the following steps: Step 210, retrieve target training data and generate a target training library.

[0029] Target training data refers to the data used for training. In some embodiments, training includes dialogue training, etc. Dialogue training means that the processor issues a question and the target user answers it.

[0030] In some embodiments, the target training data includes multiple training sub-data, etc. Training sub-data refers to a dialogue statement required for conducting dialogue training for the target user.

[0031] The target user refers to the user who needs to receive dialogue training. For example, at least one of a customer service representative, a teller, a salesperson, an e-commerce operator, etc.

[0032] The target training library refers to a database used to store the target training data.

[0033] In some embodiments, the processor generates the target training library through various methods. For example, the processor obtains the target training data corresponding to the target user through a storage device, records the target training data in the target training library corresponding to the target user, and stores the target training library in the storage device for use in training the target user.

[0034] In some embodiments, the target training data corresponding to different target users can be preset according to experience and stored in a storage device. One target training library can correspond to one or more target users.

[0035] In some embodiments, the processor can determine reference difficulty data based on multiple historical training results of a target user's multiple historical trainings, and retrieve target training data based on the reference difficulty data.

[0036] The historical training result refers to the comprehensive training result corresponding to the historical training. For the description of the comprehensive training result, see the following text and its related descriptions.

[0037] The reference difficulty data refers to the reference value of the difficulty data of the training sub-data to be retrieved.

[0038] In some embodiments, the processor calculates the mean value of multiple historical training results of a target user's multiple historical trainings. If the mean value meets a preset condition, the processor reduces the initial reference difficulty by a preset adjustment amount to obtain the reference difficulty data. If the mean value does not meet the preset condition, the processor determines the initial reference difficulty as the reference difficulty data.

[0039] The preset condition includes that the mean value is not greater than the quality threshold. The quality threshold and the preset adjustment amount are preset based on historical experience.

[0040] The initial reference difficulty refers to the initially set reference difficulty data. In some embodiments, the initial reference difficulty is preset based on historical experience.

[0041] In some embodiments, the preset adjustment amount can be related to the training frequency of the target user. For example, the preset adjustment amount is negatively correlated with the training frequency of the target user.

[0042] The training frequency refers to the frequency at which the target user conducts training. The greater the training frequency, the shorter the interval between the target user's multiple trainings, indicating that the target user's own level has not changed significantly and there is no need to significantly adjust the original reference difficulty data, and thus the preset adjustment amount is smaller.

[0043] By the training frequency of the target user, a preset adjustment amount that is more suitable for the target user can be adjusted, avoiding that when the target user trains frequently, the preset adjustment amount is too large, resulting in the target user undergoing multiple trainings with a large difficulty span in a short period of time and reducing the effectiveness of training.

[0044] In some embodiments, the processor may eliminate multiple training sub-data in the target training data that do not meet the preset retrieval conditions, and generate a target training library based on the target training data after elimination. The preset retrieval conditions include that the difference between the difficulty data of the training sub-data and the reference difficulty data is less than the first difference threshold. The first difference threshold is preset based on historical experience.

[0045] According to multiple historical training results of the target user, evaluate the target user's own level, and then dynamically adjust the retrieved training sub-data, so that the target training data better conforms to the current ability of the target user and improves the training effect.

[0046] Step 220, in response to the target user training based on the target training data in the target training library, the training includes multiple rounds of conversations, and steps 221-step 225 are executed for each round of conversation.

[0047] It can be understood that the training is equivalent to simulating a real conversation scenario, and the two parties in the real conversation scenario will have multiple rounds of conversations to complete the communication.

[0048] Step 221, obtain the round conversation information of the target user.

[0049] The round conversation information refers to the content answered by the target user in this round of conversation. In some embodiments, the round conversation information includes one or more conversation sub-data and the corresponding conversation sub-features of each conversation sub-data, etc. The conversation sub-data refers to a statement in which the target user answers the training sub-data. The conversation sub-feature refers to the relevant information and features of the conversation sub-data.

[0050] In some embodiments, the conversation sub-features include the answer waiting time, answer speech rate, and answer volume corresponding to the conversation sub-data, etc.

[0051] In some embodiments, the processor can record the conversation sub-data, answer waiting time, answer speech rate, and answer volume of the target user in real time during the conversation to obtain the round conversation information.

[0052] Step 222, generate quality data corresponding to the round conversation information based on the round training data and the round conversation information.

[0053] In some embodiments, the round training data refers to the training data used in this round of conversation. For example, the training sub-data used by the processor in this round of conversation.

[0054] In some embodiments, the turn-based dialogue information corresponds to turn-based training data. It can be understood that there is a certain word order among multiple training sub-data in the target training data. When performing training, the processor selects a training sub-data from the target training library according to the word order as the turn-based training data. Correspondingly, the target user makes an answer based on the turn-based training data, obtaining one or more dialogue sub-data.

[0055] Exemplarily, the target training data is used to train the target user for after-sales service. The after-sales service may include product situation description, refund or replacement of products, and service scoring, etc. Then the processor can set the training sub-data used in the dialogue according to the word order of product situation description, refund or replacement of products, and service scoring.

[0056] In some embodiments, after generating the target training library, the processor can construct a response graph based on the target training data in the target training library.

[0057] The response graph refers to a graph structure that reflects the association relationship between the target training data and the corresponding answers. The graph structure is a data structure composed of nodes and edges. The edges connect the nodes, and the nodes and edges can have features.

[0058] In some embodiments, the response graph includes question nodes and answer nodes.

[0059] The question node represents a training sub-data. The node features of the question node include difficulty data, etc. The difficulty data refers to the data used to characterize the difficulty level of the training sub-data. In some embodiments, the difficulty data is preset based on historical experience.

[0060] The answer node refers to the answer statement corresponding to the training sub-data. The node features of the answer node include the answer waiting time, answer speech rate, and answer volume of the answer statement, etc. The answer waiting time refers to the time interval between receiving the training sub-data and making an answer.

[0061] In some embodiments, the processor queries multiple question-and-answer statements corresponding to the training sub-data in the preset question-and-answer table, uses the multiple question-and-answer statements as the answer nodes corresponding to the training sub-data, and uses the answer waiting time, answer speech rate, and answer volume corresponding to each question-and-answer statement as the node features of the answer node. The preset question-and-answer table is preset based on historical experience, includes multiple training sub-data and multiple answer statements corresponding to different training sub-data, and also includes the answer waiting time, answer speech rate, and answer volume corresponding to each answer statement.

[0062] In some embodiments, there is an edge pointing from the question node to the corresponding answer node between the question node and the corresponding answer node.

[0063] In some embodiments, since there is a certain word order among multiple training sub-data in the target training data, after answering a training sub-data, the processor can select the subsequent training sub-data to be presented based on the answer. Therefore, some answer nodes also have out-edges, and the out-edges point to one or more question nodes used in subsequent questions.

[0064] Quality data refers to data used to characterize the answer quality of the target user in the current round of conversation. In some embodiments, the processor determines the initial quality data based on the similarity between one or more dialogue sub-data in the round of conversation information and the corresponding multiple answer nodes. For example, the processor concatenates one or more dialogue sub-data in the round of conversation information into a long sentence, calculates the average similarity between the long sentence and the corresponding multiple answer nodes, and uses the average similarity as the initial quality data. The multiple answer nodes corresponding to the long sentence refer to the multiple answer nodes pointed to by the question nodes representing the training sub-data corresponding to the round of conversation information in the answer graph.

[0065] In some embodiments, the processor can calculate the similarity between the long sentence and the corresponding answer nodes through methods such as the BERT model, SentenceBERT model, and SimCSE model.

[0066] In some embodiments, the processor can also adjust the initial quality data based on the mean of the node features of multiple dialogue sub-features in the round of conversation information and the corresponding multiple answer nodes to obtain the final quality data. For example, the processor calculates the mean of each data item in the multiple dialogue sub-features, calculates the ratio of the mean of each data item to the corresponding data in the mean of the node features, and adjusts the initial quality data based on the multiple ratios and the mean of the node features through a preset algorithm. Exemplarily, the preset algorithm is shown in the following formula (1): B = A × [1 - (k1 × C1 / D1 + …… + kn × Cn / Dn) / (k1 × D1 + …… + kn × Dn)] (1) Wherein, A represents the initial quality data, B represents the final quality data, C1, …, Cn represent the mean of the first item of data to the nth item of data in the multiple dialogue sub-features, D1, …, Dn represent the first item of data to the nth item of data in the mean of the node features of the multiple answer nodes, and k1, …, kn represent the weight of the first item of data to the nth item of data. The weights of different data can be preset according to experience. n is the number of data items included in the dialogue sub-features. For example, 3.

[0067] In some embodiments, the processor can also determine the quality data based on the multiple sub-quality data corresponding to the multiple dialogue sub-data. For more details, see Figure 3 and its related descriptions.

[0068] Step 223: Generate multiple candidate training data based on the round-based dialogue information.

[0069] Candidate training data refers to the training sub-data to be determined for use in the next round of dialogue. It can be understood that there may be multiple answers to a training sub-data, and the subsequent question statements corresponding to each answer are different. Therefore, the processor can select some training sub-data as candidate training data from multiple training sub-data (i.e., subsequent question statements) for use in the next round of dialogue based on the current round-based dialogue information (i.e., the current one or more answer statements), and determine the training sub-data finally used in the next round of dialogue from the candidate training data.

[0070] In some embodiments, the processor can obtain candidate training data in multiple ways. For example, the processor determines the answer nodes that meet the candidate conditions based on the similarity between the round-based dialogue information and the corresponding multiple answer nodes, and uses the multiple question nodes pointed to by the out-edges of the answer nodes as multiple candidate training data. The candidate conditions include that the similarity between the round-based dialogue information and the answer nodes is the highest.

[0071] In some embodiments, the processor can also determine the standard difficulty data based on the change characteristics corresponding to multiple quality data and the difficulty data corresponding to the round-based training data, and generate multiple candidate training data based on the standard difficulty data and the round-based dialogue information.

[0072] It should be noted that after at least one round of dialogue, the processor can determine the change characteristics corresponding to multiple quality data and then generate multiple candidate training data.

[0073] The change characteristics refer to the data used to characterize the change situation of the quality data. In some embodiments, the change characteristics include change curves, etc. The x-axis of the change curve is the training duration, and the y-axis is the quality data.

[0074] In some embodiments, the processor can record multiple training time points in multiple rounds of dialogue, calculate the quality data corresponding to the training time points, and construct a change curve based on the multiple training time points and the quality data corresponding to different training time points. Among them, the training time point refers to the time point at the end of a round of dialogue.

[0075] The standard difficulty data refers to the reference value of the difficulty data of the candidate sub-data. In some embodiments, the processor may take the first derivative of multiple sampling points on the change curve. If the first derivatives of N consecutive sampling points from the current time point satisfy the first adjustment condition, the difficulty data of the training sub-data for this round of conversation is increased by a difficulty adjustment amount as the standard difficulty data. If the first derivatives of N consecutive sampling points from the current time point satisfy the second adjustment condition, the difficulty data of the training sub-data for this round of conversation is decreased by a difficulty adjustment amount as the standard difficulty data. Among them, the difficulty adjustment amount and multiple sampling points can be preset based on historical experience. The difficulty data of the training sub-data is obtained through the response graph. N is a preset quantity.

[0076] The first adjustment condition includes that the first derivatives are all greater than 0, and the second adjustment condition includes that the first derivatives are all less than 0. If the first derivatives of N consecutive sampling points from the current time point do not satisfy the first adjustment condition and the second adjustment condition, the processor takes the difficulty data of the training sub-data for this round of conversation as the standard difficulty data.

[0077] In some embodiments, the value of N can also be related to the number of reversals of the first derivatives of multiple sampling points on the change curve. For example, the value of N is positively correlated with the number of reversals. The number of reversals refers to the number of times the first derivative changes from a positive number to a negative number. The larger the number of reversals, the higher the difficulty of judging whether the quality data is in an upward trend or a downward trend. At this time, a larger value of N is required to improve the accuracy of judging whether the quality data is in an upward trend or a downward trend.

[0078] In some embodiments, the processor may select answer nodes that meet the candidate conditions in the response graph, determine multiple question nodes pointed to by the out-edges of the answer nodes, and take the multiple question nodes whose difference between the difficulty data and the standard difficulty data is not greater than the second difference threshold as multiple candidate training data. Among them, the second difference threshold can be preset according to experience.

[0079] Based on the change situation of the quality data of the previous rounds of training and the difficulty data of the current training sub-data, a more applicable standard difficulty data can be determined, and then multiple candidate training data that are more suitable for the current target user level can be generated, improving the training effect.

[0080] Step 224, determine the round training data for the next round of conversation based on the difficulty data and quality data corresponding to multiple candidate training data.

[0081] In some embodiments, the processor may calculate the ratio of the quality data of the current round of conversation to the qualified value, and calculate the average value of the difficulty data of multiple candidate training data. The processor multiplies the obtained ratio by the average value of the difficulty data to obtain the comprehensive difficulty data, and screens the candidate training data with the smallest absolute value of the difference between the difficulty data and the comprehensive difficulty data among the multiple candidate training data, and uses this candidate training data as the round training data for the next round of conversation. Among them, the qualified value can be preset according to experience.

[0082] Step 225, send the round training data and quality data of the next round of conversation to the user terminal.

[0083] In some embodiments, the processor may send the round training data of the next round of conversation and the quality data of the current round of conversation to the user terminal for the target user to view and conduct the next round of conversation. The processor can complete the training of multiple rounds of conversations by sending round training data to the target user multiple times, recording the corresponding round conversation information, and determining the quality data.

[0084] The comprehensive training result refers to the data used to characterize the quality of the target user's completion of the training.

[0085] In some embodiments, in response to the target user's completion of the training, the processor may generate a comprehensive training result corresponding to the training based on the multiple quality data corresponding to multiple rounds of conversations. For example, the processor may determine the average value of the multiple quality data as the comprehensive training result corresponding to the training. For another example, the processor may determine the weighted sum value of the multiple quality data as the comprehensive training result corresponding to the training. Among them, the weights of different rounds of conversations are preset based on historical experience.

[0086] By synthesizing the quality data of multiple rounds of conversations of the target user and generating an intuitive comprehensive training result after the training is completed, it is beneficial to comprehensively evaluate the training effect, provide growth feedback for the target user, and is beneficial to determining more target training data that is more in line with the target user's level in subsequent training.

[0087] By constructing a corresponding target training library for the target user and conducting training, it is possible to provide more intuitive and effective training for the target user by simulating real conversation scenarios, and at the same time, it is possible to effectively monitor and statistically analyze the training results, make real-time feedback and evaluation, and improve the efficiency and quality of the training.

[0088] It should be noted that the above description of process 200 is only for illustration and explanation, and does not limit the scope of application of this specification. For those skilled in the art, various modifications and changes can be made to the process under the guidance of this specification. However, these modifications and changes are still within the scope of this specification.

[0089] Figure 3An exemplary schematic diagram for determining quality data as shown in some embodiments of this specification.

[0090] In some embodiments, the round-based dialogue information includes multiple dialogue sub-data and multiple dialogue sub-features. The target training data further includes a target application scenario and simulated dialogue information. Among them, one training sub-data can correspond to one target application scenario and simulated dialogue information. For more descriptions about the round-based dialogue information, see Figure 2 And its related descriptions.

[0091] In some embodiments, for a single dialogue sub-data among the multiple dialogue sub-data, the processor can generate sub-quality data corresponding to the dialogue sub-data based on the dialogue sub-data (such as dialogue sub-data 311-1, …, dialogue sub-data 311-n, where n is the number of dialogue sub-data), the dialogue sub-features corresponding to the dialogue sub-data (such as dialogue sub-features 312-1, …, dialogue sub-features 312-n), the target application scenario (such as target application scenario 313), and the simulated dialogue information (such as simulated dialogue information 314-1, …, simulated dialogue information 314-n). The processor determines quality data 330 based on the multiple sub-quality data corresponding to the multiple dialogue sub-data (such as sub-quality data 320-1, …, sub-quality data 320-n). For descriptions about the target training data, dialogue sub-data, dialogue sub-features, and quality data, see Figure 2 And its related descriptions.

[0092] The target application scenario refers to the application scenario corresponding to the training sub-data. In some embodiments, the target application scenario includes the dialogue environment, mood state, and / or customer portrait, etc., of the simulated customer corresponding to the training sub-data. Among them, the customer refers to the person who simulates having a conversation with the target user. One target application scenario corresponds to one or more training sub-data.

[0093] In some embodiments, the dialogue environment includes at home or at the company, etc. The mood state includes good or average, etc. The customer portrait refers to the information used to describe the customer. The customer portrait includes the customer's gender, age, occupation, etc.

[0094] In some embodiments, the processor can pre-store multiple different target application scenarios. Before sending the training sub-data to the user terminal of the target user, randomly select one target application scenario to integrate into the training sub-data, and send the adjusted training sub-data to the user terminal of the target user for training.

[0095] In some embodiments, integrating the target application scenario into the training sub-data includes: the processor integrating one or more question keywords corresponding to the target application scenario into the training sub-data. In some embodiments, the processor obtains the question keywords corresponding to the target application scenario through a preset scenario table, etc.

[0096] In some embodiments, the preset scenario table is pre-set based on historical experience and includes multiple target application scenarios and one or more question keywords corresponding to different target application scenarios. One or more question keywords corresponding to a target application scenario are determined by technicians based on experience.

[0097] The simulated dialogue information refers to the information related to the simulated answers of the training sub-data. In some embodiments, the simulated dialogue information includes one or more simulated sub-data and the simulated sub-features corresponding to different simulated sub-data. The simulated sub-data refers to the simulated dialogue sub-data. The simulated sub-feature refers to the dialogue sub-feature corresponding to the simulated sub-data.

[0098] One training sub-data corresponds to one or more simulated sub-data.

[0099] In some embodiments, the processor generates the simulated dialogue information in multiple ways. For example, the processor can obtain the simulated speech segments input by the technician, split the simulated speech segments into multiple more colloquial or easy-to-read simulated sentences as multiple simulated sub-data, and determine the simulated sub-features corresponding to the simulated sub-data based on the waiting time, speech rate, and volume of the simulated speech segments input by voice.

[0100] In some embodiments, the processor can also generate the simulated dialogue information based on the historical training data and historical quality scores of the target user.

[0101] The historical training data refers to the relevant data in the historical training process. In some embodiments, the historical training data includes multiple historical training sub-data used by the target user and the corresponding multiple historical round dialogue information.

[0102] The historical quality data refers to the quality data corresponding to the historical dialogue sub-data. In some embodiments, the historical quality data includes the quality data corresponding to the multiple historical round dialogue information in the historical training data.

[0103] In some embodiments, the processor obtains the historical training data and historical quality data from the storage device.

[0104] In some embodiments, the processor can insert the dialogue defects included in the historical round dialogue information into the initial simulated dialogue information based on the obviousness of the defects and the number of defects to generate the simulated dialogue information.

[0105] The dialogue defect refers to the defect in the answer of the target user. For example, the dialogue defects include inappropriate word use, too slow speech rate, too fast speech rate, semantic repetition, logical contradiction, etc. The dialogue defects included in the historical round dialogue information are determined by manual annotation of the historical round dialogue information or automatically detected by the processor.

[0106] The defect obviousness refers to the data used to characterize the obviousness of the dialogue defect. The defect obviousness can be represented by means of a numerical value, etc. The larger the numerical value, the greater the defect obviousness.

[0107] In some embodiments, the processor calculates the mean value of multiple quality data in the historical quality data, and determines the defect obviousness and the number of defects by querying a preset defect table.

[0108] In some embodiments, the preset defect table is pre-set based on historical experience, and includes the mean values of multiple quality data, as well as the defect obviousness and the number of defects corresponding to different mean values of quality data. Among them, the defect obviousness and the number of defects can be determined by manually annotating the historical round of dialogue information corresponding to the mean value of quality data.

[0109] In some embodiments, for each historical training sub-data among multiple historical training sub-data, the processor can also, based on the historical training sub-data, find a question node corresponding to the historical training sub-data in the response graph. The processor randomly selects a target application scenario, incorporates one or more question keywords corresponding to the target application scenario into multiple answer nodes corresponding to the question node, and generates initial simulated dialogue information based on the multiple answer nodes and the corresponding node features. For more descriptions of the response graph, see Figure 2 and its related descriptions.

[0110] In some embodiments, the processor inserts the same number of dialogue defects included in the historical round of dialogue information as the number of defects into the initial simulated dialogue information based on the number of defects. The processor can also insert corresponding dialogue defects into the initial simulated dialogue information based on the defect obviousness. For example, if the defect obviousness is large, the dialogue defects are improper word use, logical contradictions, etc.; if the defect obviousness is low, the dialogue defects are too slow speech rate, too fast speech rate, etc.

[0111] By aggregating the historical training data and historical quality data of the target user and simulating the dialogue defects of the target user, realistic simulated dialogue information can be generated, which is beneficial to more accurately evaluating the training quality of the target user subsequently.

[0112] Sub-quality data refers to the quality data corresponding to the dialogue sub-data.

[0113] In some embodiments, the processor obtains sub-quality data in multiple ways. For example, the processor adjusts the initial sub-quality data based on the scene fitting degree between the dialogue sub-data and the target application scenario and the defect compensation degree of the dialogue sub-data for the simulated dialogue information to obtain the sub-quality data.

[0114] In some embodiments, the processor calculates the similarity between the dialogue sub-data and the corresponding multiple answer nodes, and takes the highest similarity as the initial sub-quality data of the dialogue sub-data. For the description of calculating similarity, seeFigure 2 and its related descriptions.

[0115] The scene fitness is used to characterize the degree of fit between the dialogue sub - data and the target application scenario. In some embodiments, the processor calculates the overlap degree between the dialogue sub - data and one or more answer keywords corresponding to the target application scenario, and determines the obtained overlap degree as the scene fitness. Among them, the processor determines the ratio of the number of data of answer keywords included in the dialogue sub - data to the total number of answer keywords corresponding to the target application scenario as the overlap degree.

[0116] The method for determining the answer keywords is similar to the method for determining the question keywords.

[0117] The defect compensation degree is used to characterize the degree of compensation of the dialogue sub - data for the simulated dialogue information. In some embodiments, the processor compares the dialogue sub - data, the dialogue sub - features with the simulated dialogue information, and determines the ratio of the number of dialogue defects compensated by the dialogue sub - data or the dialogue sub - features in the simulated dialogue information to the total number of dialogue defects in the simulated dialogue information as the defect compensation degree. Exemplarily, if the simulated dialogue information lacks the term "working day" and the dialogue sub - data contains "working day", it is determined that the dialogue sub - data compensates for one defect in the simulated dialogue information. If the answer speed in the simulated dialogue information is relatively fast and the answer speed in the dialogue sub - features is moderate, it is determined that the dialogue sub - features compensate for one defect in the simulated dialogue information.

[0118] In some embodiments, the processor can calculate the product of the scene fitness and the defect compensation degree, multiply the obtained product plus 1 by the initial sub - quality data to obtain the sub - quality data.

[0119] In some embodiments, the processor can also generate the sub - quality data corresponding to the dialogue sub - data through a quality evaluation model. For the description of this part, see Figure 4 and its related content.

[0120] In some embodiments, the processor can perform weighted summation on multiple sub - quality data corresponding to multiple dialogue sub - data to obtain the quality data. Among them, the weights of different dialogue sub - data are preset based on historical experience.

[0121] By deeply analyzing the dialogue sub - data, the target application scenario, and the simulated dialogue information, more accurate and comprehensive sub - quality data can be determined, and then more accurate quality data can be determined.

[0122] Figure 4 is an exemplary schematic diagram of the quality evaluation model shown in some embodiments of this specification.

[0123] In some embodiments, the processor may generate sub-quality data corresponding to the dialogue sub-data through a quality assessment model based on the dialogue sub-data, the dialogue sub-features corresponding to the dialogue sub-data, the target application scenario, the simulated dialogue information, and the target training data corresponding to the dialogue sub-data (i.e., the training sub-data corresponding to the dialogue sub-data). For the descriptions of the dialogue sub-data, the dialogue sub-features corresponding to the dialogue sub-data, the target application scenario, the simulated dialogue information, and the training sub-data, see Figures 2 - 3 and its related descriptions.

[0124] The quality assessment model 420 refers to a model for evaluating sub-quality data. In some embodiments, the quality assessment model may be a machine learning model. For example, the quality assessment model is one or a combination of a deep neural network (DNN) model or a custom model, etc.

[0125] In some embodiments, the input of the quality assessment model includes the training sub-data, the target application scenario, the simulated dialogue information, the dialogue sub-data, and the dialogue sub-features, and the output includes the sub-quality data corresponding to the dialogue sub-data.

[0126] In some embodiments, the quality assessment model can be obtained by training with the training samples and the corresponding labels in the training sample set. The training samples include sample training sub-data, sample target application scenarios, sample simulated dialogue information, sample dialogue sub-data, and sample dialogue sub-features. The labels corresponding to the training samples include the actual sub-quality data of the sample dialogue sub-data. In some embodiments, the training samples are obtained based on historical data. The labels are pre-annotated and determined by technicians.

[0127] In some embodiments, the quality assessment model includes an information generation layer 421 and a quality assessment layer 422. For example, the information generation layer and the quality assessment layer may include one or a combination of a deep neural network (DNN) model or a custom model, etc.

[0128] In some embodiments, the input of the information generation layer 421 may include the training sub-data 411, the target application scenario (such as the target application scenario B 412), and the simulated dialogue information (such as the simulated dialogue information B 413), and the output may include the standard dialogue sub-data 431 and the standard dialogue sub-features 432. The standard dialogue sub-data refers to the standard answer generated by the information generation layer. The standard dialogue sub-features refer to the dialogue sub-features corresponding to the standard answer generated by the information generation layer.

[0129] In some embodiments, the inputs of the quality assessment layer 422 may include standard dialogue sub-data 431, standard dialogue sub-features 432, dialogue sub-data (such as dialogue sub-data B 441), and dialogue sub-features (such as dialogue sub-features B 442), and the outputs may include sub-quality data corresponding to the dialogue sub-data (such as sub-quality data B 450).

[0130] In some embodiments, the information generation layer and the quality assessment layer may be obtained by training with the training samples and the corresponding labels in the training sample set respectively.

[0131] In some embodiments, the training sample set includes first training samples for training the information generation layer and first labels corresponding to the first training samples. The first training samples include sample training sub-data, sample target application scenarios, and sample simulated dialogue information. The first labels include standard dialogue sub-data and standard dialogue sub-features corresponding to the sample training sub-data.

[0132] In some embodiments, the first training samples are obtained based on historical data. The processor may select multiple historical dialogue sub-data corresponding to the sample training sub-data, calculate the scene fitness degree of each historical dialogue sub-data with the sample target application scenario and the defect compensation degree for the sample simulated dialogue information, and use the historical dialogue sub-data with the highest weighted sum of the defect compensation degree and the scene fitness degree as the standard dialogue sub-data corresponding to the sample training sub-data, and use the historical dialogue sub-features corresponding to the historical dialogue sub-data as the standard dialogue sub-features corresponding to the sample training sub-data. Among them, the weights of the scene fitness degree and the defect compensation degree are preset based on historical experience. For the description of the scene fitness degree and the defect compensation degree, see Figure 3 and its related descriptions.

[0133] In some embodiments, the training sample set includes second training samples for training the quality assessment layer and second labels corresponding to the second training samples. The second training samples include sample standard dialogue sub-data and sample standard dialogue sub-features output by the information generation layer, sample dialogue sub-data, and sample dialogue sub-features. The second labels include the actual sub-quality data of the sample dialogue sub-data. In some embodiments, the second training samples are obtained based on historical data and the information generation layer. The second labels are determined by pre-annotation by technicians.

[0134] In some embodiments, one first training sample corresponds to one second training sample.

[0135] In some embodiments, the information generation layer and the quality evaluation layer can be trained through the following joint training method: input a plurality of first training samples with a first label into the initial information generation layer, and input the output of the initial information generation layer and a second training sample with a second label into the initial quality evaluation layer. A loss function is constructed based on the second label and the prediction result of the initial quality evaluation layer, and the initial information generation layer and the initial quality evaluation layer are synchronously updated iteratively based on the loss function. When the loss function meets the preset iteration condition, the quality evaluation model training is completed. Among them, the preset iteration condition can be that the loss function converges, the number of iterations reaches a set value, etc.

[0136] In some embodiments, the labels of multiple training samples corresponding to different application scenario classifications in the training sample set meet the preset training conditions.

[0137] Application scenario classification refers to the classification of the target application scenario. In some embodiments, the application scenario classification can be preset based on historical experience. For example, for multiple target application scenarios, the processor can divide the multiple target application scenarios into different application scenario classifications based on at least one of the conversation environment, mood state, and / or customer portrait, etc. Exemplarily, the processor divides the multiple target application scenarios into three categories: at home, at the company, and on the street based on the conversation environment.

[0138] The training sample corresponding to the application scenario classification refers to the first training sample where the target application scenario corresponding to the application scenario classification is located.

[0139] The preset training condition includes that the standard deviation of the second label of the second training sample corresponding to the first training sample is greater than the preset distribution threshold. In some embodiments, the preset distribution threshold is related to the historical quality data of the target user. For example, the greater the range of multiple quality data in the historical quality data, the greater the preset score threshold. For the description of the historical quality data, see Figure 3 and its related descriptions.

[0140] Since the greater the range of the quality data of the target user, it indicates that the quality of the target user's answers may fluctuate greatly, and the accuracy of the output of the quality evaluation model will be correspondingly reduced. By introducing training samples with large differences under different application scenario classifications, the generalization performance and adaptability of the model can be enhanced, thereby optimizing the training effect and improving the accuracy of determining the sub-quality data.

[0141] By adopting a machine learning model, the influence of multi-dimensional data on the sub-quality data is evaluated more carefully, and the accuracy of determining the sub-quality data is improved. At the same time, through the hierarchical structure, targeted processing can be performed on different input data, thereby improving the accuracy of the final model output. Joint training solves problems such as insufficient training samples and difficult-to-obtain labels.

[0142] In some embodiments of this specification, an intelligent AI scenario simulation training device is further provided. The device includes at least one processor and at least one memory. The at least one memory is used to store computer instructions. The at least one processor is used to execute at least some of the computer instructions to implement the intelligent AI scenario simulation training method described in any one of the above embodiments.

[0143] In some embodiments of this specification, a computer-readable storage medium is further provided. The storage medium stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes the intelligent AI scenario simulation training method described in any one of the above embodiments.

[0144] In addition, certain features, structures, or characteristics in one or more embodiments of this specification may be appropriately combined.

[0145] In some embodiments, numbers are used to describe components and attribute quantities. It should be understood that such numbers used for embodiment descriptions are modified by the modifiers "about", "approximately", or "substantially" in some examples. Unless otherwise stated, "about", "approximately", or "substantially" indicate that the stated number allows a ±20% variation.

[0146] If there are inconsistencies or conflicts between the descriptions, definitions, and / or uses of terms in the materials cited in this specification and the content described in this specification, the descriptions, definitions, and / or uses of terms in this specification shall prevail.

Claims

1. An intelligent AI scene simulation training method, characterized in that: The method is executed by a processor, and the method includes: Retrieve target training data and generate a target training library; In response to a target user performing training based on target training data in the target training library, the training includes multiple rounds of dialogues, and for each round of dialogue: Acquire round dialogue information of the target user, where the round dialogue information corresponds to round training data, and the round training data is training data used in the current round dialogue; Based on the round training data and the round dialogue information, generating quality data corresponding to the round dialogue information; Based on the round-by-round dialogue information, generating a plurality of candidate training data; Determining round training data for a next round of dialogue based on the difficulty data corresponding to the plurality of candidate training data and the quality data; The round training data of the next round of dialogue and the quality data are sent to the user terminal.

2. The method according to claim 1, characterized in that The target training data includes a target application scenario and simulated dialogue information, the turn dialogue information includes a plurality of dialogue sub-data and a plurality of dialogue sub-features, and the generating of quality data corresponding to the turn dialogue information based on the turn training data and the turn dialogue information includes: For a single dialogue sub-data among the plurality of dialogue sub-data, generating sub-quality data corresponding to the dialogue sub-data based on the dialogue sub-data, dialogue sub-features corresponding to the dialogue sub-data, the target application scenario and the simulated dialogue information; The quality data is determined based on a plurality of the sub-quality data corresponding to the plurality of the conversation sub-data.

3. The method according to claim 2, characterized in that For a single dialogue sub-data among the plurality of dialogue sub-data, generating sub-quality data corresponding to the dialogue sub-data based on the dialogue sub-data, dialogue sub-features corresponding to the dialogue sub-data, the target application scenario and the simulated dialogue information comprises: Based on the conversation sub-data, conversation sub-features corresponding to the conversation sub-data, the target application scenario, the simulated conversation information and the target training data corresponding to the conversation sub-data, sub-quality data corresponding to the conversation sub-data is generated through a quality assessment model, and the quality assessment model is a machine learning model.

4. The method according to claim 1, characterized in that The step of retrieving target training data and generating a target training library comprises: Determining reference difficulty data based on multiple historical training results of multiple historical trainings of the target user and preset adjustment amounts; Based on the reference difficulty data, the target training data is retrieved.

5. The method according to claim 1, characterized in that The method further comprises: In response to the target user completing the training, a comprehensive training result corresponding to the training is generated based on the plurality of quality data corresponding to the plurality of rounds of conversations.

6. An intelligent AI scene simulation training system, characterized in that: The system includes a generation module and a training module; The generating module is configured to retrieve target training data and generate a target training library; The training module is configured to perform training based on target training data in the target training library in response to a target user, wherein the training includes multiple rounds of dialogues, and for each round of dialogue: Acquire round dialogue information of the target user, where the round dialogue information corresponds to round training data, and the round training data is training data used in the current round dialogue; Based on the round training data and the round dialogue information, generating quality data corresponding to the round dialogue information; Based on the round-by-round dialogue information, generating a plurality of candidate training data; Determining round training data for a next round of dialogue based on the difficulty data corresponding to the plurality of candidate training data and the quality data; The round training data of the next round of dialogue and the quality data are sent to the user terminal.

7. The system according to claim 6, characterized in that The target training data includes a target application scenario and simulated dialogue information, the turn dialogue information includes a plurality of dialogue sub-data and a plurality of dialogue sub-features, and the training module is further configured as follows: For a single dialogue sub-data among the plurality of dialogue sub-data, generating sub-quality data corresponding to the dialogue sub-data based on the dialogue sub-data, dialogue sub-features corresponding to the dialogue sub-data, the target application scenario and the simulated dialogue information; The quality data is determined based on a plurality of the sub-quality data corresponding to the plurality of the conversation sub-data.

8. The system according to claim 7, characterized in that The training module is further configured to: Based on the conversation sub-data, conversation sub-features corresponding to the conversation sub-data, the target application scenario and the simulated conversation information, sub-quality data corresponding to the conversation sub-data is generated through a quality assessment model, and the quality assessment model is a machine learning model.

9. An intelligent AI scene simulation training device, characterized in that: The apparatus comprises at least one processor and at least one memory; The at least one memory is used to store computer instructions; The at least one processor is configured to execute at least part of the computer instructions to implement the method of any one of claims 1 to 5.

10. A computer-readable storage medium, characterized in that: The storage medium stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Training method and device based on virtual reality, virtual reality equipment and storage medium

    CN116741010A

  • Question and answer model construction method and project question and answer method

    CN117633196A

  • Customer service information generation method and system based on historical sessions

    CN119311840A

  • Method for incremental training of large model

    CN119338014A

  • Model training method and sample determination method for model training

    CN119357679A

Cited By

  • Data transmission method and related device

    CN121619314A

  • A data transmission method and related apparatus

    CN121619314B