An intelligent AI scene simulation training method, system, device and medium

By simulating real conversation scenarios and generating and feeding back quality data, it addresses the shortcomings of authenticity and interactivity in traditional conversation training methods and achieves a more efficient and high-quality training experience.

CN120067273BActive Publication Date: 2025-09-16SICHUAN RONGCHENG LEIMING TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510504160.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-09-16
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

Traditional dialogue training methods lack authenticity and interactivity, and are unable to restore the complexity and uncertainty of actual conversations, resulting in poor training results and user experience.

Method used

By generating a target training library, simulating real conversation scenarios, obtaining user turn-by-turn conversation information, generating quality data and candidate training data, determining the training data for the next round of conversation based on difficulty and quality data, and providing real-time feedback to the user terminal.

Benefits of technology

Provide a more intuitive and effective training experience, be able to monitor and evaluate training results in real time, improve training efficiency and quality, and enhance user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067273B_ABST
    Figure CN120067273B_ABST
Patent Text Reader

Abstract

The embodiments of this specification provide an intelligent AI scenario simulation training method, system, device, and medium. The method includes retrieving target training data and generating a target training library; responding to a target user performing training based on the target training data in the target training library, wherein the training includes multiple rounds of dialogue, and for each round of dialogue: obtaining the target user's round dialogue information; generating quality data corresponding to the round dialogue information based on the round training data and the round dialogue information; generating multiple candidate training data based on the round dialogue information; determining the round training data for the next round of dialogue based on the difficulty data and quality data corresponding to the multiple candidate training data; and sending the round training data and quality data for the next round of dialogue to a user terminal. This specification can provide a more intuitive and effective training experience, can effectively monitor and statistically analyze training results, provide real-time feedback and evaluation, and improve the efficiency and quality of training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of artificial intelligence technology, and in particular to an intelligent AI scene simulation training method, system, device and medium. Background Art

[0002] Traditional conversation training methods mainly rely on artificially simulated conversation scenarios or simple recordings for training, which lack authenticity and interactivity and make it difficult to restore the complexity and uncertainty of actual conversations, resulting in poor training results and user experience.

[0003] Therefore, we aim to propose an intelligent AI scenario simulation training method, system, device, and medium that simulates real-world conversation scenarios to provide a more intuitive and effective training experience. This system can effectively monitor and statistically analyze training results, provide real-time feedback and evaluation, improve training efficiency and quality, and enhance the user experience. Summary of the Invention

[0004] One or more embodiments of this specification provide an intelligent AI scenario simulation training method. The method includes: retrieving target training data and generating a target training library; in response to a target user performing training based on the target training data in the target training library, the training includes multiple rounds of dialogue, and for each round of dialogue: obtaining the round dialogue information of the target user, the round dialogue information corresponding to the round training data, and the round training data being the training data used for the current round of dialogue; based on the round training data and the round dialogue information, generating quality data corresponding to the round dialogue information; based on the round dialogue information, generating multiple candidate training data; based on the difficulty data and the quality data corresponding to the multiple candidate training data, determining the round training data for the next round of dialogue; and sending the round training data and the quality data for the next round of dialogue to the user terminal.

[0005] One or more embodiments of this specification provide an intelligent AI scenario simulation training system. The system includes a generation module and a training module; the training module is configured to retrieve target training data and generate a target training library; the scoring module is configured to respond to the target user for training based on the target training data in the target training library, and the training includes multiple rounds of dialogue. For each round of dialogue: the round dialogue information of the target user is obtained, the round dialogue information corresponds to the round training data, and the round training data is the training data used for this round of dialogue; based on the round training data and the round dialogue information, quality data corresponding to the round dialogue information is generated; based on the round dialogue information, multiple candidate training data are generated; based on the difficulty data and the quality data corresponding to the multiple candidate training data, the round training data for the next round of dialogue is determined; the round training data for the next round of dialogue and the quality data are sent to the user terminal.

[0006] One or more embodiments of this specification provide an intelligent AI scenario simulation training device, which includes at least one processor and at least one memory; the at least one memory is used to store computer instructions; and the at least one processor is used to execute at least part of the computer instructions to implement the intelligent AI scenario simulation training method described in the above embodiments.

[0007] One or more embodiments of this specification provide a computer-readable storage medium that stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes the intelligent AI scenario simulation training method described in the above embodiments.

[0008] The beneficial effects of this specification include but are not limited to: (1) providing a more intuitive and effective training experience; (2) enabling effective monitoring and statistical analysis of training results, real-time feedback and evaluation, and improving the efficiency and quality of training; and (3) using machine learning models to more carefully evaluate the impact of multi-dimensional data on sub-quality data and improve the accuracy of determining sub-quality data. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] This specification will be further described in the form of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting, and in these embodiments, like numbers represent like structures, wherein:

[0010] Figure 1 is an exemplary schematic diagram of an intelligent AI scenario simulation training system according to some embodiments of this specification;

[0011] Figure 2 is an exemplary flow chart of an intelligent AI scenario simulation training method according to some embodiments of this specification;

[0012] Figure 3 is an exemplary schematic diagram of determining quality data according to some embodiments of this specification;

[0013] Figure 4 is an exemplary schematic diagram of a quality assessment model according to some embodiments of this specification.

[0014] Explanation of the accompanying drawings: 110: generation module; 120: training module; 311-1, ..., 311-n: multiple dialogue sub-data; 312-1, ..., 312-n: multiple dialogue sub-features; 313: target application scenario A; 314-1, ..., 314-n: multiple simulated dialogue information; 320-1, ..., 320-n: multiple sub-quality data; 330: quality data; 411: training sub-data; 412: target application scenario B; 413: simulated dialogue information B; 420: quality assessment model; 421: information generation layer; 422: quality assessment layer; 431: standard dialogue sub-data; 432: standard dialogue sub-feature; 441: dialogue sub-data B; 442: dialogue sub-feature B; 450: sub-quality data B. DETAILED DESCRIPTION

[0015] In order to more clearly illustrate the technical solutions of the embodiments of this specification, the following briefly introduces the drawings required for describing the embodiments. The drawings do not represent all implementation methods.

[0016] It should be understood that the terms "system," "device," "unit," and / or "module" used herein are a method for distinguishing different components, elements, parts, portions, or assemblies at different levels. If other terms can achieve the same purpose, the terms may be replaced by other expressions.

[0017] Unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not refer to the singular but include the plural. Generally speaking, the terms "include" and "comprise" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.

[0018] When operations are performed according to the step descriptions in the embodiments of this specification, unless otherwise specified, the order of the steps is interchangeable, steps can be omitted, and other steps can be included in the operation process.

[0019] Figure 1 This is an exemplary schematic diagram of an intelligent AI scenario simulation training system according to some embodiments of this specification.

[0020] In some embodiments, the intelligent AI scenario simulation training system 100 includes a generation module 110 and a training module 120 .

[0021] In some embodiments, the generation module may be configured to retrieve target training data and generate a target training library.

[0022] In some embodiments, the training module can be configured to obtain the target user's round dialogue information in response to the target user performing training based on the target training data in the target training library; generate quality data corresponding to the round dialogue information based on the round training data and the round dialogue information; generate multiple candidate training data based on the round dialogue information; determine the round training data for the next round of dialogue based on the difficulty data and quality data corresponding to the multiple candidate training data; and send the round training data and quality data for the next round of dialogue to the user terminal.

[0023] In some embodiments, for a single dialogue sub-data among multiple dialogue sub-data, the training module is further configured to: generate sub-quality data corresponding to the dialogue sub-data based on the dialogue sub-data, dialogue sub-features corresponding to the dialogue sub-data, target application scenarios, and simulated dialogue information; and determine quality data based on multiple sub-quality data corresponding to multiple dialogue sub-data.

[0024] In some embodiments, the training module is further configured to: generate sub-quality data corresponding to the dialogue sub-data through a quality assessment model based on the dialogue sub-data, dialogue sub-features corresponding to the dialogue sub-data, the target application scenario and the simulated dialogue information.

[0025] The user terminal refers to the terminal device used by the target user, such as a mobile phone, tablet computer, etc. In some embodiments, the intelligent AI scenario simulation training system is communicatively connected to the user terminal.

[0026] In some embodiments, the intelligent AI scenario simulation training system may further include a processor and a storage device. The processor is configured to process data from at least one module of the intelligent AI scenario simulation training system 100 or an external data source. The processor includes a central processing unit (CPU), an application-specific integrated circuit (ASIC), a controller, a microcontroller unit, a microprocessor, or any combination thereof.

[0027] The storage device is configured to store data, instructions, and / or information related to the intelligent AI scenario simulation training system 100. For example, the storage device stores target training data and a target training library. In some embodiments, the storage device can be integrated into the processor or provided in the cloud.

[0028] In some embodiments, the generation module and the training module may be implemented by a processor.

[0029] For more information on the above, see Figure 2-Figure 4 and its related descriptions.

[0030] The intelligent AI scenario simulation training system generates a target training library through generation modules, which can provide targeted training for different users. By training target users through training modules, the training results can be fed back and evaluated in real time to improve the efficiency and quality of training.

[0031] It should be understood that Figure 1 The intelligent AI scene simulation training system and its modules shown can be implemented in various ways. It should be noted that the above description of the intelligent AI scene simulation training system and its modules is only for the convenience of description and does not limit this specification to the scope of the embodiments given. It is understandable that for those skilled in the art, after understanding the principles of the system, it is possible to arbitrarily combine the modules or form a subsystem to connect with other modules without deviating from this principle. In some embodiments, Figure 1 The generation module 110 and training module 120 disclosed herein may be separate modules within a single system, or a single module may implement the functions of two or more of the aforementioned modules. For example, the modules may share a storage module, or each module may have its own storage module. Such variations are within the scope of this specification.

[0032] Figure 2 This is an exemplary flow chart of the intelligent AI scenario simulation training method according to some embodiments of this specification. In some embodiments, process 200 is executed by a processor. Figure 2 As shown, process 200 includes the following steps:

[0033] Step 210: retrieve target training data and generate a target training library.

[0034] Target training data refers to data used for training. In some embodiments, training includes dialogue training, etc. Dialogue training refers to training in which the processor asks questions and the target user answers them.

[0035] In some embodiments, the target training data includes a plurality of training sub-data, etc. The training sub-data refers to a dialogue sentence required for dialogue training for the target user.

[0036] Target users are those who need conversation training, for example, at least one of customer service representatives, tellers, salespeople, and e-commerce users.

[0037] The target training database refers to the database used to store target training data.

[0038] In some embodiments, the processor generates the target training library using various methods. For example, the processor obtains target training data corresponding to the target user from a storage device, adds the target training data to the target training library corresponding to the target user, and stores the target training library in the storage device for use in training the target user.

[0039] In some embodiments, target training data corresponding to different target users can be pre-set based on experience and stored in a storage device. A target training library can correspond to one or more target users.

[0040] In some embodiments, the processor may determine reference difficulty data based on multiple historical training results of multiple historical trainings of the target user, and retrieve target training data based on the reference difficulty data.

[0041] Historical training results refer to the comprehensive training results corresponding to historical training. For an explanation of comprehensive training results, please refer to the following and related descriptions.

[0042] The reference difficulty data refers to the reference value of the difficulty data of the training sub-data to be retrieved.

[0043] In some embodiments, the processor calculates an average of the multiple historical training results from multiple historical training sessions for the target user. If the average meets a preset condition, the processor reduces the initial reference difficulty by a preset adjustment amount to obtain reference difficulty data. If the average does not meet the preset condition, the processor determines the initial reference difficulty as the reference difficulty data.

[0044] The preset conditions include the mean value being no greater than the quality threshold. The quality threshold and the preset adjustment amount are pre-set based on historical experience.

[0045] The initial reference difficulty refers to the initially set reference difficulty data. In some embodiments, the initial reference difficulty is pre-set based on historical experience.

[0046] In some embodiments, the preset adjustment amount may be related to the training frequency of the target user. For example, the preset adjustment amount is negatively correlated to the training frequency of the target user.

[0047] Training frequency refers to how often the target user undergoes training. A higher training frequency indicates shorter intervals between training sessions for the target user. This means the target user's skill level remains stable, so there's no need to significantly adjust the original reference difficulty data, resulting in a smaller preset adjustment amount.

[0048] The preset adjustment amount can be adjusted to better match the target user through the training frequency of the target user, to avoid the preset adjustment amount being too large when the target user is trained frequently, resulting in the target user undergoing multiple trainings with a wide range of difficulty in a short period of time, which reduces the effectiveness of the training.

[0049] In some embodiments, the processor may eliminate multiple training sub-data from the target training data that do not meet a preset retrieval condition, and generate a target training library based on the eliminated target training data. The preset retrieval condition includes the difference between the difficulty data of the training sub-data and the reference difficulty data being less than a first difference threshold. The first difference threshold is preset based on historical experience.

[0050] Based on the target user's multiple historical training results, the target user's own level is evaluated, and then the retrieved training sub-data is dynamically adjusted to make the target training data more in line with the target user's current ability and improve the training effect.

[0051] Step 220 , in response to the target user performing training based on the target training data in the target training library, the training includes multiple rounds of dialogue, and steps 221 to 225 are executed for each round of dialogue.

[0052] It is understandable that training is equivalent to simulating a real dialogue scenario, in which the two parties in the real dialogue scenario will conduct multiple rounds of dialogue to complete the communication.

[0053] Step 221: Obtain the target user's turn-based conversation information.

[0054] Turn-based conversation information refers to the target user's responses in this conversation turn. In some embodiments, turn-based conversation information includes one or more conversation sub-data and the corresponding conversation sub-features for each conversation sub-data. A conversation sub-data refers to a single sentence in which the target user responds to the training sub-data. A conversation sub-feature refers to the relevant information and features of the conversation sub-data.

[0055] In some embodiments, the conversation sub-features include the answer waiting time, answer speaking speed, and answer volume corresponding to the conversation sub-data.

[0056] In some embodiments, the processor may record the target user's conversation sub-data, answer waiting time, answer speed, and answer volume in real time during the conversation to obtain turn-based conversation information.

[0057] Step 222: Generate quality data corresponding to the round dialogue information based on the round training data and the round dialogue information.

[0058] In some embodiments, round training data refers to training data used in the current round of dialogue, for example, training sub-data used by the processor in the current round of dialogue.

[0059] In some embodiments, the turn-based conversation information corresponds to turn-based training data. It is understood that the multiple training sub-data in the target training data have a specific word order. During training, the processor selects a training sub-data from the target training database based on the word order as the turn-based training data. Accordingly, the target user responds based on the turn-based training data, resulting in one or more conversation sub-data.

[0060] For example, the target training data is used to train target users to perform after-sales service, which may include product description, refund or replacement of products, and service ratings, etc. The processor can set the training sub-data used in the dialogue according to the word order of product description, refund or replacement of products, and service ratings.

[0061] In some embodiments, after generating the target training library, the processor may construct a dialogue graph based on the target training data in the target training library.

[0062] The answer graph is a graph structure that reflects the relationship between target training data and corresponding answers. The graph structure is a data structure composed of nodes and edges. Edges connect nodes, and nodes and edges can have features.

[0063] In some embodiments, the dialogue graph includes question nodes and answer nodes.

[0064] The question node represents a training sub-data. The node characteristics of the question node include difficulty data, etc. The difficulty data refers to data used to characterize the difficulty level of the training sub-data. In some embodiments, the difficulty data is pre-set based on historical experience.

[0065] A response node is a response sentence corresponding to the training sub-data. Node features of a response node include the response wait time, response rate, and response volume. Response wait time is the time interval between receiving the training sub-data and making a response.

[0066] In some embodiments, the processor queries a preset question-and-answer table for multiple question-and-answer statements corresponding to the training sub-data, identifies the multiple question-and-answer statements as answer nodes corresponding to the training sub-data, and identifies the answer waiting time, answer speaking rate, and answer volume corresponding to each question-and-answer statement as node features of the answer node. The preset question-and-answer table is pre-set based on historical experience and includes multiple training sub-data and multiple answer statements corresponding to different training sub-data, as well as the answer waiting time, answer speaking rate, and answer volume corresponding to each answer statement.

[0067] In some embodiments, there is an edge between the question node and the corresponding answer node, from the question node to the answer node.

[0068] In some embodiments, because the multiple training sub-data in the target training data have a certain word order, after answering a training sub-data, the processor can select the subsequent training sub-data to ask based on the answer. Therefore, some answer nodes also have outgoing edges, which point to one or more question nodes used in subsequent questions.

[0069] Quality data refers to data used to characterize the quality of the target user's responses in the current conversation round. In some embodiments, the processor determines initial quality data based on the similarity between one or more dialogue sub-data in the conversation round information and the corresponding multiple answer nodes. For example, the processor calculates the similarity between one or more dialogue sub-data in the conversation round information and connects them into a long sentence, calculates the average similarity between the long sentence and the corresponding multiple answer nodes, and uses the average similarity as the initial quality data. The multiple answer nodes corresponding to the long sentence are the multiple answer nodes pointed to by the question node representing the training sub-data corresponding to the conversation round information in the answer graph.

[0070] In some embodiments, the processor may calculate the similarity between the long sentence and the corresponding answer node through a BERT model, a SentenceBERT model, a SimCSE model, or the like.

[0071] In some embodiments, the processor may further adjust the initial quality data based on the multiple dialogue sub-features in the turn-based dialogue information and the node feature mean of the corresponding multiple answer nodes to obtain the final quality data. For example, the processor calculates the mean of each data item in the multiple dialogue sub-features, and calculates the ratio of the mean of each data item to the corresponding data item in the node feature mean, and adjusts the initial quality data based on the multiple ratios and the node feature mean using a preset algorithm. Exemplarily, the preset algorithm is shown in the following formula (1):

[0072] B=A×[1 - (k1 × C1 / D1 + ……+ kn × Cn / Dn) / (k1 ×D1 + …… + kn ×Dn)] (1)

[0073] Where A represents the initial quality data, B represents the final quality data, C1, ..., Cn represent the mean of the first to nth data items in the conversation sub-features, D1, ..., Dn represent the mean of the first to nth data items in the node feature of the multiple answer nodes, and k1, ..., kn represent the weight of the first to nth data items. The weights of different data items can be preset based on experience. n represents the data items included in the conversation sub-feature, for example, 3.

[0074] In some embodiments, the processor may further determine the quality data based on a plurality of sub-quality data corresponding to the plurality of conversation sub-data. For more information, see Figure 3 and its related descriptions.

[0075] Step 223: Generate multiple candidate training data based on the turn-based dialogue information.

[0076] Candidate training data refers to the training sub-data to be used in the next round of dialogue. It's understandable that there may be multiple responses to a particular training sub-data, each corresponding to a different subsequent question. Therefore, based on the current round of dialogue information (i.e., the current one or more responses), the processor can select some training sub-data from the multiple training sub-data for the next round of dialogue (i.e., the subsequent question statements) as candidate training sub-data, and determine the final training sub-data to be used in the next round of dialogue from this candidate training sub-data.

[0077] In some embodiments, the processor can obtain candidate training data in various ways. For example, based on the similarity between the turn-based dialogue information and the corresponding multiple answer nodes, the processor determines an answer node that meets a candidate condition, and uses the multiple question nodes pointed to by the outgoing edges of the answer node as multiple candidate training data. The candidate condition includes the highest similarity between the turn-based dialogue information and the answer node.

[0078] In some embodiments, the processor may further determine standard difficulty data based on the variation characteristics corresponding to the plurality of quality data and the difficulty data corresponding to the round training data, and generate a plurality of candidate training data based on the standard difficulty data and the round dialogue information.

[0079] It should be noted that at least one round of dialogue is required before the processor can determine the change characteristics corresponding to multiple quality data and then generate multiple candidate training data.

[0080] The change feature refers to data used to characterize the change of quality data. In some embodiments, the change feature includes a change curve, etc. The x-axis of the change curve is the training time, and the y-axis is the quality data.

[0081] In some embodiments, the processor may record multiple training time points during multiple rounds of conversation, calculate quality data corresponding to the training time points, and construct a change curve based on the multiple training time points and the quality data corresponding to different training time points. A training time point refers to the time point at the end of a round of conversation.

[0082] Standard difficulty data refers to the reference value of the difficulty data of the candidate sub-data. In some embodiments, the processor can take the first-order derivatives of multiple sampling points on the change curve. If the first-order derivatives of N consecutive sampling points from the current time point meet the first adjustment condition, the difficulty data of the training sub-data of the current round of dialogue is increased by a difficulty adjustment amount as the standard difficulty data. If the first-order derivatives of N consecutive sampling points from the current time point meet the second adjustment condition, the difficulty data of the training sub-data of the current round of dialogue is reduced by a difficulty adjustment amount as the standard difficulty data. The difficulty adjustment amount and the multiple sampling points can be pre-set based on historical experience. The difficulty data of the training sub-data is obtained through the dialogue graph. N is a preset number.

[0083] The first adjustment condition includes that all first-order derivatives are greater than 0, and the second adjustment condition includes that all first-order derivatives are less than 0. If the first-order derivatives of N consecutive sampling points from the current time point do not meet the first adjustment condition and the second adjustment condition, the processor uses the difficulty data of the training sub-data of the current round of dialogue as the standard difficulty data.

[0084] In some embodiments, the N value may also be related to the number of reversals of the first-order derivative at multiple sampling points on the change curve. For example, the N value is positively correlated with the number of reversals. The number of reversals refers to the number of times the first-order derivative changes from positive to negative. The greater the number of reversals, the more difficult it is to determine whether the quality data is in an upward or downward trend. In this case, a larger N value is required to improve the accuracy of determining whether the quality data is in an upward or downward trend.

[0085] In some embodiments, the processor may select an answer node that meets the candidate criteria in the answer graph, determine multiple question nodes to which outgoing edges of the answer node point, and select multiple question nodes whose difficulty data differs from the standard difficulty data by no more than a second difference threshold as multiple candidate training data. The second difference threshold may be preset based on experience.

[0086] By using the changes in quality data from previous rounds of training and the difficulty data of the current training sub-data, more applicable standard difficulty data can be determined, and then multiple candidate training data that are more suitable for the current target user level can be generated to improve the training effect.

[0087] Step 224 : determining round training data for the next round of dialogue based on the difficulty data and quality data corresponding to the plurality of candidate training data.

[0088] In some embodiments, the processor may calculate the ratio of the quality data of the current round of conversation to the qualified value, and calculate the average of the difficulty data of multiple candidate training data. The processor multiplies the obtained ratio by the average of the difficulty data to obtain a comprehensive difficulty data. The processor then selects the candidate training data with the smallest absolute difference between the difficulty data and the comprehensive difficulty data from the multiple candidate training data, and uses this candidate training data as the round training data for the next round of conversation. The qualified value can be preset based on experience.

[0089] Step 225: Send the round training data and quality data of the next round of dialogue to the user terminal.

[0090] In some embodiments, the processor can send round training data for the next round of conversation and quality data of the current round of conversation to the user terminal for the target user to review and conduct the next round of conversation. The processor can complete multi-round conversation training by sending round training data to the target user multiple times, recording the corresponding round of conversation information, and determining the quality data.

[0091] Comprehensive training results refer to data used to characterize the quality of training completed by target users.

[0092] In some embodiments, in response to the target user completing training, the processor may generate a comprehensive training result corresponding to the training based on multiple quality data corresponding to multiple rounds of conversations. For example, the processor may determine the average of the multiple quality data as the comprehensive training result corresponding to the training. In another example, the processor may determine the weighted sum of the multiple quality data as the comprehensive training result corresponding to the training. The weights for different rounds of conversations are pre-set based on historical experience.

[0093] By integrating the quality data of multiple rounds of conversations with target users, intuitive comprehensive training results are generated after the training is completed. This is conducive to a comprehensive evaluation of the training effect, providing growth feedback for target users, and facilitating the determination of target training data that is more in line with the target user's level in subsequent training.

[0094] By building a corresponding target training library for target users and conducting training, we can provide target users with more intuitive and effective training by simulating real dialogue scenarios. At the same time, we can effectively monitor and statistically analyze the training results, provide real-time feedback and evaluation, and improve the efficiency and quality of training.

[0095] It should be noted that the above description of process 200 is for illustration and purpose only and does not limit the scope of application of this specification. Those skilled in the art may make various modifications and changes to the process under the guidance of this specification. However, such modifications and changes are still within the scope of this specification.

[0096] Figure 3is an exemplary schematic diagram of determining quality data according to some embodiments of this specification.

[0097] In some embodiments, the turn-based conversation information includes multiple conversation sub-data and multiple conversation sub-features. The target training data also includes target application scenarios and simulated conversation information. One training sub-data may correspond to one target application scenario and simulated conversation information. For more information about turn-based conversation information, see Figure 2 and its related descriptions.

[0098] In some embodiments, for a single conversation sub-data among the plurality of conversation sub-data, the processor may generate sub-quality data corresponding to the conversation sub-data based on the conversation sub-data (e.g., conversation sub-data 311-1, ..., conversation sub-data 311-n, where n is the number of conversation sub-data), conversation sub-features corresponding to the conversation sub-data (e.g., conversation sub-features 312-1, ..., conversation sub-features 312-n), target application scenarios (e.g., target application scenarios 313), and simulated conversation information (e.g., simulated conversation information 314-1, ..., simulated conversation information 314-n). The processor then determines quality data 330 based on the plurality of sub-quality data corresponding to the plurality of conversation sub-data (e.g., sub-quality data 320-1, ..., sub-quality data 320-n). For a description of target training data, conversation sub-data, conversation sub-features, and quality data, see [ 330 ]. Figure 2 and its related descriptions.

[0099] The target application scenario refers to the application scenario corresponding to the training sub-data. In some embodiments, the target application scenario includes the simulated customer conversation environment, mood, and / or customer profile corresponding to the training sub-data. The customer refers to the simulated person who will have a conversation with the target user. One target application scenario corresponds to one or more training sub-data.

[0100] In some embodiments, the conversation environment includes home or work. Mood status includes good or average. Customer profile refers to information used to describe the customer. The customer profile includes the customer's gender, age, occupation, etc.

[0101] In some embodiments, the processor can pre-store multiple different target application scenarios, and before sending the training sub-data to the user terminal of the target user, randomly select a target application scenario to integrate into the training sub-data, and send the adjusted training sub-data to the user terminal of the target user for training.

[0102] In some embodiments, integrating the target application scenario into the training sub-data includes: the processor integrating one or more question keywords corresponding to the target application scenario into the training sub-data. In some embodiments, the processor obtains the question keywords corresponding to the target application scenario from a preset scenario table.

[0103] In some embodiments, the preset scenario table is pre-set based on historical experience and includes multiple target application scenarios and one or more question keywords corresponding to different target application scenarios. The one or more question keywords corresponding to the target application scenarios are determined by technicians based on experience.

[0104] Simulated conversation information refers to information related to simulated responses to training sub-data. In some embodiments, the simulated conversation information includes one or more simulated sub-data and simulated sub-features corresponding to the different simulated sub-data. Simulated sub-data refers to simulated conversation sub-data. Simulated sub-features refer to conversation sub-features corresponding to the simulated sub-data.

[0105] One training sub-data corresponds to one or more simulation sub-data.

[0106] In some embodiments, the processor generates simulated conversation information in various ways. For example, the processor may obtain a simulated speech segment input by a technician, split the simulated speech segment into multiple simulated sentences that are more colloquial or easier to read as multiple simulated sub-data, and determine the simulated sub-features corresponding to the simulated sub-data based on the waiting time, speech rate, and volume of the simulated speech segment input.

[0107] In some embodiments, the processor may further generate simulated conversation information based on historical training data and historical quality scores of the target user.

[0108] Historical training data refers to relevant data in the historical training process. In some embodiments, the historical training data includes multiple historical training sub-data used by the target user and corresponding multiple historical rounds of conversation information.

[0109] Historical quality data refers to quality data corresponding to historical conversation sub-data. In some embodiments, the historical quality data includes quality data corresponding to multiple historical rounds of conversation information in the historical training data.

[0110] In some embodiments, the processor retrieves historical training data and historical quality data from a storage device.

[0111] In some embodiments, the processor may insert the dialogue defects included in the historical round dialogue information into the initial simulated dialogue information based on the degree of defect visibility and the number of defects to generate simulated dialogue information.

[0112] Conversational flaws refer to deficiencies in the target user's responses. For example, these include inappropriate word choice, slow or fast speech, repetitive semantics, and logical inconsistencies. Conversational flaws in historical conversational turn information are identified through manual annotation or automated detection by the processor.

[0113] Defect conspicuity refers to data used to characterize the degree of conspicuity of a dialog defect. Defect conspicuity can be represented by a numerical value, for example, where a larger numerical value indicates a greater degree of defect conspicuity.

[0114] In some embodiments, the processor calculates an average of multiple quality data in the historical quality data, and determines the degree of defect significance and the number of defects by querying a preset defect table.

[0115] In some embodiments, the preset defect table is pre-set based on historical experience and includes multiple quality data means and the defect visibility and defect quantity corresponding to different quality data means. The defect visibility and defect quantity can be determined by manually annotating the historical rounds of conversation information corresponding to the quality data means.

[0116] In some embodiments, for each of the multiple historical training sub-data, the processor can also search for a question node corresponding to the historical training sub-data in the dialogue graph based on the historical training sub-data. The processor randomly selects a target application scenario, integrates one or more question keywords corresponding to the target application scenario into multiple answer nodes corresponding to the question node, and generates initial simulated dialogue information based on the multiple answer nodes and corresponding node features. For more information about the dialogue graph, see Figure 2 and related instructions.

[0117] In some embodiments, based on the number of defects, the processor inserts into the initial simulated conversation information a number of dialogue defects included in the historical rounds of conversation information equal to the number of defects. The processor may also insert corresponding dialogue defects into the initial simulated conversation information based on the degree of defect visibility. For example, if the defect visibility is high, the dialogue defect may be inappropriate wording, logical contradictions, etc., while if the defect visibility is low, the dialogue defect may be too slow or too fast a speech rate, etc.

[0118] By collecting the target user's historical training data and historical quality data and simulating the target user's conversation defects, we can generate realistic simulated conversation information, which is conducive to more accurate evaluation of the target user's training quality in the future.

[0119] Sub-quality data refers to quality data corresponding to the dialogue sub-data.

[0120] In some embodiments, the processor obtains the sub-quality data in various ways. For example, the processor adjusts the initial sub-quality data based on the degree of fit between the dialogue sub-data and the target application scenario and the degree to which the dialogue sub-data compensates for defects in the simulated dialogue information to obtain the sub-quality data.

[0121] In some embodiments, the processor calculates the similarity between the dialogue sub-data and the corresponding multiple answer nodes, and uses the highest similarity as the initial sub-quality data of the dialogue sub-data. For an explanation of calculating similarity, see Figure 2 and related instructions.

[0122] Scenario fit is used to characterize the degree of fit between the conversation sub-data and the target application scenario. In some embodiments, the processor calculates the overlap between the conversation sub-data and one or more answer keywords corresponding to the target application scenario, and determines the resulting overlap as the scenario fit. The processor determines the overlap as the ratio of the number of answer keywords contained in the conversation sub-data to the total number of answer keywords corresponding to the target application scenario.

[0123] The method for determining answer keywords is similar to that for determining question keywords.

[0124] The defect compensation degree is used to characterize the degree to which the conversation sub-data compensates for the simulated conversation information. In some embodiments, the processor compares the conversation sub-data and conversation sub-features with the simulated conversation information and determines the defect compensation degree as the ratio of the number of conversation defects in the simulated conversation information compensated by the conversation sub-data or conversation sub-features to the total number of conversation defects in the simulated conversation information. For example, if the simulated conversation information lacks the word "weekday" but the conversation sub-data contains "weekday," the conversation sub-data is determined to have compensated for a defect in the simulated conversation information. If the response speed in the simulated conversation information is fast, but the response speed in the conversation sub-features is moderate, the conversation sub-feature is determined to have compensated for a defect in the simulated conversation information.

[0125] In some embodiments, the processor may calculate the product of the scene fit and the defect compensation degree, add 1 to the obtained product, and multiply the result by the initial sub-quality data to obtain the sub-quality data.

[0126] In some embodiments, the processor may also generate sub-quality data corresponding to the conversation sub-data through the quality assessment model. Figure 4 and its related contents.

[0127] In some embodiments, the processor may perform weighted summation of multiple sub-quality data corresponding to multiple conversation sub-data to obtain quality data, wherein the weights of different conversation sub-data are pre-set based on historical experience.

[0128] By conducting in-depth analysis of conversation sub-data, target application scenarios, and simulated conversation information, more accurate and comprehensive sub-quality data can be determined, and thus more accurate quality data can be determined.

[0129] Figure 4 is an exemplary schematic diagram of a quality assessment model according to some embodiments of this specification.

[0130] In some embodiments, the processor may generate sub-quality data corresponding to the conversation sub-data using a quality assessment model based on the conversation sub-data, the conversation sub-features corresponding to the conversation sub-data, the target application scenario, the simulated conversation information, and the target training data corresponding to the conversation sub-data (i.e., the training sub-data corresponding to the conversation sub-data). For a description of the conversation sub-data, the conversation sub-features corresponding to the conversation sub-data, the target application scenario, the simulated conversation information, and the training sub-data, see Figure 2-Figure 3 and its related descriptions.

[0131] The quality assessment model 420 is a model used to assess the sub-quality data. In some embodiments, the quality assessment model can be a machine learning model. For example, the quality assessment model can be a deep neural network (DNN) model or a custom model, or a combination thereof.

[0132] In some embodiments, the input of the quality assessment model includes training sub-data, target application scenarios, simulated conversation information, conversation sub-data, and conversation sub-features, and the output includes sub-quality data corresponding to the conversation sub-data.

[0133] In some embodiments, the quality assessment model can be trained using training samples and corresponding labels from a training sample set. The training samples include sample training sub-data, sample target application scenarios, sample simulated conversation information, sample conversation sub-data, and sample conversation sub-features. The labels corresponding to the training samples include actual sub-quality data of the sample conversation sub-data. In some embodiments, the training samples are obtained based on historical data. The labels are pre-annotated and determined by technical personnel.

[0134] In some embodiments, the quality assessment model includes an information generation layer 421 and a quality assessment layer 422. For example, the information generation layer and the quality assessment layer may include a combination of one or more Deep Neural Networks (DNN) models or custom models.

[0135] In some embodiments, the inputs to information generation layer 421 may include training sub-data 411, target application scenarios (e.g., target application scenario B 412), and simulated conversation information (e.g., simulated conversation information B 413). The outputs may include standard conversation sub-data 431 and standard conversation sub-features 432. Standard conversation sub-data refers to standard responses generated by the information generation layer. Standard conversation sub-features refer to conversation sub-features corresponding to the standard responses generated by the information generation layer.

[0136] In some embodiments, the input of the quality assessment layer 422 may include standard dialogue sub-data 431, standard dialogue sub-features 432, dialogue sub-data (such as dialogue sub-data B441) and dialogue sub-features (such as dialogue sub-features B442), and the output may include sub-quality data corresponding to the dialogue sub-data (such as sub-quality data B450).

[0137] In some embodiments, the information generation layer and the quality assessment layer can be obtained through training samples in the training sample set and corresponding label training respectively.

[0138] In some embodiments, the training sample set includes a first training sample from a training information generation layer and a first label corresponding to the first training sample. The first training sample includes sample training sub-data, a sample target application scenario, and sample simulated conversation information. The first label includes standard conversation sub-data and standard conversation sub-features corresponding to the sample training sub-data.

[0139] In some embodiments, the first training sample is obtained based on historical data. The processor can select multiple historical conversation sub-data corresponding to the sample training sub-data, calculate the scene fit of each historical conversation sub-data with the sample target application scene and the defect compensation degree of the sample simulated conversation information, and use the historical conversation sub-data with the highest weighted sum of the defect compensation degree and the scene fit as the standard conversation sub-data corresponding to the sample training sub-data, and use the historical conversation sub-features corresponding to the historical conversation sub-data as the standard conversation sub-features corresponding to the sample training sub-data. The weights of the scene fit and the defect compensation degree are preset based on historical experience. For an explanation of the scene fit and defect compensation degree, see Figure 3 and its related descriptions.

[0140] In some embodiments, the training sample set includes a second training sample from the training quality assessment layer and a second label corresponding to the second training sample. The second training sample includes the sample standard conversation sub-data and sample standard conversation sub-features output by the information generation layer, as well as the sample conversation sub-data and sample conversation sub-features. The second label includes the actual sub-quality data of the sample conversation sub-data. In some embodiments, the second training sample is obtained based on historical data and the information generation layer. The second label is pre-annotated and determined by a technician.

[0141] In some embodiments, one first training sample corresponds to one second training sample.

[0142] In some embodiments, the information generation layer and the quality assessment layer can be trained by the following joint training method: multiple first training samples with a first label are input into the initial information generation layer, and the output of the initial information generation layer and second training samples with a second label are input into the initial quality assessment layer, a loss function is constructed using the second label and the prediction result of the initial quality assessment layer, and the initial information generation layer and the initial quality assessment layer are synchronously updated based on the iteration of the loss function. When the loss function satisfies a preset iteration condition, the quality assessment model training is completed. The preset iteration condition may be the convergence of the loss function, the number of iterations reaching a set value, etc.

[0143] In some embodiments, labels of multiple training samples corresponding to different application scenario categories in the training sample set meet preset training conditions.

[0144] Application scenario classification refers to the classification of target application scenarios. In some embodiments, the application scenario classification can be pre-set based on historical experience. For example, for multiple target application scenarios, the processor can classify the multiple target application scenarios into different application scenario classifications based on at least one of the conversation environment, mood state, and / or customer profile. Exemplarily, the processor classifies the multiple target application scenarios into three categories: home, work, and street based on the conversation environment.

[0145] The training sample corresponding to the application scenario classification refers to the first training sample where the target application scenario corresponding to the application scenario classification is located.

[0146] The preset training condition includes that the standard deviation of the second label of the second training sample corresponding to the first training sample is greater than the preset distribution threshold. In some embodiments, the preset distribution threshold is related to the historical quality data of the target user. For example, the greater the range of multiple quality data in the historical quality data, the greater the preset score threshold. For a description of historical quality data, see Figure 3 and its related descriptions.

[0147] Since the greater the range of target user quality data, the greater the fluctuation in the quality of their responses, the lower the accuracy of the quality assessment model output. By introducing training samples with significant differences in different application scenarios, the model's generalization and adaptability can be enhanced, thereby optimizing training results and improving the accuracy of sub-quality data.

[0148] By employing machine learning models, we can more meticulously assess the impact of multi-dimensional data on sub-quality data, improving the accuracy of sub-quality data determination. Furthermore, the hierarchical structure allows for targeted processing of different input data, thereby improving the accuracy of the final model output. Joint training addresses issues such as insufficient training samples and difficulty obtaining labels.

[0149] Some embodiments of this specification also provide an intelligent AI scenario simulation training device, comprising at least one processor and at least one memory. The at least one memory is configured to store computer instructions. The at least one processor is configured to execute at least some of the computer instructions to implement any of the intelligent AI scenario simulation training methods described in any of the above embodiments.

[0150] Some embodiments of this specification also provide a computer-readable storage medium that stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes the intelligent AI scenario simulation training method described in any one of the above embodiments.

[0151] Furthermore, certain features, structures, or characteristics in one or more embodiments of this specification may be appropriately combined.

[0152] In some embodiments, numbers describing the quantities of components and attributes are used. It should be understood that such numbers used in the description of the embodiments are modified in some examples by the modifiers "about," "approximately," or "substantially." Unless otherwise specified, "about," "approximately," or "substantially" indicate that the numbers allow a variation of ±20%.

[0153] If there is any inconsistency or conflict between the descriptions, definitions, and / or usage of terms in the materials cited in this specification and the contents described in this specification, the descriptions, definitions, and / or usage of terms in this specification shall prevail.

Claims

1. An intelligent AI scenario simulation training method, characterized in that: The method is executed by a processor, and includes: Retrieve target training data and generate target training library; In response to a target user performing training based on target training data in the target training library, the target training data includes a target application scenario and simulated dialogue information, the simulated dialogue information is determined based on historical training data and historical quality data of the target user, the training includes multiple rounds of dialogue, and for each round of dialogue: Obtaining round-dialogue information of the target user, where the round-dialogue information corresponds to round-dialogue training data, which is training data used for the current round of dialogue, and the round-dialogue information includes a plurality of dialogue sub-data and a plurality of dialogue sub-features; For a single conversation sub-data among the plurality of conversation sub-data, generating sub-quality data corresponding to the conversation sub-data using a quality assessment model based on the conversation sub-data, the conversation sub-features corresponding to the conversation sub-data, the target application scenario, the simulated conversation information, and the target training data corresponding to the conversation sub-data, wherein the quality assessment model is a machine learning model; Determining quality data corresponding to the round dialogue information based on the plurality of sub-quality data corresponding to the plurality of dialogue sub-data; Determining standard difficulty data based on the change characteristics corresponding to the plurality of quality data and the difficulty data corresponding to the round training data; generating a plurality of candidate training data based on the standard difficulty data and the turn-based dialogue information; Determining round training data for a next round of dialogue based on the difficulty data and the quality data corresponding to the plurality of candidate training data; The round training data of the next round of dialogue and the quality data are sent to the user terminal.

2. The method according to claim 1, wherein The acquiring of target training data and generating a target training library comprises: Determining reference difficulty data based on multiple historical training results of multiple historical trainings of the target user and preset adjustment amounts; Based on the reference difficulty data, the target training data is retrieved.

3. The method according to claim 1, wherein The method further comprises: In response to the target user completing the training, a comprehensive training result corresponding to the training is generated based on the plurality of quality data corresponding to the multiple rounds of dialogues.

4. An intelligent AI scenario simulation training system, characterized in that: The system includes a generation module and a training module; The generating module is configured to retrieve target training data and generate a target training library; The training module is configured to perform training based on target training data in the target training library in response to a target user, wherein the target training data includes a target application scenario and simulated dialogue information, wherein the simulated dialogue information is determined based on historical training data and historical quality data of the target user, and wherein the training includes multiple rounds of dialogue, and for each round of dialogue: Obtaining round-dialogue information of the target user, where the round-dialogue information corresponds to round-dialogue training data, which is training data used for the current round of dialogue, and the round-dialogue information includes a plurality of dialogue sub-data and a plurality of dialogue sub-features; For a single conversation sub-data among the plurality of conversation sub-data, generating sub-quality data corresponding to the conversation sub-data using a quality assessment model based on the conversation sub-data, the conversation sub-features corresponding to the conversation sub-data, the target application scenario, the simulated conversation information, and the target training data corresponding to the conversation sub-data, wherein the quality assessment model is a machine learning model; Determining quality data corresponding to the round dialogue information based on the plurality of sub-quality data corresponding to the plurality of dialogue sub-data; Determining standard difficulty data based on the change characteristics corresponding to the plurality of quality data and the difficulty data corresponding to the round training data; generating a plurality of candidate training data based on the standard difficulty data and the turn-based dialogue information; Determining round training data for a next round of dialogue based on the difficulty data and the quality data corresponding to the plurality of candidate training data; The round training data of the next round of dialogue and the quality data are sent to the user terminal.

5. An intelligent AI scene simulation training device, characterized in that: The apparatus comprises at least one processor and at least one memory; The at least one memory is for storing computer instructions; The at least one processor is configured to execute at least part of the computer instructions to implement the method according to any one of claims 1 to 3.

6. A computer-readable storage medium, characterized in that The storage medium stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Training method and device based on virtual reality, virtual reality equipment and storage medium

    CN116741010A

  • Method for incremental training of large model

    CN119338014A

  • Model training method and sample determination method for model training

    CN119357679A