Training method, interaction method, device and equipment for intelligent interaction model
By introducing training methods of central control sub-model and sub-interaction model in the intelligent interaction model, combined with matching scoring and reinforcement learning, the problems of poor context coherence and low dialogue quality in multiple rounds of interaction are solved, and more efficient multi-round interaction performance and user experience are achieved.
Patent Information
- Application Number
- CN202111447279.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-11-29
AI Technical Summary
The existing intelligent interaction model performs poorly in multiple rounds of interactive tasks, resulting in poor context coherence and low conversation quality, which affects the user experience.
A training method of intelligent interaction model is adopted. By obtaining multiple input statements, the central control sub-model and sub-interaction model are used for interaction prediction and matching analysis, and reinforcement learning training is carried out based on the matching score to improve the multi-round interaction performance of the model.
It effectively improves the performance of the intelligent interaction model in multiple rounds of interactive tasks, improves user experience, and improves conversation quality and context coherence.
Smart Images

Figure CN114118451B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of human-computer interaction technologies, and in particular to a training method, an interaction method, a device, and a device for an intelligent interaction model. Background Art
[0002] In recent years, with the rapid development of artificial intelligence technologies, various types of machine learning models have achieved relatively good application effects in fields such as image classification, face recognition, and autonomous driving. Among them, in the field of human-computer interaction, an intelligent interaction model can be built through artificial intelligence technologies to collect relevant information and process tasks based on human-computer conversations.
[0003] In related technologies, when interacting with a user based on an intelligent interaction model, the interaction effect mainly depends on the training effect of the intelligent interaction model. However, due to certain limitations in the abundance of training data, the setting of interaction rule strategies, etc., when the number of interaction rounds is large, problems such as poor context coherence and low dialogue quality often occur, resulting in a poor interaction effect and affecting the user experience.
[0004] In summary, the problems existing in related technologies urgently need to be solved. Summary of the Invention
[0005] An object of the present application is to solve at least to some extent one of the technical problems existing in related technologies.
[0006] To this end, an object of an embodiment of the present application is to provide a training method for an intelligent interaction model, which can effectively improve the performance of the trained intelligent interaction model when processing multi-round interaction tasks, and is beneficial to improving the user experience when the intelligent interaction model is put into operation.
[0007] To achieve the above technical object, the technical solutions adopted in the embodiments of the present application include:
[0008] On the one hand, an embodiment of the present application provides a training method for an intelligent interaction model, where the intelligent interaction model includes a central control sub-model and a plurality of trained different sub-interaction models; the training method includes:
[0009] Obtain first input information, where the first input information includes a plurality of first input statements;
[0010] Input the first input statement into the intelligent interaction model, perform interaction prediction on the first input statement through each of the sub-interaction models to obtain a plurality of initial output statements, and determine a target sub-interaction model according to the first input statement through the central control sub-model, and determine the initial output statement output by the target sub-interaction model as the target output statement corresponding to the first input statement;
[0011] In response to a matching instruction, perform matching analysis on several groups of the first input statements and the target output statements corresponding to the first input statements to obtain a matching score; the matching instruction is used to guide at least one of content matching analysis or scenario matching analysis;
[0012] Determine a reward value according to the matching score, and perform reinforcement learning training on the central control sub-model through the reward value to obtain a trained intelligent interaction model; the magnitude of the matching score is positively correlated with the magnitude of the reward value.
[0013] In addition, according to the training method of an intelligent interaction model in the above embodiments of the present application, the following additional technical features may also be provided:
[0014] Further, in an embodiment of the present application, the inputting the first input statement into the intelligent interaction model, and performing interaction prediction on the first input statement through each sub-interaction model includes:
[0015] Input the first input statement of the current round of interaction into the intelligent interaction model, and determine historical interaction information according to the interaction round of the current round of interaction; the historical interaction information includes the first input statements and target output statements of a preset number of rounds before the current round of interaction;
[0016] According to the historical interaction information, perform interaction prediction on the first input statement of the current round of interaction through each sub-interaction model.
[0017] Further, in an embodiment of the present application, the performing matching analysis on several groups of the first input statements and the target output statements corresponding to the first input statements includes:
[0018] Determine the interaction rounds where each group of the first input statements and the target output statements corresponding to the first input statements are located;
[0019] According to the preset matching rounds, select multiple groups of the first input statements and the target output statements corresponding to the first input statements with later interaction rounds for matching analysis.
[0020] Further, in an embodiment of the present application, the performing matching analysis on several groups of the first input statements and the target output statements corresponding to the first input statements to obtain a matching score includes:
[0021] Perform matching analysis on each group of the first input statements and the target output statements corresponding to the first input statements to obtain an initial score;
[0022] Weight the initial score according to the interaction round in which each group of the first input statement and the target output statement corresponding to the first input statement are located, to obtain the matching score; wherein, the weighting weight of the initial score of each group of the first input statement and the target output statement corresponding to the first input statement is positively correlated with the size of the interaction round.
[0023] Further, in an embodiment of the present application, the performing matching analysis on each group of the first input statement and the target output statement corresponding to the first input statement to obtain an initial score includes:
[0024] Extract the first feature information of the target output statement;
[0025] Perform semantic analysis according to the first feature information to obtain the text detection result of the target output statement; the text detection result is used to characterize whether the text content in the target output statement belongs to a natural language in a predetermined format;
[0026] Determine the initial score according to the text detection result.
[0027] Further, in an embodiment of the present application, the performing matching analysis on each group of the first input statement and the target output statement corresponding to the first input statement to obtain an initial score includes:
[0028] Obtain the standard interaction statement corresponding to the first input statement;
[0029] Extract the first feature information of the target output statement and extract the second feature information of the standard interaction statement;
[0030] Determine the similarity between the first feature information and the second feature information;
[0031] Determine the initial score according to the similarity; the initial score is positively correlated with the size of the similarity.
[0032] Further, in an embodiment of the present application, the first input statement carries a scene label; the performing matching analysis on each group of the first input statement and the target output statement corresponding to the first input statement to obtain an initial score includes:
[0033] Extract the first feature information of the target output statement;
[0034] Perform scene analysis according to the first feature information to obtain the scene detection result of the target output statement; the scene detection result is used to characterize the preset scene category to which the target output statement belongs;
[0035] Determine the initial score according to the scene detection result and the scene label.
[0036] Further, in an embodiment of the present application, the scenario detection result of the target output statement obtained by performing scenario analysis based on the first feature information includes:
[0037] Input the first feature information into an intent analysis model to obtain an intent prediction result output by the intent analysis model; the intent prediction result is used to characterize the interaction action category of the sub - interaction model for the first input statement, and the interaction action category at least includes asking questions and giving answers;
[0038] Determine the scenario detection result according to the intent prediction result.
[0039] Further, in an embodiment of the present application, the first input statement carries a label, and the sub - interaction model is trained through a training data set corresponding to the label;
[0040] Performing matching analysis on several groups of the first input statements and the target output statements corresponding to the first input statements to obtain a matching score, including:
[0041] Determine the target sub - interaction model determined by the central control sub - model according to the first input statement according to the target output statement corresponding to the first input statement;
[0042] Determine the matching score according to the matching relationship between the target sub - interaction model and the label of the first input statement.
[0043] Further, in an embodiment of the present application, performing matching analysis on several groups of the first input statements and the target output statements corresponding to the first input statements to obtain a matching score, including:
[0044] Performing content matching analysis on multiple groups of the first input statements and the target output statements corresponding to the first input statements to obtain a first score;
[0045] Performing scenario matching analysis on multiple groups of the first input statements and the target output statements corresponding to the first input statements to obtain a second score;
[0046] Weight the first score and the second score to obtain the matching score.
[0047] On the other hand, an embodiment of the present application provides an interaction method, and the method includes the following steps:
[0048] Collect voice data;
[0049] Perform speech recognition on the text content of the voice data to obtain a third input information;
[0050] Input the third input information into the intelligent interaction model obtained by training the above-mentioned intelligent interaction model training method to obtain the target output statement output by the intelligent interaction model;
[0051] Convert the target output statement into audio data for output.
[0052] On the other hand, an embodiment of the present application provides a training device for an intelligent interaction model. The intelligent interaction model includes a central control sub-model and multiple trained different sub-interaction models; the training device includes:
[0053] An acquisition module for acquiring first input information, where the first input information includes multiple first input statements;
[0054] A prediction module for inputting the first input statement into the intelligent interaction model, performing interaction prediction on the first input statement through each sub-interaction model to obtain multiple initial output statements, and determining a target sub-interaction model according to the first input statement through the central control sub-model, and determining the initial output statement output by the target sub-interaction model as the target output statement corresponding to the first input statement;
[0055] A scoring module for performing matching analysis on several groups of the first input statements and the target output statements corresponding to the first input statements in response to a matching instruction to obtain a matching score; the matching instruction is used to guide at least one of content matching analysis or scenario matching analysis;
[0056] An update module for determining a reward value according to the matching score, and performing reinforcement learning training on the central control sub-model through the reward value to obtain a trained intelligent interaction model; the matching score is positively correlated with the reward value.
[0057] On the other hand, an embodiment of the present application provides a computer device, including:
[0058] At least one processor;
[0059] At least one memory for storing at least one program;
[0060] When the at least one program is executed by the at least one processor, the at least one processor implements the above-mentioned intelligent interaction model training method or interaction method.
[0061] On the other hand, an embodiment of the present application further provides a computer-readable storage medium, in which a processor-executable program is stored, and the processor-executable program is used to implement the above-mentioned intelligent interaction model training method or interaction method when executed by a processor.
[0062] The advantages and beneficial effects of the present application will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present application:
[0063] A training method for an intelligent interaction model disclosed in an embodiment of the present application is applied to an intelligent interaction model including a central control sub-model and multiple sub-interaction models. The method obtains first input information including multiple first input statements, inputs the first input statements into the intelligent interaction model, performs interaction prediction on the first input statements through each sub-interaction model to obtain multiple initial output statements, and selects corresponding target output statements from the multiple initial output statements according to the first input statements through the central control sub-model, so as to obtain the interaction content of multiple rounds of interaction; then, performs matching analysis on several groups of first input statements and the target output statements corresponding to the first input statements to obtain a matching score; and determines a reward value according to the matching score, and performs reinforcement learning training on the central control sub-model through the reward value to obtain a trained intelligent interaction model. By analyzing the interaction content of multiple rounds of interaction at the content or scenario level and performing reinforcement learning training on the central control sub-model according to the analysis results, the method can effectively improve the performance of the central control sub-model in accurately selecting appropriate target output statements from multiple initial output statements, thereby improving the performance of the intelligent interaction model in processing multiple rounds of interaction tasks. When the trained intelligent interaction model is put into operation, the user experience can be effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following introduces the accompanying drawings of the related technical solutions in the embodiments of the present application or the prior art. It should be understood that the accompanying drawings in the following introduction are only for conveniently and clearly presenting some embodiments of the technical solutions in the present invention, and those skilled in the art can also obtain other accompanying drawings based on these drawings without creative efforts.
[0065] Figure 1 It is a schematic diagram of the implementation environment of a training method for an intelligent interaction model provided in an embodiment of the present application;
[0066] Figure 2 It is a schematic diagram of the interaction principle of an intelligent interaction model provided in an embodiment of the present application;
[0067] Figure 3 It is a schematic flowchart of a training method for an intelligent interaction model provided in an embodiment of the present application;
[0068] Figure 4 It is a schematic diagram for determining the reinforcement learning reward in a training method for an intelligent interaction model provided in an embodiment of the present application;
[0069] Figure 5Provided in the embodiments of the present application Figure 3 A specific flowchart of step 120 in
[0070] Figure 6 Provided in the embodiments of the present application Figure 3 A specific flowchart of step 130 in
[0071] Figure 7 Provided in the embodiments of the present application Figure 3 Another specific flowchart of step 130 in
[0072] Figure 8 A schematic diagram of content matching analysis in a training method of an intelligent interaction model provided in the embodiments of the present application
[0073] Figure 9 A schematic diagram of scenario matching analysis in a training method of an intelligent interaction model provided in the embodiments of the present application
[0074] Figure 10 A flowchart of an interaction method provided in the embodiments of the present application
[0075] Figure 11 A schematic diagram of the structure of a training device of an intelligent interaction model provided in the embodiments of the present application
[0076] Figure 12 A schematic diagram of the structure of a computer device provided in the embodiments of the present application Specific embodiments
[0077] The present application will be further described below in conjunction with the accompanying drawings of the specification and specific embodiments. The described embodiments should not be regarded as limitations of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0078] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.
[0079] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0080] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.
[0081] 1) Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0082] 2) Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. The natural language involved in this field is the language used by people in daily life, so it also has a close connection with the research of linguistics. Natural language processing technologies usually include technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs.
[0083] 3) Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence. Machine learning (deep learning) usually includes technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0084] 4) Reinforcement Learning (RL): Also known as re-inforcement learning, evaluation learning, or enhancement learning, it is one of the paradigms and methodologies of machine learning, used to describe and solve the problem of an agent achieving maximum reward or a specific goal through learning strategies during its interaction with the environment.
[0085] 5) Blockchain is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Essentially, blockchain is a decentralized database, a series of data blocks associated using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. Blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer. The blockchain underlying platform can include processing modules such as user management, basic services, smart contracts, and operation monitoring. Among them, the user management module is responsible for the identity information management of all blockchain participants, including maintaining the generation of public and private keys (account management), key management, and the maintenance of the correspondence between the real identity of the user and the blockchain address (permission management). And under authorization, it supervises and audits the transaction situations of certain real identities, providing the rule configuration for risk control (risk control audit); the basic service module is deployed on all blockchain node devices, used to verify the validity of business requests, and after reaching a consensus on valid requests, record them in storage. For a new business request, the basic service first performs interface adaptation parsing and authentication processing (interface adaptation), then encrypts the business information through the consensus algorithm (consensus management), and after encryption, transmits it intact and consistently to the shared ledger (network communication), and records and stores it; the smart contract module is responsible for the registration and issuance of contracts, as well as contract triggering and contract execution. Developers can define contract logic through a certain programming language, publish it to the blockchain (contract registration), trigger the execution according to the logic of the contract terms by calling keys or other events, complete the contract logic, and at the same time also provide functions for contract upgrade and cancellation; the operation monitoring module is mainly responsible for the deployment, configuration modification, contract setting, cloud adaptation during the product release process, and the visual output of the real-time state during product operation, such as: alarming, monitoring network conditions, monitoring the health status of node devices, etc. The platform product service layer provides the basic capabilities and implementation frameworks of typical applications. Developers can build on these basic capabilities and overlay the characteristics of the business to complete the blockchain implementation of the business logic. The application service layer provides application services based on the blockchain solution for business participants to use.
[0086] 6) Conversational AI refers to intelligent behaviors manifested through conversations and interactions. Usually, intelligent systems interact with users or the environment and achieve learning and modeling during the interaction. It mainly includes, but is not limited to, the following aspects of research: general question-answering systems, including automatic question answering, reading comprehension, etc.; task or goal-oriented dialogue systems; open-domain chat systems. Among them, the general question-answering system aims to find accurate information from structured (such as knowledge bases, tables) and unstructured (such as documents) to answer users' questions; task or goal-oriented dialogue systems need to achieve a specific task or goal through interaction, such as various intelligent assistants, ticket-booking, and meal-ordering systems; open-domain chat systems focus on chatting with users, emotional communication, and companionship, which are important bases and prerequisites for social robots to enter thousands of households. These interactive systems not only use natural language as a carrier but also comprehensively apply multimedia information such as images and voices, enabling machines to understand the environment they are in and exhibit intelligent behaviors that conform to the situation.
[0087] In recent years, with the rapid development of artificial intelligence technology, various types of machine learning models have achieved relatively good application effects in fields such as image classification, face recognition, and autonomous driving. Among them, in the field of human-computer interaction, an intelligent interaction model can be built through artificial intelligence technology to collect relevant information and process tasks based on human-computer conversations. This technology is called Conversational AI.
[0088] In related technologies, when implementing human-computer interaction based on conversational AI, the built intelligent interaction model performs well when dealing with single-round conversations. For example, the general question-answering system that finds accurate information from structured (such as knowledge bases, tables) and unstructured (such as documents) to answer users' questions can relatively easily identify the content required by users from a large amount of information. However, when facing the task requirements of multi-round conversations, the performance of the intelligent interaction model often fails to meet expectations. On the one hand, due to the large amount of data processing in multi-round conversation tasks, it is difficult to train the intelligent interaction model. On the other hand, many current intelligent interaction models in the industry rely on rule-based dialogue strategies, and the stacking and judgment of rules become extremely complex during multi-round conversations. The learning ability and scalability of the intelligent interaction model are relatively poor. These factors above lead to problems such as poor context coherence and low dialogue quality when the intelligent interaction model processes a large number of interaction rounds, which has a very adverse impact on the application of task or goal-oriented dialogue systems and open-domain chat systems, resulting in poor interaction effects and a bad user experience.
[0089] In order to solve the problems in the related art that when the number of interaction rounds of the intelligent interaction model is large, the context coherence is often poor, the dialogue quality is low, etc., resulting in poor interaction effects and affecting the user experience, the embodiments of the present application provide a training method, an interaction method, a device and a device for an intelligent interaction model. The training method is applied to an intelligent interaction model including a central control sub-model and multiple sub-interaction models. By obtaining first input information including multiple first input statements, inputting the first input statements into the intelligent interaction model, performing interaction prediction on the first input statements through each sub-interaction model to obtain multiple initial output statements, and selecting corresponding target output statements from the multiple initial output statements according to the first input statements through the central control sub-model, the interaction content of multiple rounds of interaction is obtained; then, matching analysis is performed on several groups of first input statements and the target output statements corresponding to the first input statements to obtain a matching score; and a reward value is determined according to the matching score, and the central control sub-model is trained by reinforcement learning through the reward value to obtain a trained intelligent interaction model. This training method analyzes the interaction content of multiple rounds of interaction at the content or scenario level, and trains the central control sub-model by reinforcement learning according to the analysis results, which can effectively improve the performance of the central control sub-model to accurately select appropriate target output statements from multiple initial output statements, and further improve the performance of the intelligent interaction model in processing multiple rounds of interaction tasks. When the trained intelligent interaction model is put into operation, the user experience can be effectively improved.
[0090] Figure 1 It is a schematic diagram of the implementation environment of a training method for an intelligent interaction model provided by an embodiment of the present application. Refer to Figure 1 , the main software and hardware entities of this implementation environment mainly include an operation terminal 101 and a server 102, and the operation terminal 101 is communicatively connected to the server 102. Among them, the training method of this intelligent interaction model can be configured to be executed by the operation terminal 101 alone, or can be configured to be executed by the server 102 alone, or can be executed based on the interaction between the operation terminal 101 and the server 102. Specifically, it can be appropriately selected according to the actual application situation, and this embodiment does not make specific limitations on this. In addition, the operation terminal 101 and the server 102 can be nodes in the blockchain, and this embodiment does not make specific limitations on this.
[0091] Specifically, the operation terminal 101 in the present application may include, but is not limited to, any one or more of a smart watch, a smart phone, a computer, a personal digital assistant (PDA), a smart voice interaction device, a smart home appliance, or a vehicle-mounted terminal. The server 102 may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. A communication connection may be established between the operation terminal 101 and the server 102 through a wireless network or a wired network. The wireless network or the wired network uses standard communication technologies and / or protocols. The network may be set to the Internet or any other network, such as any combination including, but not limited to, a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network, or a virtual private network.
[0092] Before describing the training method of the intelligent interaction model provided in the embodiments of the present application, first, the composition structure and application principle of the intelligent interaction model in the present application are introduced.
[0093] Please refer to Figure 2 , Figure 2 which is a schematic diagram of the interaction principle of the intelligent interaction model in the present application. Specifically, the intelligent interaction model includes a central control sub-model and multiple different sub-interaction models. Figure 2 The intelligent interaction model shown in includes three sub-interaction models, namely sub-interaction model A, sub-interaction model B, and sub-interaction model C. In an actual intelligent interaction model, the number of sub-interaction models may be any integer greater than or equal to 2. In the intelligent interaction model, each sub-interaction model is used to perform interaction prediction on the input information to obtain the interaction information corresponding to each sub-interaction model. Generally speaking, the interaction process is mainly completed in the form of a statement dialogue, that is, the input information may include multiple statements. For each input statement, each sub-interaction model can give a corresponding interaction statement. For example Figure 2Among them, the input information, in the order of interaction, successively includes statement S1, statement S2, and statement S3. The sub-interaction models A, B, and C respectively give corresponding interaction contents for these input statements. Among them, the interaction content of sub-interaction model A for statement S1 is statement A1, the interaction content for statement S2 is statement A2, and the interaction content for statement S3 is statement A3. Sub-interaction models B and C are similar and will not be elaborated here. The central control sub-model is used to select appropriate statements from the interaction contents given by each of the sub-interaction models A, B, and C according to the statements in the input information as the interaction output of the intelligent interaction model for the input information. For example, Figure 2 Among them, for statement S1 in the input information, the central control sub-model selects statement B1 output by sub-interaction model B. For statement S2 in the input information, the central control sub-model selects statement C2 output by sub-interaction model C. For statement S3 in the input information, the central control sub-model selects statement A3 output by sub-interaction model A as the interaction content, so as to form a multi-round continuous dialogue of "statement S1 - statement B1 - statement S2 - statement C2 - statement S3 - statement A3".
[0094] The above is a brief description of the composition structure and application principle of the intelligent interaction model in this application. In this intelligent interaction model, the central control sub-model and each sub-interaction model can be built based on machine learning algorithms in the field of artificial intelligence. The machine learning model structures selected for the central control sub-model and each sub-interaction model in this application are not limited.
[0095] It should be noted that in order to simulate the normal interaction process well, when the central control sub-model in this application selects the output statement of the sub-interaction model corresponding to the input statement, it can determine the output statement by comprehensively considering the previous interaction content. For example, Figure 2 Among them, for input statement S2, the central control sub-model can simultaneously consider the previous dialogue "statement S1 - statement B1" before statement S2, and then combine the content of statement S2 to determine the output statement. Of course, in order to reduce the pressure of data processing, it can be set that the central control sub-model determines the output statement of this time based on the current input statement and the interaction content of several previous rounds. The specific number of rounds can be flexibly set according to needs.
[0096] It should be noted that each component in the intelligent interaction model in the embodiment of this application can be integrally arranged in one place or can adopt a distributed layout. For example, the sub-interaction model in the intelligent interaction model can be a node in the blockchain. This embodiment does not make specific limitations on this.
[0097] Figure 3It is a flowchart of a method for training an intelligent interaction model provided by an embodiment of the present application. The execution subject of this method can be at least one of an operation terminal or a server. Figure 3 Taking the case where the training method of the intelligent interaction model is configured on the operation terminal as an example for description. Refer to Figure 3 , the training method of the intelligent interaction model includes but is not limited to steps 110 to 140.
[0098] Step 110: Obtain first input information, where the first input information includes a plurality of first input statements.
[0099] In this step, when training the central control sub-model of the intelligent interaction model, obtain the training data to be input into the intelligent interaction model, denoted as the first input information. The first input information includes a plurality of statements, denoted as the first input statements. In the present application, the source channel for obtaining the first input information is not limited. For example, in some embodiments, the first input information can be downloaded from a relevant resource server, or transmitted through a hardware port, or obtained from the environment through a voice collection and recognition device.
[0100] It should be noted that in the first input information of the present application, each first input statement is logically separated in semantics, that is, other interaction contents can be inserted between each first input statement in the first input information. And, in some embodiments, each first input statement can have a temporal order in semantic logic, and can be arranged in sequence according to this temporal order in advance. As an optional way to construct the first input information, the speech record of a certain participant can be intercepted from normal conversation content, so as to organize and obtain the first input information.
[0101] It should be noted that when obtaining the first input information in this application, it can be obtained in advance at one time, or the first input statement can be obtained each time in combination with the target output statement in the subsequent step 120. In other words, in some embodiments, the central control sub-model can be trained using pre-set training data, which is the first input information containing multiple first input statements. Each first input statement is pre-set and does not change due to the target output statement given by the intelligent interaction model for each first input statement. This training method is friendly to the acquisition process of training data and can significantly improve the training efficiency of the model. In other embodiments, when training the central control sub-model, in each interaction round, the first input statement obtained can be determined according to the target output statement given by the intelligent interaction model for the previous first input statement, that is, the first input statement is not pre-set but selected according to the content requirements during the interaction process. Although this training method increases the processing complexity of the training data to a certain extent, it can more accurately continuously simulate the entire interaction process. The first input statements in the training data will not deviate significantly from the interaction theme, which can effectively improve the stability of training.
[0102] Step 120: Input the first input statement into the intelligent interaction model, perform interaction prediction on the first input statement through each sub-interaction model to obtain multiple initial output statements, and determine the target sub-interaction model according to the first input statement through the central control sub-model, and determine the initial output statement output by the target sub-interaction model as the target output statement corresponding to the first input statement.
[0103] In this step, the first input statement in the first input information is input into the intelligent interaction model for interaction prediction. After receiving the input first input statement, the intelligent interaction model can perform interaction prediction on the first input statement through each of its sub-interaction models respectively to obtain the initial output statement corresponding to each sub-interaction model. Then, the intelligent interaction model determines the target sub-interaction model according to the first input statement through the central control sub-model, and takes the initial output statement output by the target sub-interaction model as the target output statement corresponding to the first input statement selected from these initial output statements, and then takes the target output statement as the interaction output of the intelligent interaction model for the first input statement.
[0104] It should be noted that since the first input information includes multiple first input statements, for the first input information, the foregoing step 120 can be executed multiple times. That is, for each input statement in the first input information, step 120 can be executed once, so as to obtain the target output statements corresponding to each first input statement in the first input information. Of course, in some embodiments, for the first input information obtained in advance at one time, some first input statements can also be selected to execute step 120, and the specific selection method can be determined according to the number of statements or the proportion in the first input information. It should be noted that when selecting some first input statements to participate in the training, generally the original word order of each first input statement can be maintained, so that the intelligent interaction model can better generate continuous multi-round interaction content according to the input first input statement.
[0105] It should be noted that when the central control sub-model determines the target sub-interaction model according to the first input statement, it can first perform serial number encoding on each sub-interaction model, and then determine the corresponding target sub-interaction model from the encoded data output by the central control sub-model. In some embodiments, for example, the first input statement can be processed as feature information and input into the central control sub-model. The central control sub-model performs calculation or mapping processing on the feature information and can output a group of vectors to represent the target sub-interaction model determined by the central control sub-model according to the first input statement. Taking Figure 2 the intelligent interaction model in as an example, for example, the vector (0, 1, 0) is output for the statement S1, which means that the central control sub-model determines the second sub-interaction model, that is, sub-interaction model B, as the target sub-interaction model for the statement S1, so that the statement B1 can be determined as the target output statement corresponding to the statement S1; of course, the data format output by the central control sub-model can be flexibly set according to needs, and this application does not limit this.
[0106] Step 130: In response to the matching instruction, perform matching analysis on several groups of first input statements and the target output statements corresponding to the first input statements to obtain a matching score; the matching instruction is used to guide at least one of content matching analysis or scenario matching analysis.
[0107] In this step, as described above, after the intelligent interaction model performs interaction prediction on multiple first input statements in the first input information, the target output statement corresponding to each first input statement can be obtained. At this time, each first input statement and the target output statement corresponding to the first input statement can form a round of interaction content. In this application, the obtained multiple groups of first input statements and the target output statements corresponding to the first input statements can be subjected to matching analysis, that is, the obtained multiple rounds of interaction content can be subjected to matching analysis to obtain a matching score. Specifically, when determining the matching score, for a round of interaction content, that is, each first input statement and the target output statement corresponding to the first input statement, they can be first subjected to matching analysis to obtain the initial score corresponding to this round of interaction content, and then the initial scores corresponding to multiple rounds of interaction content are weighted to obtain the matching score evaluated on the whole of multiple rounds of interaction content.
[0108] It should be noted that in this application, the matching analysis of the first input statement and the target output statement can be performed in response to a matching instruction. The matching instruction can be set in advance by the user before the intelligent interaction model is trained, or can be temporarily output. The matching instruction can be used to guide the content matching analysis or the scenario matching analysis of the first input statement and the target output statement separately, or can be used to guide the content matching analysis and the scenario matching analysis of the first input statement and the target output statement. In other words, when performing the matching analysis on the first input statement and the target output statement corresponding to the first input statement, in some embodiments, the content matching analysis of the first input statement and the target output statement can be separately selected to determine whether there is a non-corresponding situation at the content level between the two. If so, the matching score can be determined as a lower value; otherwise, if not, the matching score can be determined as a higher value. In some embodiments, the scenario matching analysis of the first input statement and the target output statement can be separately selected to determine whether there is a non-matching situation in the interaction context between the two. Similarly, if so, the matching score can be determined as a lower value; otherwise, if not, the matching score can be determined as a higher value. In some embodiments, the content matching analysis and the scenario matching analysis can also be performed on the first input statement and the target output statement respectively, and then the matching score is determined according to the matching conditions in the two aspects.
[0109] Step 140: Determine the reward value according to the matching score, and perform reinforcement learning training on the central control sub-model through the reward value to obtain a trained intelligent interaction model.
[0110] In this step, the matching score determined in the previous step 130 can be used to quantitatively represent the strength of the multi-round interaction ability of the intelligent interaction model. In other words, the matching score can reflect the coherence and stability of the intelligent interaction model in achieving effective interaction when facing complex interaction requirements. The higher the matching score, the better the intelligent interaction model can perform in multi-round interaction. On the contrary, the lower the matching score, the worse the intelligent interaction model can perform in multi-round interaction. Combining the foregoing Figure 2 and the description of the interaction principle of the intelligent interaction model of the present application, it can be understood that the performance of the intelligent interaction model mainly depends on the performance of each sub-interaction model and the central control sub-model. More specifically, it depends on the interaction prediction performance of each sub-interaction model and the performance of the central control sub-model in accurately selecting a suitable target output statement from multiple initial output statements. For the former, the interaction prediction performance of each sub-interaction model is related to the structure and training method of its own model, and the present application does not discuss this; while the performance of the central control sub-model in accurately selecting a suitable target output statement from multiple initial output statements is exactly the main problem to be solved by the training method of the intelligent interaction model proposed in the present application. That is, in the present application, the performance of the central control sub-model in accurately selecting a suitable target output statement from multiple initial output statements can be improved through the training of the central control sub-model, thereby improving the performance of the intelligent interaction model in processing multi-round interaction tasks.
[0111] Specifically, in this step, the central control sub-model can be trained based on reinforcement learning. The basic principle of reinforcement learning is that if a certain behavior strategy of an agent (i.e., the central control sub-model in the present application) causes the environment to generate a positive reward (reinforcement signal), then the tendency of the agent to generate this behavior strategy in the future will be strengthened. The goal of reinforcement learning is to find the optimal strategy at each discrete state to maximize the expected discounted reward sum. This algorithm regards learning as a trial and evaluation process. An action is taken on the environment, and the state of the environment changes after accepting the action. At the same time, a reinforcement signal (reward or punishment) is generated and fed back to the agent. The agent then selects the next action based on the reinforcement signal and the current state of the environment. The selection principle is to increase the probability of receiving positive reinforcement (reward). For the present application, every time the central control sub-model selects the target output statements corresponding to multiple first input statements, it is equivalent to performing an action on the environment. And the matching score obtained by performing matching analysis on several groups of first input statements and the target output statements corresponding to the first input statements can be considered equivalent to the reinforcement signal feedback by the environment. If the matching score of the current time becomes larger, it means that the target output statement currently selected by the central control sub-model is better; if the matching score of the current time becomes smaller, it means that the target output statement currently selected by the central control sub-model is worse. Therefore, the reinforcement signal feedback by the environment, that is, the reward value of reinforcement learning, can be determined according to the matching score, and the size of the reward value should be positively correlated with the size of the matching score.
[0112] Based on the above principle, referring to Figure 4 , in this application, the process of the central control sub-model cyclically selecting a corresponding target output statement from multiple initial output statements according to the first input statement, and performing matching analysis on several groups of the first input statements and the target output statements corresponding to the first input statements to obtain a matching score can be executed. For example, Figure 4 , in Figure 4 , for the first input statement S1 in the input information, the corresponding target output statement is statement B1, for the first input statement S2 in the input information, the corresponding target output statement is statement C2, and for the first input statement S3 in the input information, the corresponding target output statement is statement A3. Then, the interaction content of the first round, the interaction content of the second round, and the interaction content of the third round can be scored respectively. After obtaining the scores of the three rounds of interactions, the total matching score is determined by weighting, and then the reward value of the reinforcement learning is determined according to the matching score, and the central control sub-model is trained based on this reward value. So that the central control sub-model can effectively learn and update its own parameters in the direction of selecting and outputting more appropriate target output statements, thereby obtaining a trained central control sub-model. Matching with each trained sub-interaction model, a trained intelligent interaction model can be obtained.
[0113] It should be noted that in the embodiments of this application, the interaction content of each round does not always start with the first input statement, and can be flexibly selected according to needs. For example, in Figure 4 , there is also a context relationship between statement B1 and statement S2 itself, and they can also be used as the interaction content of one round alone, and the subsequent interaction content can be selected by analogy.
[0114] In the foregoing embodiments, the training process of the central control sub-model of the intelligent interaction model of the present application was mainly described. It was defaulted that in this process, each sub-interaction model of the intelligent interaction model was a trained sub-interaction model. Here, the trained sub-interaction model can be obtained from other interaction models or can be self-trained. In the embodiments of the present application, the sub-interaction model can perform interaction prediction on the first input statement to obtain multiple initial output statements. There is no specific limitation on the type of the sub-interaction model. For example, in some embodiments, the sub-interaction model can be a retrieval-type model, which is preset with a statement library. For the input first input statement, the sub-interaction model can, according to the feature information obtained from the first input statement, obtain an output result that can represent a certain statement in the statement library through calculation or mapping, and then determine the statement corresponding to the statement library as the initial output statement according to the output result; in some embodiments, the sub-interaction model can also be a generation-type model, such as a Seq2Seq model adopting an encoder-decoder structure. For the input first input statement, the first input statement can be first converted into sequence data. The encoder in the model can encode sequence data of any length into a vector, and the decoder can then automatically generate a corresponding context vector according to the vector and decode it into sequence data. The text content output by the model can be further obtained according to the sequence data output by the decoder, that is, the initial output statement output by the sub-interaction model.
[0115] In some embodiments, for the training process of the intelligent interaction model, it may also include the training of the sub-interaction model. Therefore, the method in the present application may further include but is not limited to steps 150 to 180:
[0116] Step 150: Obtain second input information, where the second input information includes multiple second input statements and labels corresponding to the second input statements; the labels include at least one of content labels or scene labels;
[0117] Step 160: Input the second input statement into the sub-interaction model, and perform interaction prediction on the second input statement through the sub-interaction model to obtain a predicted output statement;
[0118] Step 170: Determine the training loss value according to the predicted output statement and the label;
[0119] Step 180: Train the sub-interaction model according to the loss value to obtain a trained sub-interaction model.
[0120] In the embodiments of the present application, for each sub - interaction model in the intelligent interaction model, common machine learning algorithms can be used for construction, and they can be trained based on supervised learning. Here, they can be trained with a labeled training data set, and the training data in the training data set can also be information such as statements. For example, obtain the information used to train the sub - interaction model, denoted as the second input information. The second input information can also include multiple input statements, and each input statement is denoted as the second input statement. For each second input statement, it can carry a label, and the label can be at least one of a content label or a scene label. Among them, the content label can be used to characterize the content that the sub - interaction model needs to perform interaction prediction output for the second input statement. The specific form of the content label is not limited here. In some embodiments, the content label can include multiple appropriate optional statements, or the content label can also include a certain (or multiple) keyword(s) used to limit the content that the sub - interaction model needs to perform interaction prediction output. Similarly, the scene label can be used to characterize the context scene of the content output when the sub - interaction model needs to perform interaction prediction for the second input statement. For example, when the scene label is the "reply scene", the context of the statement content output by the sub - interaction model should be the reply scene. If the context of the statement content output by the sub - interaction model at this time is the question scene (such as the statement output by the sub - interaction model is a question sentence), it means that there are areas for improvement in the interaction prediction of the sub - interaction model at the context scene level, and this problem can be improved by adjusting the model parameters.
[0121] In this application, during the training process of the sub-interaction model, the second input statement can be input into the initialized sub-interaction model. In some embodiments, if the data format of the second input statement is in text format, it can be encoded and converted before being input into the model to convert the unstructured text data into structured data that is easy for the model to process. For example, the second input statement can be tokenized to obtain the phrases that make up the statement. Here, there are various tokenization algorithms that can be used. For example, in some embodiments, a dictionary-based tokenization algorithm can be used. First, the statement is segmented into words according to the dictionary, and then the best combination of words is found. In some embodiments, a character-based tokenization algorithm can also be used. First, the statement is divided into individual characters, and then the characters are combined into words to find the optimal combination. After the statement is tokenized, the word embedding vector corresponding to each word in the phrase can be determined through a pre-established dictionary. Of course, in some embodiments, the word embedding vector can be obtained by mapping the word to a vector space with a unified lower dimension. The strategies for generating such a mapping include neural networks, dimensionality reduction of word co-occurrence matrices, probability models, and interpretable knowledge base methods, etc. Taking the word embedding vector as an example of the structured data obtained by encoding the word, after obtaining the word embedding vector corresponding to each word in the second input statement, these word embedding vectors can be accumulated. The accumulated vector can be denoted as the phrase vector. After normalizing the phrase vector, the vector corresponding to the second input statement can be obtained. For example, when normalizing, it can be set that the sum of the elements in the vector corresponding to the statement is 1.
[0122] For the second input statement data input into the sub-interaction model, its feature information can be extracted, and interaction prediction can be performed based on the feature information to obtain the predicted output statement output by the sub-interaction model. After obtaining the predicted output statement, the accuracy of the prediction of the machine learning model can be evaluated according to the predicted output statement and the foregoing label, so as to perform backpropagation training on the model and update its internal relevant parameters. Specifically, for a machine learning model, its prediction accuracy can be measured by a loss function. The loss function is defined on a single training data and is used to measure the prediction error of a training data. Specifically, the loss value of a training data is determined by the label of the single training data and the prediction result of the model on the training data. During actual training, a training data set has many training data. Therefore, a cost function is generally used to measure the overall error of the training data set. The cost function is defined on the entire training data set and is used to calculate the average value of the prediction errors of all training data, which can better measure the prediction effect of the model. For a general machine learning model, based on the foregoing cost function, plus a regularization term that measures the complexity of the model, it can be used as the objective function for training. Based on this objective function, the loss value of the entire training data set can be obtained. There are many types of commonly used loss functions. For example, the 0-1 loss function, the squared loss function, the absolute loss function, the logarithmic loss function, the cross-entropy loss function, etc. can all be used as the loss function of the machine learning model, which will not be elaborated one by one here. In the embodiments of the present application, any one of the loss functions can be selected to determine the loss value of the training, that is, the loss value between the predicted output statement and the label. Based on the loss value of the training, the parameters of the model are updated using the backpropagation algorithm, and the trained sub-interaction model can be obtained by iterating a preset number of rounds.
[0123] It should be noted that the above training process of the sub-interaction model is only used to exemplify an optional implementation scheme of the sub-interaction model in the intelligent interaction model, and does not mean to limit its specific training method. It can be understood that each sub-interaction model of the intelligent interaction model in this application can be trained by different training methods, and there can be differences in the structure, training algorithm, and training data set of the model itself. For example, in some embodiments, each sub-interaction model can be built and trained according to functional requirements. For example, in the intelligent interaction model, one (or more) sub-interaction models can be trained specifically for the requirements in the reply scenario, and another (or more) sub-interaction models can be trained specifically for the requirements in the question scenario. In this way, on the one hand, it can enable a single sub-interaction model to focus on specific task requirements, greatly reducing its own training difficulty and improving the adaptability of the intelligent interaction model to various interaction contents and scenarios; on the other hand, it can reduce the functional overlap of the sub-interaction models in the intelligent interaction model, reduce unnecessary resource waste, and modularize the functions of the sub-interaction models, which is convenient for the construction and splitting of the intelligent interaction model.
[0124] Referring to Figure 5 As shown, in an embodiment of the present application, step 120 is further described. Step 120 may include but is not limited to step 121 and step 122:
[0125] Step 121: Input the first input statement of the current round of interaction into the intelligent interaction model, and determine the historical interaction information according to the interaction round of the current round of interaction; the historical interaction information includes the first input statements and target output statements of a preset number of rounds before the current round of interaction;
[0126] Step 122: According to the historical interaction information, perform interaction prediction on the first input statement of the current round of interaction through each sub-interaction model.
[0127] In the embodiments of the present application, in order to simulate the normal interaction process well, when the sub-interaction model makes an interaction prediction to output the initial output statement according to the first input statement, it can also be determined by integrating the previous interaction content. Specifically, at this time, for each round of interaction process, the first input statement in this round of interaction is input into the intelligent interaction model, and the model can obtain the historical interaction information before this round of interaction. The historical interaction information may include the first input statements in several rounds before this round of interaction and the target output statements determined for these first input statements. Here, the selected several rounds in the historical interaction information may be preset rounds. If the total number of rounds of the current interaction does not exceed the preset rounds, then all the current interaction content is taken as the historical interaction information. For example, suppose the current intelligent interaction model and the user are in the 11th round of interaction, and the user has previously input the first input statement 10 times, and the intelligent interaction model has given the target output statement corresponding to the first input statement. Then, assuming that the preset round is 5, during the 11th round of interaction, after the user inputs the first input statement, the intelligent interaction model first determines that the round number of this round of interaction is 11, and then determines the historical interaction information according to the preset round, that is, it can be determined that the historical interaction information includes all the content during the 6th round to the 10th round of interaction between the intelligent interaction model and the user. Thus, the intelligent interaction model can use the historical interaction information as part of the reference data, and through each sub-interaction model, make an interaction prediction for the first input statement in the 11th round of this interaction, so as to determine the initial output statements that each needs to output in the 11th round.
[0128] It should be noted that in the present application, there is no limitation on the data application method of the historical interaction information in the interaction prediction process. In some embodiments, the feature information of these historical interaction information can be extracted, and the feature information of these historical interaction information and the feature information of the first input statement in this round of interaction are input into the sub-interaction model for interaction prediction together; in some embodiments, before inputting into the sub-interaction model, the feature information of the first input statement in this round of interaction can also be fused according to the feature information obtained by extracting the historical interaction information to obtain the fused feature information, and then the fused feature information is input into the sub-interaction model for interaction prediction. The specific fusion methods of the feature information may include data splicing, data weighting, etc.
[0129] It can be understood that in the present application, during each round of interaction process, the sub-interaction model can not only perform interaction prediction based on the first input statement of the current round, but also combine the historical interaction information of several rounds before the current round as a reference, so as to effectively connect the context information in the interaction process and improve the coherence and fluency of multi-round interaction. At the same time, in order to reduce the pressure of data processing, a parameter of a preset number of rounds can be set in the present application to prevent the situation that the sub-interaction model needs to process a large amount of redundant data when the number of interaction rounds is too large, and improve the resource utilization rate and data processing efficiency. Specifically, the present application does not limit the actual set value of the preset number of rounds, and it can be flexibly set according to actual needs.
[0130] Referring to Figure 6 as shown, in an embodiment of the present application, step 130 is further described. Step 130 may include but is not limited to step 131 and step 132:
[0131] Step 131: Determine the interaction rounds in which each group of the first input statements and the target output statements corresponding to the first input statements are located;
[0132] Step 132: According to the preset matching rounds, select multiple groups of the first input statements with later interaction rounds and the target output statements corresponding to the first input statements for matching analysis.
[0133] In the embodiment of the present application, combined with the foregoing analysis and elaboration of Figure 5 , it can be known that when each sub-interaction model in the present application predicts the initial output statement in the interaction, it not only performs interaction prediction based on the first input statement of the current round, but also combines the historical interaction information of several rounds before the current round as a reference. Then, it can be understood that the target initial statement further selected by the central control sub-model is also an output obtained by integrating the historical interaction information. In step 130 of the present application, it is responsible for performing matching analysis on the obtained multi-round interaction content to obtain a matching score, so as to quantify the multi-round interaction ability of the intelligent interaction model. The multi-round interaction ability of the intelligent interaction model is mainly reflected in the coherence and stability of the interaction. When determining the matching score, in order to reduce the complexity and amount of data processing, the interaction content of some interaction rounds can be selected from all the interaction rounds, and the matching score of these interaction contents is judged as the global matching score.
[0134] It can be understood that for the overall interaction content, the later interaction rounds can better reflect the processing performance of the intelligent interaction model for multi-round interaction tasks. Therefore, in this application, when determining the matching score, the part of the content participating in the evaluation can be first selected from all the interaction content, that is, the interaction rounds where the corresponding first input statement and target output statement of each group can be determined. Then, according to the pre-set number of interaction rounds required to participate in the matching (denoted as the matching rounds), multiple groups of first input statements and target output statements with later interaction rounds are selected for matching analysis. For example, during a certain training process, a total of 20 rounds of interaction simulations are performed on the intelligent interaction model, that is, 20 first input statements are sequentially input to the intelligent interaction model, and the model outputs 20 target output statements for each first input statement. Among them, the interaction round where the first input first input statement and its corresponding target output statement are located is the 1st round, and the interaction round where the last input first input statement and its corresponding target output statement are located is the 20th round. Assuming that the pre-set matching rounds are 10 rounds, the interaction content in the interaction rounds from the 11th round to the 20th round can be selected as the evaluation content, and the first input statements and target output statements in these interaction rounds are subjected to matching analysis. In this way, both the consumption of computing resources can be reduced, and the performance of the intelligent interaction model in processing multi-round interaction can be more accurately evaluated.
[0135] Referring to Figure 7 As shown, in an embodiment of the present application, step 130 is further described. Step 130 may further include but is not limited to step 133 and step 134:
[0136] Step 133: Perform matching analysis on each group of first input statements and the target output statements corresponding to the first input statements to obtain an initial score;
[0137] Step 134: Weight the initial score according to the interaction rounds where each group of first input statements and the target output statements corresponding to the first input statements are located to obtain a matching score; wherein, the weighting weight of the initial score of each group of first input statements and the target output statements corresponding to the first input statements is positively correlated with the size of the interaction rounds.
[0138] In the embodiment of the present application, in combination with the foregoing description of Figure 6From the above analysis and elaboration, it can be known that when determining the matching score, some parts of the interaction content involved in the evaluation can be first selected from all the interaction content. For example, multiple sets of the first input statements and target output statements at the later interaction rounds can be selected for matching analysis. During the specific matching analysis, according to the corresponding interaction rounds, the first input statement and the target output statement in each group are subjected to matching analysis, and the initial score corresponding to the interaction content of this group can be obtained. Then, based on the initial scores corresponding to the first input statements and target output statements of each group, the overall matching score of the selected evaluation content can be determined. For example, in some embodiments, the average value of the initial scores corresponding to the first input statements and target output statements of each group can be calculated, and the obtained average value can be used as the matching score. In this way of obtaining the matching score, the weights of the interaction content in each round are the same. However, generally speaking, due to the continuous progress of the interaction, in the later interaction rounds, the information that the intelligent interaction model needs to comprehensively consider, the requirements for the coherence of the interaction, etc. are more complex, and it is more difficult to give a good target output statement. Therefore, in order to better determine the ability of the intelligent interaction model to handle multi-round interactions, the content in the later interaction rounds should be emphasized.
[0139] Therefore, in the embodiments of the present application, when determining the overall matching score of the evaluation content according to the initial score, the weighted weights of the initial scores corresponding to each group of the first input statements and target output statements can be first determined according to the interaction rounds in which they are located, where the weighted weights of the initial scores of each group of the first input statements and target output statements are positively correlated with the magnitudes of the interaction rounds in which they are located. That is, for the first input statement and the target output statement with a larger value of the interaction round, the weighted weights of their corresponding initial scores are also larger. On the contrary, for the first input statement and the target output statement with a smaller value of the interaction round, the weighted weights of their corresponding initial scores are also smaller. In this way, the intelligent interaction model can pay more attention to the coherence and stability of multi-round interactions, and can improve the performance of the trained intelligent interaction model in handling multi-round interaction tasks.
[0140] Next, in combination with the accompanying drawings and some embodiments, the implementation manner of step 133 in the present application will be specifically described.
[0141] Referring to Figure 8 , when performing matching analysis on the first input statement and the target output statement in step 133 of the present application to obtain the initial score, the analysis can be performed based on the matching relationship between the first input statement and the target output statement at the content level. For example, Figure 8 shows a situation where in a certain round of interaction, the same first input statement "The weather is nice today" is input, and the intelligent interaction model outputs different target output statements.
[0142] In the first case, the target output statement output by the intelligent interaction model is "Yes, it's very suitable for playing", which well connects with the first input statement at the content level and can form a relatively smooth and coherent round of interaction logically, facilitating the continuous progress of the interaction. In this case, it shows that the current central control sub-model of the intelligent interaction model has good content judgment performance in processing this round of interaction and can select an appropriate target output statement. Therefore, a relatively high initial score can be given for the content of this round of interaction.
[0143] In the second case, the target output statement output by the intelligent interaction model is "I've had breakfast", which has almost no connection with the first input statement "The weather is nice today" at the content level, resulting in very poor logic in the interaction content and affecting the further progress of the interaction. In this case, it shows that the current central control sub-model of the intelligent interaction model has poor content judgment performance in processing this round of interaction and it is difficult to select an appropriate target output statement. Therefore, a relatively low initial score can be given for the content of this round of interaction. The above is an overview of the implementation principle of content matching analysis in this application. Next, some specific implementations of content matching analysis will be described in detail.
[0144] In some embodiments, step 133 may include but is not limited to steps 1330 to 1332:
[0145] Step 1330: Extract the first feature information of the target output statement;
[0146] Step 1331: Perform semantic analysis based on the first feature information to obtain the text detection result of the target output statement; the text detection result is used to represent whether the text content in the target output statement belongs to a natural language in a predetermined format;
[0147] Step 1332: Determine the initial score according to the text detection result.
[0148] In the embodiments of the present application, when performing matching analysis on each group of first input statements and the target output statements corresponding to the first input statements to obtain an initial score, it can be determined based on whether the semantics of the target output statement itself conforms to natural language in a predetermined format. Among them, the predetermined format here is generally the same type of language as the first input statement. For example, generally, when the first input statement is in Chinese format, it is expected that the model can analyze and give a target output statement in Chinese format so that when the intelligent interaction model is put into operation, it can effectively complete the interaction according to the way given by the user and reduce the situation of mismatched interaction content. Of course, the predetermined format can also include personalized settings for rules such as the type of language, speech rate, or grammar collocation. In addition, in the present application, the semantic content of the target output statement is also detected, which can effectively measure whether the target output statement given by the intelligent interaction model is natural language. Here, natural language refers to the language that conforms to the daily usage rules of people, which can facilitate determining whether there are problems such as meaningless or unconventional grammar in the target output statement that lead to content mismatches in semantic logic. For these semantic-level matching analyses, in the embodiments of the present application, a semantic analysis model can be used to detect the target output statement to determine whether the text content in the target output statement belongs to natural language in a predetermined format.
[0149] Specifically, the target output statement can be input into the semantic analysis model to extract the feature information of the target output statement, denoted as the first feature information. Then, after the semantic analysis model processes the first feature information, a text detection result is output. Here, the text detection result can be either a classification result or a numerical result. Moreover, the semantic analysis model can be further subdivided. For example, a language type analysis model can be established to detect whether the text content in the target output statement belongs to the language type defined in this interaction, and a logical analysis model can be established to detect whether the text content in the target output statement conforms to the usage mode of natural language, etc. For these semantic analysis models, a statistical language model or a language model based on deep learning is an optional implementation method, which will not be elaborated here.
[0150] In the embodiments of the present application, after obtaining the text detection result, the initial score can be determined according to the text detection result. Specifically, when the text detection result indicates that the text content in the target output statement belongs to the natural language of a predetermined format, it shows that the interaction content given by the central control sub-model basically meets the requirements semantically, and a relatively high initial score can be given to this group of first input statements and the target output statement; on the contrary, when the text detection result indicates that the text content in the target output statement does not belong to the natural language of a predetermined format, it shows that the interaction content given by the central control sub-model does not meet the requirements semantically, and at this time, a relatively low initial score can be given to this group of first input statements and the target output statement. In the present application, the specific setting method of the initial score value and the score size corresponding to various situations can be flexibly adjusted according to needs, and no limitation is made here.
[0151] In some embodiments, step 133 may include but is not limited to steps 1333 to 1336:
[0152] Step 1333: Obtain the standard interaction statement corresponding to the first input statement;
[0153] Step 1334: Extract the first feature information of the target output statement and extract the second feature information of the standard interaction statement;
[0154] Step 1335: Determine the similarity between the first feature information and the second feature information;
[0155] Step 1336: Determine the initial score according to the similarity; the initial score is positively correlated with the size of the similarity.
[0156] In the embodiments of the present application, when performing matching analysis on each group of first input statements and the target output statements corresponding to the first input statements to obtain the initial score, a standard interaction statement library corresponding to some of the first input statements can also be established in advance. Here, the standard interaction statement refers to the daily-used or standard interaction content for the first input statement. For example, for some frequently-occurring interaction contents, some of the statements can be selected as standard interaction statements according to the occurrence frequency; for some topics with clear answers or responses, the standard answers and responses can also be used as standard interaction statements. Of course, it should be noted that in the present application, there is no limit to the length of the standard interaction statement. In some embodiments, the standard interaction statement can also only include words.
[0157] It should be noted that in the embodiments of the present application, for a first input statement, there can be multiple standard interaction statements corresponding to it, and the present application does not limit the specific quantity.
[0158] When the standard interaction statement library is established, the interaction performance of the intelligent interaction model can be evaluated based on it. Specifically, for a set of first input statements and target output statements, first, the corresponding standard interaction statements can be found and determined from the standard interaction statement library according to the first input statements, and then the first feature information of the target output statements and the feature information of the standard interaction statements are extracted, denoted as the second feature information. Next, the similarity between the first feature information and the second feature information is determined. Here, it should be noted that in order to facilitate the determination of the similarity between the first feature information and the second feature information, the same data structure can be used to represent these feature information when extracting. For example, the first feature information and the second feature information can be characterized by embedding vectors. The first feature information is denoted as the first vector, and the second feature information is denoted as the second vector. Then, algorithms such as the cosine similarity algorithm, the Pearson correlation coefficient method, or the Jaccard similarity coefficient method can be used to calculate the similarity based on the first vector and the second vector.
[0159] Specifically, for example, the length of the first vector can be determined first, denoted as the first length, and the length of the second vector, denoted as the second length. Then, the product of the first length and the second length is calculated as the first value, and the inner product of the first vector and the second vector is calculated as the second value. Then, the quotient of the first value and the second value is calculated as the similarity between the first vector and the second vector, that is, the similarity between the first feature information and the second feature information.
[0160] It can be understood that the higher the similarity between the first feature information and the second feature information, the more similar the target output statement and the standard interaction statement are, the closer the interaction content given by the intelligent interaction model is to the standard interaction content, and the better the interaction effect is relatively. Therefore, in the embodiments of the present application, after obtaining the above similarity data, the initial score of the first input statement and the target output statement can be further determined according to the similarity. For example, the value of the initial score can be determined based on the similarity and a predetermined function, and this function makes the similarity and the initial score have a positive correlation relationship. The specific function adopted in the present application is not limited.
[0161] Refer to Figure 9 , in step 133 of the present application, when performing matching analysis on the first input statement and the target output statement to obtain the initial score, it can also be analyzed based on the matching relationship between the first input statement and the target output statement at the scenario level. For example, Figure 9 shows a situation where in a certain round of interaction, the same first input statement "What's the weather like tomorrow?" is input, and the intelligent interaction model outputs different target output statements. Obviously, at this time, the first input statement is a question sentence. Correspondingly, it is hoped that the intelligent interaction model can give a normal reply based on this question sentence, that is, the first input statement limits that the target output statement should be a statement in a "reply scenario".
[0162] In the first case, the target output statement output by the intelligent interaction model is "It may rain". For the first input statement "What's the weather like tomorrow?", it belongs to an output of the "reply scenario" type, constituting a normal round of interaction content of question and answer. At this time, the target output statement can better connect with the first input statement, forming a relatively smooth and coherent round of interaction, which is conducive to the continuous progress of the interaction. Similarly, in this case, it shows that the current central control sub-model of the intelligent interaction model has good performance in scene judgment for processing this round of interaction and can select an appropriate target output statement. Therefore, for the interaction content of this round, a relatively high initial score can be given.
[0163] In the second case, the target output statement output by the intelligent interaction model is "Did you buy an umbrella?". It can be seen that when the first input statement is in a "question scenario", the ideal target output statement should be the content in the "reply scenario". However, the target output statement "Did you buy an umbrella?" is obviously not an interaction content in the "reply scenario", but belongs to the "question scenario" together with the first input statement. In this case, the intelligent interaction model does not answer the question in the first input statement, but instead raises a new question. Therefore, this problem that does not match the predetermined scenario will lead to very poor logic in the interaction content and affect the smooth progress of the interaction. In this case, it shows that the current central control sub-model of the intelligent interaction model has poor performance in scene judgment for processing this round of interaction and is difficult to select an appropriate target output statement. Similarly, for the interaction content of this round, a relatively low initial score can be given. The above is an overview of the implementation principle of scene matching analysis in this application. Next, some specific implementations of scene matching analysis will be described in detail.
[0164] In some embodiments, the first input statement carries a scene label, and step 133 may include but is not limited to steps 1337 to 1339:
[0165] Step 1337: Extract the first feature information of the target output statement;
[0166] Step 1338: Perform scene analysis based on the first feature information to obtain the scene detection result of the target output statement; the scene detection result is used to characterize the preset scene category to which the target output statement belongs;
[0167] Step 1339: Determine the initial score according to the scene detection result and the scene label.
[0168] In the embodiments of the present application, when performing scene matching analysis, first, a corresponding scene label can be marked for the first input statement. The scene label is used to characterize the context scene of the target output statement to be output when the intelligent interaction model needs to perform interaction prediction for the first input statement. For example, when the scene label is a "reply scene", the context in which the target output statement is located should be a reply scene. If the context in which the target output statement is located at this time is a question scene (for example, the target output statement is a question sentence), it indicates that there are areas for improvement in the interaction prediction of the intelligent interaction model at the context scene level, and the problem needs to be improved by adjusting the model parameters. Here, the categories of the scene labels can be flexibly set according to needs. For example, they can include a "chat scene", a "reply scene", a "question scene", and so on.
[0169] Next, for each group of the first input statement and the target output statement corresponding to the first input statement, feature extraction can be performed on the target output statement to obtain first feature information. Then, scene analysis is performed on the first feature information to determine the scene detection result of the target output statement. The scene detection result is similar to the aforementioned scene label and is used to characterize the context scene of the target output statement actually output when the intelligent interaction model performs interaction prediction. It can be understood that the preset context scene categories to which the scene detection results belong can be set in the same way as the scene labels, which is convenient for subsequent matching analysis. When the scene detection result of the target output statement is determined, the scene judgment performance of the interaction of the intelligent interaction model can be judged according to the scene label carried by each group of the first input statement itself and the scene detection result of the corresponding target output statement. If the scene detection result is the same as the scene label, it indicates that the scene judgment performance of the current central control sub-model of the intelligent interaction model for processing this round of interaction is good, and the output target output statement meets the expectations. At this time, a relatively high initial score can be given to this group of the first input statement and the target output statement. On the contrary, if the scene detection result is different from the scene label, it indicates that the scene judgment performance of the current central control sub-model of the intelligent interaction model for processing this round of interaction is poor, and the output target output statement does not meet the expectations. At this time, a relatively low initial score can be given to this group of the first input statement and the target output statement. Similarly, in the present application, the specific setting method of the initial score value and the score sizes corresponding to various situations can be flexibly adjusted according to needs and are not limited herein.
[0170] In some more detailed embodiments, step 1338 of the present application may include, but is not limited to, steps 13381 to 13382:
[0171] Step 13381: Input the first feature information into the intention analysis model to obtain the intention prediction result output by the intention analysis model; the intention prediction result is used to characterize the interaction action category of the sub-interaction model for the first input statement, and the interaction action category includes at least question and reply;
[0172] Step 13382: Determine the scene detection result according to the intention prediction result.
[0173] In the embodiments of the present application, when the user interacts through the intelligent interaction model, according to different interaction requirements or interaction stages, the one leading each round of interaction may be the user or the intelligent interaction model. Therefore, the intelligent interaction model needs to better adapt to the interaction scene requirements in different situations. Thus, when performing scene matching analysis, the preset scene categories can also be divided into a "leading type" and a "led type". Similarly, the first input statement can be pre-labeled with scene labels of the "leading type" and the "led type". At this time, the scene label is used to characterize the context scene of the target output statement that needs to be selected by the intelligent interaction model. For example, when the desired target output statement belongs to the "leading type", the intelligent interaction model needs to select a more topical and interaction-continuable target output statement; when the desired target output statement belongs to the "led type", the intelligent interaction model needs to select a target output statement that better meets the requirements of the first input statement.
[0174] Specifically, for the target output statement or the first input statement, to determine whether it belongs to the "leading type" and the "led type", it can be judged according to the intention information represented by these statements themselves, and the intention can be determined by the interaction action category. In other words, to judge who is in the leading position, it can be determined according to the interaction action category of the user or the intelligent interaction model during the interaction process, because the action itself is a dimension of the intention. For example, the interaction action categories of the first input statement and the target output statement can include asking questions, answering, notifying, accepting, denying, suggesting, echoing, objecting, etc. If the statement of the user or the intelligent interaction model is determined to be an interaction action category such as asking questions, notifying, denying, suggesting, objecting, etc., then he (it) is in a leading position. For example, in a specific context scene, asking a question can be understood as asking the other party to answer, objecting and denying mean not agreeing with the other party's view and asking the other party to explain, and suggesting means actively proposing a new topic plan, etc. If the statement of the user or the intelligent interaction model is determined to be an interaction action category such as answering, accepting, echoing, then it is relatively passive and does not take the leading position in the specific context scene.
[0175] Therefore, in the embodiments of the present application, the first feature information of the target output statement can be extracted and input into the intent analysis model. The intent analysis model discriminates the interaction actions of the target output statement, thereby determining the intent prediction result. Then, the scene detection result corresponding to the target output statement can be determined according to the intent prediction result. For example, when the intent prediction result indicates that the interaction action of the target output statement belongs to one of question, notification, denial, suggestion, and objection, the scene detection result corresponding to the target output statement is "dominant type"; when the intent prediction result indicates that the interaction action of the target output statement belongs to one of answer, acceptance, and echo, the scene detection result corresponding to the target output statement is "dominated type".
[0176] According to the scene labels carried by each group of first input statements and the scene detection results of the corresponding target output statements, the scene judgment performance of the intelligent interaction model can be judged. For example, when the scene label carried by the first input statement itself is "dominated type", and the scene detection result corresponding to the target output statement is "dominant type", it indicates that the target output statement does not make a good action response to the first input statement (such as the situation of mutual questioning), which has an adverse impact on the coherence of the interaction; similarly, when the scene label carried by the first input statement itself is "dominant type", and the scene detection result corresponding to the target output statement is "dominated type", it indicates that the target output statement does not complete the task of continuing the interaction topic, which is likely to lead to the end of the interaction and is also not conducive to the smooth progress of the multi-round interaction task.
[0177] In the above embodiments, the implementation of obtaining the initial score by performing content matching analysis or scene matching analysis on each group of first input statements and the target output statements corresponding to the first input statements is introduced. In an embodiment of the present application, in combination with the above embodiments, step 130 is further described. Step 130 may further include but is not limited to steps 135 to 137.
[0178] Step 135: Perform content matching analysis on multiple groups of first input statements and the target output statements corresponding to the first input statements to obtain a first score;
[0179] Step 136: Perform scene matching analysis on multiple groups of first input statements and the target output statements corresponding to the first input statements to obtain a second score;
[0180] Step 137: Weight the first score and the second score to obtain a matching score.
[0181] In the embodiments of the present application, for the initial score corresponding to a single set of first input statements and target output statements, it can be obtained by content matching analysis, or by scenario matching analysis, or by weighting the scores determined by the two analysis methods. Therefore, when determining the overall matching score, some of the first input statements and target output statements can be evaluated through content matching analysis, and the obtained initial score is recorded as the first score; for another part of the first input statements and target output statements, they are evaluated through scenario matching analysis, and the obtained initial score is recorded as the second score, and then the first score and the second score are weighted to obtain the matching score. In this way, the determined matching score can better take into account the judgment and processing performance of the intelligent interaction model for interaction content and interaction scenarios, and can improve the stability of the trained model and the effect of processing multi-round interaction tasks.
[0182] In some embodiments, as pointed out in the embodiments of steps 150 to 180, each sub-interaction model can be built and trained according to functional requirements. For example, in the intelligent interaction model, one (or more) sub-interaction models can be trained specifically for the requirements in the reply scenario, and another (or more) sub-interaction models can be trained specifically for the requirements in the question scenario. In the present application, when training these sub-interaction models, a specified training data set can be used for training. For example, a sub-interaction model can be trained using the corpus specifically processed for the "reply scenario", and the association relationship between the training data set and the label of the "reply scenario" is established. In this way, when training the central control sub-model, when the first input statement in the input carries the scenario label "reply scenario", theoretically, the sub-interaction model trained using the corpus specifically processed for the "reply scenario" should have better interaction prediction performance for this first input statement.
[0183] Therefore, the performance of the central control sub-model can be judged according to whether the central control sub-model selects the initial output statement output by the sub-interaction model. For example, if the central control sub-model selects the initial output statement output by the sub-interaction model as the target output statement for this round, a relatively high matching score (initial score) can be given to the interaction in this round; conversely, if the central control sub-model does not select the initial output statement output by the sub-interaction model as the target output statement for this round, a relatively low matching score (initial score) can be given to the interaction in this round. Specifically, when performing matching analysis on the first input statement and the target output statement to determine the matching score, the sub-interaction model that outputs the target output statement can be determined first, denoted as the target sub-interaction model, and then the matching relationship between the target sub-interaction model and the label of the first input statement can be detected to determine the matching score. In this way, complex analysis and matching of the specific content of the statement are not required, and the matching score can be quickly determined only by simply comparing the matching relationship between the target sub-interaction model and the label of the first input statement, which can greatly improve the training efficiency and reduce the consumption of computing resources.
[0184] In an embodiment of the present application, an interaction method is further provided. Similarly, this interaction method can be applied in the Figure 1 shown implementation environment. Moreover, this interaction method can be configured to be executed by the operation terminal 101 alone, or can be configured to be executed by the server 102 alone, or can be executed based on the interaction between the operation terminal 101 and the server 102. Specifically, it can be appropriately selected according to the actual application situation, and this embodiment does not make specific limitations on this.
[0185] Referring to Figure 10 , a flowchart of an interaction method provided by an embodiment of the present application. In this embodiment, the operation terminal and the server are taken as the execution main body for illustration. Referring to Figure 10 , this interaction method includes but is not limited to steps 210 to 240.
[0186] Step 210: Collect voice data;
[0187] Step 220: Perform speech recognition on the text content of the voice data to obtain the third input information;
[0188] Step 230: Input the third input information into the intelligent interaction model trained by the training method of the intelligent interaction model as shown in Figure 3 to obtain the target output statement output by the intelligent interaction model;
[0189] Step 240: Convert the target output statement into audio data for output.
[0190] In an embodiment of the present application, taking the interaction between an operating terminal and a server to implement the interaction method in the present application as an example, the operating terminal at least has the functions of collecting voice data of a user, sending the voice data to the server, receiving the text data of the target output statement returned by the server, and converting the text data of the target output statement into audio data for output; the server at least has the functions of receiving the voice data sent by the operating terminal, identifying the text content of the voice data to obtain input information, inputting the input information into a trained intelligent interaction model to obtain a target output statement, and sending the text data of the target output statement to the operating terminal. In this way, the operating terminal can send the collected voice data to the server, and the intelligent interaction model in the server performs interaction prediction on the text content of the voice data, outputs the target output statement, and plays it to the user through the operating terminal, so as to perform intelligent interaction.
[0191] In an alternative implementation, the operating terminal may be a vehicle-mounted terminal installed with an automatic navigation APP. The vehicle-mounted terminal may include an audio collection component, a communication component, and a sound component. In response to the user's operation of opening the automatic navigation APP of the vehicle-mounted terminal, the vehicle-mounted terminal initiates a call to the audio collection component; after the audio collection component collects the voice data of the user, the voice data is sent to the background server of the automatic navigation APP through the communication component of the vehicle-mounted terminal. In the background server of the automatic navigation APP, the voice data of the user is converted into text information by automatic speech recognition technology (ASR), denoted as the third input information. Then, the server inputs the third input information into the trained intelligent interaction model, and the model can give the target output statement corresponding to the third input information and send the target output statement back to the communication component of the vehicle-mounted terminal, so that after receiving the target output statement, the vehicle-mounted terminal plays the corresponding audio data through the sound component, thereby realizing the automatic navigation function based on voice interaction.
[0192] In another alternative implementation, the operating terminal can be a human-machine dialogue robot, which can also include an audio acquisition component, a communication component, and a sound component. Users can interact with the human-machine dialogue robot to implement functions such as information consultation and scenario dialogue simulation. For example, in the application scenario of hospital diagnosis, a human-machine dialogue robot can be set up to help answer relevant medical condition consultations and process handling matters; in the application scenario of assisted learning, the human-machine dialogue robot can help students understand a large amount of information consultation and answer relevant difficult questions to assist students in learning; in the application scenario of service training, a human-machine dialogue robot can be used to simulate customers to exercise the professional qualities and communication abilities of service personnel. Of course, the above implementation scenarios are only used to illustrate some specific implementation methods of the interaction method provided in this application, and do not mean to limit its specific implementation.
[0193] Referring to Figure 11 , the embodiment of the present application also provides a training device for an intelligent interaction model. The intelligent interaction model includes a central control sub-model and multiple trained different sub-interaction models. The training device includes:
[0194] An acquisition module 1110, configured to acquire first input information, where the first input information includes multiple first input statements;
[0195] A prediction module 1120, configured to input the first input statement into the intelligent interaction model, perform interaction prediction on the first input statement through each sub-interaction model to obtain multiple initial output statements, and determine a target sub-interaction model according to the first input statement through the central control sub-model, and determine the initial output statement output by the target sub-interaction model as the target output statement corresponding to the first input statement;
[0196] A scoring module 1130, configured to perform matching analysis on several groups of first input statements and the target output statements corresponding to the first input statements in response to a matching instruction to obtain a matching score; the matching instruction is used to guide at least one of content matching analysis or scenario matching analysis;
[0197] An update module 1140, configured to determine a reward value according to the matching score, perform reinforcement learning training on the central control sub-model through the reward value to obtain a trained intelligent interaction model; the matching score and the reward value are positively correlated.
[0198] It can be understood that Figure 3 The content in the embodiment of the training method of the intelligent interaction model shown is applicable to the embodiment of the training device of the intelligent interaction model. The functions specifically implemented by the embodiment of the training device of the intelligent interaction model are the same as those in Figure 3 the embodiment of the training method of the intelligent interaction model shown, and the beneficial effects achieved are the same as those in Figure 3The beneficial effects achieved by the embodiments of the training method of the intelligent interaction model shown are also the same.
[0199] Referring to Figure 12 , an embodiment of the present application also discloses a computer device, including:
[0200] At least one processor 1210;
[0201] At least one memory 1220, configured to store at least one program;
[0202] When the at least one program is executed by the at least one processor 1210, the at least one processor 1210 is caused to implement the embodiments of the training method of the intelligent interaction model as Figure 3 shown or Figure 10 the embodiments of the interaction method shown.
[0203] It can be understood that the content in the embodiments of the training method of the intelligent interaction model as Figure 3 shown or Figure 10 the embodiments of the interaction method shown is applicable to this embodiment of the computer device. The functions specifically implemented by this embodiment of the computer device are the same as those in the embodiments of the training method of the intelligent interaction model as Figure 3 shown or Figure 10 the embodiments of the interaction method shown, and the beneficial effects achieved are also the same as those in the embodiments of the training method of the intelligent interaction model as Figure 3 shown or Figure 10 the embodiments of the interaction method shown.
[0204] An embodiment of the present application also discloses a computer-readable storage medium, in which a program executable by a processor is stored, and the program executable by the processor is used to implement the embodiments of the training method of the intelligent interaction model as Figure 3 shown or Figure 10 the embodiments of the interaction method shown.
[0205] It can be understood that the content in the embodiments of the training method of the intelligent interaction model as Figure 3 shown or Figure 10 the embodiments of the interaction method shown is applicable to this embodiment of the computer-readable storage medium. The functions specifically implemented by this embodiment of the computer-readable storage medium are the same as those in the embodiments of the training method of the intelligent interaction model as Figure 3 shown or Figure 10 the embodiments of the interaction method shown, and the beneficial effects achieved are also the same as those in the embodiments of the training method of the intelligent interaction model as Figure 3 shown or Figure 10 the embodiments of the interaction method shown.
[0206] In some alternative embodiments, the functions / operations recited in the block diagrams may not occur in the order presented in the operational illustrations. For example, depending on the functions / operations involved, two blocks shown in succession may actually be executed substantially simultaneously or the blocks can sometimes be executed in reverse order. Further, the embodiments presented and described in the flowcharts of the present application are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logic flows presented herein. Alternative embodiments are contemplated in which the order of various operations is altered and in which sub-operations described as part of a larger operation are performed independently.
[0207] Moreover, although the present application has been described in the context of functional modules, it should be understood that one or more of the functions and / or features, unless otherwise stated to the contrary, may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present application. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skill of an engineer. Accordingly, those of ordinary skill in the art will be able to implement the present application as set forth in the claims without undue experimentation. It should also be understood that the particular concepts disclosed are illustrative only and are not intended to limit the scope of the present application, the scope of which is determined by the full scope of the appended claims and their equivalents.
[0208] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present application, in essence or the part that contributes to the prior art or part of the technical solution, may be embodied in the form of a software product stored in a storage medium, including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a portable hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0209] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered as a definitional sequence of executable instructions for implementing a logical function, and can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with the instruction execution system, apparatus, or device.
[0210] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which a program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then storing it in a computer memory.
[0211] It should be understood that various parts of the present application can be implemented in hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.
[0212] In the above description of this specification, the descriptions referring to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0213] Although the embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present application. The scope of the present application is defined by the claims and their equivalents.
[0214] The above has specifically described the preferred embodiments of the present application, but the present application is not limited to the embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present application, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present application.
[0215] In the description of this specification, the description with reference to terms such as "one embodiment", "another embodiment", or "certain embodiments" means that the specific features, structures, materials, or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0216] Although the embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present application. The scope of the present application is defined by the claims and their equivalents.
Claims
1. A training method for an intelligent interaction model, characterized in that, the intelligent interaction model includes a central control sub-model and multiple trained different sub-interaction models; the training method includes: obtaining first input information, where the first input information includes multiple first input statements; inputting the first input statements into the intelligent interaction model, performing interaction prediction on the first input statements through each of the sub-interaction models to obtain multiple initial output statements, and determining a target sub-interaction model according to the first input statements through the central control sub-model, and determining the initial output statement output by the target sub-interaction model as the target output statement corresponding to the first input statement; in response to a matching instruction, performing matching analysis on several groups of the first input statements and the target output statements corresponding to the first input statements to obtain a matching score; the matching instruction is used to guide at least one of content matching analysis or scenario matching analysis; determining a reward value according to the matching score, and performing reinforcement learning training on the central control sub-model through the reward value to obtain a trained intelligent interaction model; the size of the matching score is positively correlated with the size of the reward value; wherein, the performing matching analysis on several groups of the first input statements and the target output statements corresponding to the first input statements to obtain a matching score includes: performing content matching analysis on multiple groups of the first input statements and the target output statements corresponding to the first input statements to obtain a first score; performing scenario matching analysis on multiple groups of the first input statements and the target output statements corresponding to the first input statements to obtain a second score; weighting the first score and the second score to obtain the matching score.
2. The training method for an intelligent interaction model according to claim 1, characterized in that, the inputting the first input statements into the intelligent interaction model and performing interaction prediction on the first input statements through each of the sub-interaction models includes: inputting the first input statements of the current round of interaction into the intelligent interaction model, and determining historical interaction information according to the interaction round of the current round of interaction; the historical interaction information includes the first input statements and target output statements of a preset number of rounds before the current round of interaction; performing interaction prediction on the first input statements of the current round of interaction through each of the sub-interaction models according to the historical interaction information.
3. The training method for an intelligent interaction model according to claim 2, characterized in that, the performing matching analysis on several groups of the first input statements and the target output statements corresponding to the first input statements includes: determining the interaction round in which each group of the first input statements and the target output statements corresponding to the first input statements are located; selecting multiple groups of the first input statements and the target output statements corresponding to the first input statements with later interaction rounds for matching analysis according to a preset matching round.
4. The training method for an intelligent interaction model according to claim 3, characterized in that, Performing matching analysis on several groups of the first input statements and the target output statements corresponding to the first input statements to obtain a matching score further includes: Performing matching analysis on each group of the first input statements and the target output statements corresponding to the first input statements to obtain an initial score; Weighting the initial score according to the interaction rounds in which each group of the first input statements and the target output statements corresponding to the first input statements are located to obtain the matching score; wherein, the weighting weight of the initial score of each group of the first input statements and the target output statements corresponding to the first input statements is positively correlated with the size of the interaction rounds.
5. The training method of an intelligent interaction model according to claim 4, characterized in that Performing matching analysis on each group of the first input statements and the target output statements corresponding to the first input statements to obtain an initial score includes: Extracting first feature information of the target output statement; Performing semantic analysis according to the first feature information to obtain a text detection result of the target output statement; the text detection result is used to characterize whether the text content in the target output statement belongs to natural language in a predetermined format; Determining the initial score according to the text detection result.
6. The training method of an intelligent interaction model according to claim 4, characterized in that Performing matching analysis on each group of the first input statements and the target output statements corresponding to the first input statements to obtain an initial score includes: Obtaining a standard interaction statement corresponding to the first input statement; Extracting first feature information of the target output statement and extracting second feature information of the standard interaction statement; Determining the similarity between the first feature information and the second feature information; Determining the initial score according to the similarity; the initial score is positively correlated with the size of the similarity.
7. The training method of an intelligent interaction model according to claim 4, characterized in that The first input statement carries a scene label; Performing matching analysis on each group of the first input statements and the target output statements corresponding to the first input statements to obtain an initial score includes: Extracting first feature information of the target output statement; Performing scene analysis according to the first feature information to obtain a scene detection result of the target output statement; the scene detection result is used to characterize the preset scene category to which the target output statement belongs; Determining the initial score according to the scene detection result and the scene label.
8. The training method of an intelligent interaction model according to claim 7, characterized in that Performing scene analysis according to the first feature information to obtain a scene detection result of the target output statement includes: Inputting the first feature information into an intention analysis model to obtain an intention prediction result output by the intention analysis model; the intention prediction result is used to characterize the interaction action category of the sub-interaction model for the first input statement, and the interaction action category at least includes asking questions and answering; Determining the scene detection result according to the intention prediction result.
9. A training method for an intelligent interaction model according to any one of claims 1-8, characterized in that, the first input statement is labeled, and the sub-interaction model is trained through a training data set corresponding to the label; the matching analysis of a plurality of groups of the first input statements and the target output statements corresponding to the first input statements to obtain a matching score further includes: determining a target sub-interaction model determined by the central control sub-model according to the first input statement according to the target output statement corresponding to the first input statement; determining the matching score according to the matching relationship between the target sub-interaction model and the label of the first input statement.
10. An interaction method, characterized in that, the method includes the following steps: collecting voice data; performing speech recognition on the text content of the voice data to obtain third input information; inputting the third input information into an intelligent interaction model trained by the training method of the intelligent interaction model according to any one of claims 1-9 to obtain a target output statement output by the intelligent interaction model; converting the target output statement into audio data for output.
11. A training device for an intelligent interaction model, characterized in that, the intelligent interaction model includes a central control sub-model and a plurality of trained different sub-interaction models; the training device includes: an acquisition module for acquiring first input information, where the first input information includes a plurality of first input statements; a prediction module for inputting the first input statement into the intelligent interaction model, performing interaction prediction on the first input statement through each sub-interaction model to obtain a plurality of initial output statements, and determining a target sub-interaction model according to the first input statement through the central control sub-model, and determining the initial output statement output by the target sub-interaction model as the target output statement corresponding to the first input statement; a scoring module for performing matching analysis on a plurality of groups of the first input statements and the target output statements corresponding to the first input statements in response to a matching instruction to obtain a matching score; the matching instruction is used to guide at least one of content matching analysis or scenario matching analysis; an update module for determining a reward value according to the matching score, and performing reinforcement learning training on the central control sub-model through the reward value to obtain a trained intelligent interaction model; the matching score is positively correlated with the reward value; wherein, the matching analysis of a plurality of groups of the first input statements and the target output statements corresponding to the first input statements to obtain a matching score includes: performing content matching analysis on multiple groups of the first input statements and the target output statements corresponding to the first input statements to obtain a first score; performing scenario matching analysis on multiple groups of the first input statements and the target output statements corresponding to the first input statements to obtain a second score; weighting the first score and the second score to obtain the matching score.
12. A computer device, characterized in that, includes: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the training method of the intelligent interaction model according to any one of claims 1-9 or implements the interaction method according to claim 10.
13. A computer-readable storage medium storing a program executable by a processor, wherein: the program executable by the processor, when executed by the processor, is used to implement the training method of the intelligent interaction model according to any one of claims 1-9 or to implement the interaction method according to claim 10.
Citation Information
Patent Citations
A method and device for generating model
CN110737758A