Interaction method, device and system, and intelligent device
By obtaining the semantic entities of the input statement and directly outputting feedback information using the dialogue model, the problem of high code demand in the existing technology is solved, and the effect of reducing costs and handling complex dialogue interactions is achieved.
Patent Information
- Application Number
- CN202010513432.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-08
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-06-08
AI Technical Summary
In the process of dialogue decision making and implementation, the code demand is high, making it expensive and difficult to achieve complex dialogue interactions.
By obtaining semantic entities in the input statement and directly outputting feedback information using the dialogue model, the logical code writing steps in traditional methods are avoided.
It realizes reducing the code demand, reducing the cost of dialogue decision-making, and can effectively handle complex dialogue interaction processes.
Smart Images

Figure CN113836932B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Internet technologies, and in particular, to an interaction method, apparatus, and system, and an intelligent device. Background Art
[0002] In a voice interaction scenario, after a user's voice input is recognized into text by Automatic Speech Recognition (ASR), it needs to go through parts such as intent detection, entity recognition, dialogue decision-making, and text generation to obtain the text for replying to the user, and then be converted into sound through TTS and played to the user. Among them, dialogue decision-making is mainly based on carefully designed logic code written by developers. Once the designed dialogue logic is too complex, the cost of writing dialogue logic code becomes very high or even difficult to implement.
[0003] In a traditional dialogue interaction system, after a user's voice input is recognized into text by ASR, it needs to go through parts such as intent detection, entity recognition, dialogue decision-making, and text generation to obtain the text for replying to the user, and then be converted into sound through TTS and played to the user.
[0004] Among them, modules such as intent detection, entity recognition, and dialogue decision-making together constitute the core of dialogue interaction. The traditional method splits dialogue interaction into the above three parts: intent detection identifies the intent of the user interaction; entity recognition obtains the key entity information in the user request; and dialogue decision-making obtains the actions to be performed and the reply to the user according to the user's context and environmental information, and this part is mainly completed by developers writing logic code.
[0005] However, the above-mentioned existing technologies lead to an error propagation effect in the processing process of the pipeline (PipeLine). For example, if the accuracy rate of each of the three parts is 0.9, then the accuracy rate of the final system can only reach 0.729; and the dialogue decision-making system implemented through logic code brings a significant increase in maintenance and modification costs, and is extremely difficult to implement for complex dialogue interaction processes.
[0006] Aiming at the problem of large code requirements in the process of formulating and implementing dialogue decision-making in the above-mentioned existing technologies, no effective solution has been proposed yet. Summary of the Invention
[0007] Embodiments of the present invention provide an interaction method, apparatus, and system, and an intelligent device, so as to at least solve the technical problem of large code requirements in the process of formulating and implementing dialogue decision-making by the existing technology.
[0008] According to one aspect of the embodiments of the present invention, an interaction method is provided, including: obtaining a semantic entity in the input statement according to the collected input statement, where the semantic entity is used to represent the semantics in the input statement; obtaining feedback information corresponding to the input statement through a dialogue model according to the semantic entity, where the feedback information includes at least one of the following: a reply to the input statement, and a third-party service requested to be called in the input statement.
[0009] Optionally, obtaining the semantic entity in the input statement includes: segmenting the input statement to obtain at least one word segment that makes up the input statement; obtaining the category to which the at least one word segment belongs, and determining the semantic entity in the input statement according to the category.
[0010] Optionally, obtaining the feedback information corresponding to the input statement through the dialogue model according to the input statement and the semantic entity includes: obtaining the historical interaction result; obtaining the dialogue vector corresponding to the input statement according to the input statement, the semantic entity, and the historical interaction result; obtaining the feedback information corresponding to the dialogue vector through the dialogue model.
[0011] Further, optionally, the method further includes: after obtaining the feedback information, inputting the feedback information and the input statement into the dialogue model as updated data to obtain an optimized dialogue model.
[0012] Optionally, the dialogue model is trained by a process of obtaining the language feature of the input statement, and obtaining the corresponding dialogue vector according to the language feature; obtaining the category label and action feature of the input statement according to the language feature and the dialogue vector; generating the feedback information corresponding to the input statement according to the category label and the action feature; and obtaining the language feature, dialogue vector, category label, and action feature of the input statement according to the historical dialogue pair containing the feedback information.
[0013] According to one aspect of the embodiments of the present invention, another interaction method is provided, including: collecting input statements input by multiple objects, and obtaining semantic entities in the input statements, where the semantic entities are used to represent the semantics in the input statements; determining the language features of different objects based on the semantic entities of different objects; calling different dialogue models based on the language features of different objects; processing the input statements input by different objects with the corresponding dialogue models to generate feedback information corresponding to different objects, where the feedback information includes at least one of the following: feedback information on the input statement, and a third-party service requested to be called in the input statement.
[0014] According to one aspect of the embodiments of the present invention, another interactive method is provided, including: obtaining the language features of sample data, and obtaining the corresponding dialogue vector according to the language features; obtaining the category label and action features of the sample data according to the language features and the dialogue vector; generating the feedback information corresponding to the sample data according to the category label and the action features; training the process of obtaining the language features, dialogue vector, category label and action features of the sample data according to the historical dialogue pair containing the feedback information to obtain a dialogue model, where the dialogue model is used to obtain the feedback information corresponding to the input statement, and the feedback information includes at least one of the following: the feedback information on the input statement, the third-party service requested to be called in the input statement.
[0015] Optionally, obtaining the language features of the sample data includes: respectively performing entity recognition and semantic recognition on the sample data to obtain the language features.
[0016] Further, optionally, respectively performing entity recognition and semantic recognition on the sample data to obtain the language features includes: performing entity recognition on the sample data to obtain semantic entities; matching the semantic entities with preset data to obtain a first recognition result; performing semantic recognition on the sample data to obtain a second recognition result; obtaining the language features according to the first recognition result, the second recognition result and the state of the input sample data, where the language features include: state features, entity features, semantic features and word features.
[0017] Optionally, obtaining the corresponding dialogue vector according to the language features includes: associating the language features with at least one historical dialogue data to obtain the dialogue vector.
[0018] Further, optionally, associating the language features with at least one historical dialogue data to obtain the dialogue vector includes: associating the state features, entity features, semantic features and word features in the language features with at least one historical dialogue data to obtain a dialogue vector including the language features and at least one historical dialogue data.
[0019] Optionally, obtaining the category label and action features of the sample data according to the language features and the dialogue vector includes: obtaining the category label of the sample data according to the dialogue vector; obtaining the action features of the sample data according to the language features and the dialogue vector.
[0020] Further, optionally, obtaining the category label of the sample data according to the dialogue vector includes: classifying the semantic labels in the dialogue vector to obtain the category label of the sample data.
[0021] Optionally, the interactive method is applied to an end-to-end interactive system, where the end-to-end interactive system includes: an online interactive system, an intelligent interactive device, and a device embedded in the intelligent interactive device.
[0022] According to one aspect of an embodiment of the present invention, there is provided yet another interaction method, including: Step A, collecting an input statement currently input; Step B, calling a dialogue model to analyze the input statement currently input, and displaying feedback information having a semantic relationship with the input statement currently input on a touch screen; Step C, collecting an input statement input again based on the feedback information; Step D, calling the dialogue model to analyze the input statement input again, and displaying feedback information having a semantic relationship with the input statement input again on the touch screen; Step E, repeatedly executing the above Steps C and D until an end instruction is received.
[0023] According to another aspect of an embodiment of the present invention, there is provided an interaction method, including: a server receiving an input statement collected by an intelligent device and obtaining a semantic entity in the input statement, where the semantic entity is used to represent the semantics in the input statement; the server obtaining feedback information matching the input statement through a dialogue model according to the semantic entity, where the feedback information includes at least one of the following: a reply to the input statement, a third-party service requested to be called in the input statement; the server sending the feedback information to the intelligent device.
[0024] According to another aspect of an embodiment of the present invention, there is provided an intelligent device, including: a voice collection device for collecting voice information; a processor connected to the voice collection device for identifying an input statement from the voice information and processing the semantic entity of the input statement through a dialogue model to generate corresponding feedback information, where the semantic entity of the input statement represents the semantics in the input statement.
[0025] According to another aspect of an embodiment of the present invention, there is provided another intelligent device, including: a touch screen for collecting an input statement input through a touch interaction function, where the semantic entity of the input statement represents the semantics in the input statement; a processor connected to the touch screen for processing the semantic entity of the input statement through a dialogue model to generate corresponding feedback information; where the touch screen is further used to display the feedback information.
[0026] According to another aspect of an embodiment of the present invention, there is provided yet another intelligent device, including: a voice collection device for collecting voice information; a touch screen for displaying at least one statement generated by recognizing the voice information and receiving an input statement selected from the at least one statement through a touch interaction function, where the semantic entity of the input statement represents the semantics in the input statement; a processor connected to the voice collection device and the touch screen for processing the semantic entity of the input statement through a dialogue model to generate corresponding feedback information; where the touch screen is further used to display the feedback information.
[0027] According to another aspect of the embodiments of the present invention, an interaction device is provided, including: an acquisition module, configured to obtain semantic entities in an input statement according to the collected input statement, where the semantic entities are used to identify the semantics in the input statement; an interaction module, configured to obtain feedback information corresponding to the input statement through a dialogue model according to the input statement and the semantic entities, where the feedback information includes at least one of the following: a reply to the input statement, a third-party service requested to be invoked in the input statement.
[0028] According to another aspect of the embodiments of the present invention, another interaction device is provided, including: a collection module, configured to collect input statements input by multiple objects and obtain semantic entities in the input statements, where the semantic entities are used to represent the semantics in the input statements; a feature module, configured to determine language features of different objects based on the semantic entities of different objects; a call module, configured to call different dialogue models based on the language features of different objects; an interaction module, configured to process the input statements input by different objects using corresponding dialogue models to generate feedback information corresponding to different objects, where the feedback information includes at least one of the following: feedback information on the input statement, a third-party service requested to be invoked in the input statement.
[0029] According to another aspect of the embodiments of the present invention, yet another interaction device is provided, including: a first acquisition module, configured to obtain the language feature of sample data and obtain a corresponding dialogue vector according to the language feature; a second acquisition module, configured to obtain the category label and action feature of the sample data according to the language feature and the dialogue vector; a statement generation module, configured to generate feedback information corresponding to the sample data according to the category label and the action feature; a model training module, configured to train the process of obtaining the language feature, dialogue vector, category label and action feature of the sample data according to a historical dialogue pair including feedback information to obtain a dialogue model, where the dialogue model is configured to obtain feedback information corresponding to an input statement, where the feedback information includes at least one of the following: feedback information on the input statement, a third-party service requested to be invoked in the input statement.
[0030] According to another aspect of the embodiments of the present invention, an interactive system is provided, including: an entity recognition module, a semantic modeling module, a feature acquisition module, a gated recurrent module, an action decision module, a classification module, and a dialogue reply module. Among them, the entity recognition module is used to perform entity recognition on the input statement to obtain semantic entities, and match the semantic entities with preset data to obtain a first recognition result; the semantic modeling module is used to perform semantic recognition on the input statement to obtain a second recognition result; the feature acquisition module is connected to the entity recognition module and the semantic modeling module respectively, and is used to obtain language features according to the first recognition result, the second recognition result, and the state of the input statement. The language features include: state features, entity features, semantic features, and word features; the gated recurrent module is connected to the feature acquisition module and is used to associate with at least one historical dialogue data according to the state features, entity features, semantic features, and word features in the language features to obtain a dialogue vector including the language features and at least one historical dialogue data; the action decision module is connected to the feature acquisition module and the gated recurrent module and is used to obtain the action features of the input statement according to the language features and the dialogue vector; the classification module is connected to the gated recurrent module and is used to obtain the category label of the input statement according to the dialogue vector; the dialogue reply module generates feedback information corresponding to the input statement according to the category label and the action features.
[0031] According to another aspect of the embodiments of the present invention, a non-volatile storage medium is provided. The non-volatile storage medium includes a stored program. When the program runs, it controls the device where the non-volatile storage medium is located to execute the above interactive method.
[0032] According to another aspect of the embodiments of the present invention, a processor is provided. The processor is used to run a program. When the program runs, it executes the above interactive method.
[0033] In the embodiments of the present invention, by obtaining the semantic entities in the input statement according to the collected input statement, where the semantic entities are used to represent the semantics in the input statement; and according to the semantic entities, obtaining the feedback information corresponding to the input statement through a dialogue model, where the feedback information includes at least one of the following: a reply to the input statement, a third-party service requested to be called in the input statement, the purpose of directly outputting the feedback information through the dialogue model is achieved, thereby realizing the technical effect of avoiding the logical code writing steps in the traditional method, and further solving the technical problem of the large demand for code in the process of dialogue decision-making and implementation in the prior art. Description of the Drawings
[0034] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0035] Figure 1 is a hardware structure block diagram of a computer terminal for an interaction method according to an embodiment of the present invention;
[0036] Figure 2 is a flowchart of an interaction method according to an embodiment of the present invention;
[0037] Figure 3 is a flowchart of an interaction method according to the second embodiment of the present invention;
[0038] Figure 4 is a flowchart of an interaction method according to the third embodiment of the present invention;
[0039] Figure 5 is a schematic diagram of a dialogue system in an interaction method according to the third embodiment of the present invention;
[0040] Figure 6 is a flowchart of an interaction method according to the fourth embodiment of the present invention;
[0041] Figure 7 is a flowchart of an interaction method according to the fifth embodiment of the present invention;
[0042] Figure 8a is a schematic diagram of an intelligent device according to the sixth embodiment of the present invention;
[0043] Figure 8b is a schematic diagram of an intelligent device according to the seventh embodiment of the present invention;
[0044] Figure 8c is a schematic diagram of an intelligent device according to the eighth embodiment of the present invention;
[0045] Figure 9 is a schematic diagram of an interaction device according to the eleventh embodiment of the present invention;
[0046] Figure 10 is a schematic diagram of an interaction device according to the twelfth embodiment of the present invention;
[0047] Figure 11 is a schematic diagram of an interaction device according to the thirteenth embodiment of the present invention;
[0048] Figure 12 is a schematic diagram of an interaction system according to the fourteenth embodiment of the present invention;
[0049] Figure 13 is a structure block diagram of a computer terminal according to an embodiment of the present invention. Detailed implementation manners
[0050] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0051] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0052] Technical terms involved in this application:
[0053] GRU: gated recurrent unit, which is a type of recurrent neural network (RNN);
[0054] ASR: Automatic Speech Recognition;
[0055] TTS: Text To Speech;
[0056] PipeLine: In the NET Framework add-in programming model, it refers to a linear communication model that identifies the pipeline segments for exchanging data between an add-in and its host, and is literally translated as conduit;
[0057] LSTM: Long Short-Term Memory, which is also a type of recurrent neural network (RNN).
[0058] Embodiment 1
[0059] According to an embodiment of the present invention, there is also provided a method embodiment of an interaction method. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0060] The method embodiment provided in the first embodiment of this application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a computer terminal as an example, Figure 1 is a hardware structure block diagram of a computer terminal for an interaction method according to an embodiment of the present invention. As Figure 1 shown, the computer terminal 10 may include one or more (only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. Those of ordinary skill in the art can understand that, Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than those Figure 1 shown, or have a different configuration from that Figure 1 shown.
[0061] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / modules corresponding to the interaction method in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the vulnerability detection method of the above-mentioned application program. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely provided with respect to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0062] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 may be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0063] In a voice interaction scenario, after the user's voice input is recognized into text by ASR (Automatic Speech Recognition), it needs to go through parts such as intent detection, entity recognition, dialogue decision-making, and text generation to obtain the text for replying to the user, and then become voice through TTS (Text To Speech) and be played to the user. Among them, the dialogue decision-making is mainly completed based on the carefully designed logic code written by developers. Once the designed dialogue logic is too complex, the cost of writing the dialogue logic code becomes very high or even difficult to implement. In this embodiment, in a complex dialogue design scenario, an interaction solution is provided. The text input by the user is directly input into the dialogue model, and the dialogue model directly outputs the reply voice, thus avoiding the problems of complex processing processes, large amounts of code to be written, and low accuracy.
[0064] In the above operating environment, the present application provides an Figure 2 interaction method as shown. Figure 2 It is a flowchart of an interaction method according to an embodiment of the present invention. As Figure 2 shown, the method includes the following steps:
[0065] Step S202, according to the collected input statement, obtain the semantic entity in the input statement, where the semantic entity is used to represent the semantics in the input statement;
[0066] In step S202 of the present application, the input statement can be obtained by receiving the user's voice through a voice receiving device, and then converting the user's voice into the user's input text by means of voice conversion. For example, the above voice is recognized as text by using ASR technology. The above sample data can be the user's voice or the text corresponding to the user's voice.
[0067] Among them, the step of obtaining the semantic entity in the input statement includes the following:
[0068] Perform entity recognition on the input statement to obtain semantic entities. The above input statement can be first passed through an entity parsing module to recognize the semantic entities in the sample data. Specifically, the text of the above input statement is input into the entity parsing module, and the entity parsing module outputs the corresponding semantic entities.
[0069] In step S204, based on the input statement and the semantic entities, obtain the feedback information corresponding to the input statement through the dialogue model. The feedback information includes at least one of the following: a reply to the input statement, and a third-party service requested to be invoked in the input statement.
[0070] In step S204 of the present application, based on the input statement and the semantic entities obtained in step S202, obtain the feedback information corresponding to the input statement through the dialogue model. The feedback information may include: a response statement to the user's question and answer, or a reply to the user's conversation, or a response to the user's external service requirement (i.e., the third-party service in the embodiments of the present application).
[0071] Among them, the third-party server may include at least one of: weather query, online shopping, smart home appliance control, navigation, and online ticket booking.
[0072] In the embodiments of the present application, the presentation forms of the feedback information include the following ways:
[0073] Way 1: When the feedback information is a response statement to the user's question and answer;
[0074] For example, when the input statement is: "Hello, are you there?" The feedback information can be: "Yes, please go ahead."
[0075] Way 2: When the feedback information is the third-party service requested to be invoked in the input statement;
[0076] For example, when the input statement is: "Weather in place A tomorrow", the feedback information can be to call the third-party weather query service and feedback: "Weather in place A tomorrow: sunny; gentle breeze,...". And it can be extended to the weather, travel advice, and clothing advice in place A for the next N days.
[0077] Way 3: When the feedback information is a response statement to the user's question and answer and the third-party service requested to be invoked in the input statement;
[0078] Specifically, combining Way 1 and Way 2, for example, when the input statement is: "Hello, weather in place A tomorrow", the feedback information is to call the third-party weather query server and feedback: "Do you want to know the weather in place A tomorrow? Weather in place A tomorrow: sunny; gentle breeze,...".
[0079] It should be noted that the feedback information in the embodiments of the present application is only described by taking the above examples as an example, and is subject to the interactive method provided by the embodiments of the present application, and is not specifically limited.
[0080] In the embodiments of the present invention, according to the obtained input statement, a semantic entity in the input statement is obtained, where the semantic entity is used to represent the semantics in the input statement; according to the input statement and the semantic entity, a feedback information corresponding to the input statement is obtained through a dialogue model, where the feedback information includes: a reply to the input statement, and / or, a third-party service requested to be invoked in the input statement, achieving the purpose of directly outputting the feedback information through the dialogue model, thereby realizing the technical effect of avoiding the logical code writing steps in the traditional method, and further solving the technical problem of large code demand in the process of dialogue decision-making and implementation in the prior art.
[0081] Optionally, obtaining the semantic entity in the input statement in step S202 includes: performing word segmentation on the input statement to obtain at least one word segment constituting the input statement; obtaining the category to which at least one word segment belongs, and determining the semantic entity in the input statement according to the category.
[0082] Among them, the language features in the input statement may include status features, entity features, semantic features, and word features. The above status features may be the user's personalized tags and behavior tags. For example, whether the user is a new user, and whether the user's gender is male or female. The above status features can be obtained according to the user's information and can be recorded in the form of 0 / 1. The above entity features may be entities mentioned by the user in the current interaction.
[0083] For example, the user says: I like to listen to "XXX" by singer A.
[0084] Among them, singer A belongs to the singer, and "XXX" belongs to the song name. The above semantic features may be the vectorized representation obtained by the model for the user's current input. The above word features may be the specific words in the voice or text input by the current user.
[0085] The above language features may also include other features such as bigram and word vector, so as to enrich the language features and enable the dialogue model to more accurately determine the corresponding feedback information.
[0086] After obtaining the sample data, the language features are determined from the sample data. Since the language features include multiple contents, different contents are obtained in different ways. The above status features can be obtained through user information.
[0087] The above entity features can be determined based on sample data. Specifically, the above sample data can be first passed through an entity parsing module to identify the entities in the sample data and associate the entities with the entity nodes in the knowledge graph to determine the entity features.
[0088] Optionally, in step S204, obtaining the feedback information corresponding to the input statement through the dialogue model based on the input statement and semantic entities includes: obtaining the historical interaction result; obtaining the dialogue vector corresponding to the input statement based on the input statement, semantic entities, and historical interaction result; and obtaining the feedback information corresponding to the dialogue vector through the dialogue model.
[0089] Among them, the dialogue vector is obtained by associating the language features, semantic entities in the input statement with at least one historical dialogue data (i.e., the historical interaction result in the embodiments of the present application).
[0090] The dialogue vector under the above language features of the current dialogue is generated by associating the language features with the historical dialogue data of the previous time. Specifically, the result of the GRU obtained through the language features and the previous dialogue interaction can be input into the GRU to obtain the dialogue vector of the current dialogue interaction.
[0091] The above historical dialogue can be the dialogue interaction between the system and the user in the previous round, including the user's question and the system's answer to the user's question. The question and the answer constitute the dialogue interaction in the previous round. In the previous round of system dialogue, the result of the GRU corresponding to the question is determined according to the language features of the user's question. In the current round of dialogue interaction, the input of the GRU is the above language features and the result of the GRU of the previous dialogue interaction.
[0092] As an optional embodiment, obtaining the dialogue vector by associating the language features with at least one historical dialogue data includes: associating the state features, entity features, semantic features, and word features in the language features with at least one historical dialogue data to obtain a dialogue vector including the language features and at least one historical dialogue data.
[0093] The above historical dialogue can include the user's historical questions and the system's replies corresponding to the historical questions. One question and one reply constitute a historical dialogue. The above dialogue vector includes the above language features and the above historical dialogue data.
[0094] The above historical dialogue data can be the previous dialogue of the current historical dialogue. The closer the distance is to the current dialogue, the closer the association between the two dialogues may be, and the greater the influence on each other. Therefore, the above historical dialogue can be the historical dialogue closest to the current dialogue.
[0095] Further, optionally, the interaction method provided by the embodiments of the present application further includes: after obtaining the feedback information, using the feedback information and the input statement as updated data to input into the dialogue model to obtain an optimized dialogue model.
[0096] Specifically, the above dialogue model is optimized through a dialogue tree (i.e., a tree-like interaction structure composed of the feedback information, input statement, and historical interaction dialogues included in the embodiments of the present application). The dialogue tree is the data source, and what is optimized through the dialogue tree is the overall dialogue model. That is, from the dialogue tree, the interaction of each dialogue can be obtained, including the user's input question and the system's actions. The dialogue data of the dialogue tree is used as training samples to optimize the above overall end-to-end dialogue model.
[0097] Optionally, the dialogue model obtains the language features of the input statement and obtains the corresponding dialogue vector according to the language features; according to the language features and the dialogue vector, obtains the category label and action features of the input statement; according to the category label and action features, generates the feedback information corresponding to the input statement; and is trained according to the historical dialogue containing the feedback information for the process of obtaining the language features, dialogue vector, category label, and action features of the input statement.
[0098] Specifically, the training process of the dialogue model provided by the embodiments of the present application is as follows:
[0099] 1). The user's input text first passes through an entity parsing module to identify entities and associate the entities with the entity nodes in the knowledge graph. Among them, the construction method of the knowledge graph can be through crawling online data, such as music information, including singers, album names, composers, song names, etc. The nodes in the knowledge graph represent entities, such as singers and songs, and the edges in the knowledge graph represent relationships, such as a singer "performing" a song.
[0100] 2). The user's input text can obtain the semantic representation of the text through semantic modeling, that is, semantic features. In this system, the semantic representation of the input text is output through a bidirectional LSTM sentence encoder.
[0101] 3). The user's input state, together with the results of the above two steps, can obtain the features of the user's entire input, including: state features, entity features, semantic features, and word features, etc.; the input state includes the user's personalized and behavioral features, such as: whether the user is a new user, whether it is the first time to enter, male or female, etc.
[0102] The above state features are the user's personalized labels and behavioral labels, such as whether the user is a new user, male or female, etc., and are input in the form of a 0 / 1 vector;
[0103] The above entity features represent the entities mentioned by the user in the current interaction. For example, if the user mentions a singer or a song, they are also input in the form of a 0 / 1 vector;
[0104] The above semantic features are the vectorized representations obtained from the user's current input through the model;
[0105] The above word features are the words mentioned in the question input by the current user and are input in the form of a 0 / 1 vector.
[0106] 4). The features in the third step and the results of the GRU obtained from the previous interaction can obtain the vector representation of the current interaction through GRU (Gated Recurrent Unit), which is also the dialogue vector. The input of GRU is the four features in the third step: state feature, entity feature, semantic feature, and word feature; the output is the hidden layer vector, that is, the vector representation. The previous interaction is the interaction between the system and the user in the previous round, including the user's input question and the system's reply, and these two constitute a round of dialogue.
[0107] 5). The vector representation of the current interaction obtained in the fourth step passes through the classification module and the action pre-decision module that takes the aforementioned features as input to obtain the final output of the system, that is, the system's reply or the external service required by the request. The classification module is a classifier and can be considered as the category label scores obtained by decoding from the aforementioned vector. The input is the aforementioned interaction vector, and the output is the scores of the dialogue action category labels. The input of the action decision module is the four types of features obtained in the aforementioned third step and the vector representation of GRU obtained in the previous round; the output represents the action.
[0108] 6). Finally, optimize the above modules according to the overall data of the dialogue tree so that the actions of the dialogue data are consistent with the actions predicted by the above system. The dialogue tree refers to a tree-structured dialogue, where each node represents what the user said and how the system replied, and the next layer of nodes also corresponds to what the user said and how the system replied. The dialogue tree is the data source, and what is optimized through the dialogue tree is the overall dialogue model. That is, from the dialogue tree, the interactions of each dialogue, including the user's input question and the system's actions, can be obtained, and this data is used as training samples to optimize the above overall end-to-end dialogue model.
[0109] More features can also be used in the feature part in step 3) above, such as features like BiGram (binary word segmentation) and wordvector (word vector). In addition to GRU (Gated Recurrent Unit) in step 4), other commonly used units in recurrent neural networks such as LSTM can also be used.
[0110] Embodiment 2
[0111] This application provides an Figure 3 interaction method as Figure 3is a flowchart of an interaction method according to Embodiment 2 of the present invention. As Figure 3 shown, the method includes the following steps:
[0112] Step S302, collect input statements input by multiple objects, and obtain semantic entities in the input statements, where the semantic entities are used to represent the semantics in the input statements;
[0113] In the above step S302 of the present application, to obtain the input statements input by multiple objects, the user voice can be received through a voice receiving device, and then the user voice can be converted into the user's input text by means of voice conversion. For example, the above voice can be recognized as text by using ASR technology. The above sample data can be the user's voice or the text corresponding to the user's voice.
[0114] Among them, the step of obtaining semantic entities in the input statements includes the following:
[0115] Perform entity recognition on the input statements to obtain semantic entities. The above input statements can be first passed through an entity parsing module to recognize the semantic entities in the sample data. Specifically, the text of the above input statements is input into the entity parsing module, and the corresponding semantic entities are output by the entity parsing module.
[0116] Step S304, based on the semantic entities of different objects, determine the language features of different objects;
[0117] In the above step S304 of the present application, the above language features may include status features, entity features, semantic features, and word features. The above status features may be the user's personalized tags and behavior tags. For example, whether the user is a new user, and whether the user's gender is male or female. The above status features can be obtained according to the user's information and can be recorded in the form of 0 / 1. The above entity features may be the entities mentioned by the user in the current interaction.
[0118] For example, the user says: I like to listen to "XXX" by singer A.
[0119] Among them, singer A is a singer, and "XXX" is the name of a song. The above semantic features may be the vectorized representation obtained by passing the user's current input through a model. The above word features may be the specific words in the voice or text input by the current user.
[0120] The above language features may also include other features such as bigram and word vector, so as to enrich the language features and enable the dialogue model to more accurately determine the corresponding feedback information.
[0121] After obtaining the input statement, language features are determined from the input statement. Since the language features include multiple contents, different contents are obtained in different ways. The above state features can be obtained through user information.
[0122] The above entity features can be determined according to the input statement. Specifically, the above input statement can be first passed through an entity parsing module to identify the entities in the input statement and associate the entities with the entity nodes in the knowledge graph to determine the entity features.
[0123] The above semantic features can be determined by a bidirectional LSTM (Long Short-Term Memory) sentence encoder. The input statement is input into the above LSTM sentence encoder, and the semantic features of the input statement are output by the above LSTM sentence encoder.
[0124] As an optional embodiment, to determine the semantic features from the input statement, a semantic recognition model can also be established first. The above semantic recognition model can be a deep learning model or a machine learning model, trained with the input statement and the corresponding semantic features, and then the input statement is input into the semantic recognition model, and the semantic features are output by the semantic recognition model.
[0125] The above word features can directly perform text recognition on the text of the above input statement to determine the word features in the text.
[0126] Obtaining the corresponding dialogue vector according to the language features can be achieved by associating the language features with the previous historical dialogue data to generate the dialogue vector under the above language features of the current dialogue. Specifically, the result of the GRU obtained through the interaction between the language features and the previous dialogue can be input into the GRU to obtain the dialogue vector of the current dialogue interaction.
[0127] The above historical dialogue can be the dialogue interaction between the system and the user in the previous round, including the user's question and the system's answer to the user's question. The question and the answer constitute the dialogue interaction in the previous round. In the previous round of system dialogue, the result of the GRU corresponding to the question is determined according to the language features of the user's question. In the current round of dialogue interaction, the input of the GRU is the above language features and the result of the GRU of the previous dialogue interaction.
[0128] Step S306, based on the language features of different objects, call different dialogue models;
[0129] Step S308, process the input statements input by different objects with the corresponding dialogue models to generate the feedback information corresponding to different objects, where the feedback information includes at least one of the following: the feedback information on the input statement, the third-party service requested to be called in the input statement.
[0130] In step S308 of the present application, based on the input statements and semantic entities input by different objects obtained in step S302, feedback information corresponding to different objects is obtained through a dialogue model. The feedback information may include: an answer statement to a user's question and answer, or a reply to a user's conversation, or a response to a user's need for an external service (i.e., the third-party service in the embodiments of the present application).
[0131] Among them, the third-party server may include at least one of: weather query, online shopping, smart home appliance control, navigation, and online ticket booking.
[0132] The manifestation forms of the feedback information in the embodiments of the present application include the following ways:
[0133] Way 1: When the feedback information is an answer statement to a user's question and answer;
[0134] For example, if the input statement is: "Hello, are you there?" The feedback information can be: "Yes, please go ahead."
[0135] Way 2: When the feedback information is the third-party service requested to be called in the input statement;
[0136] For example, if the input statement is: "The weather in place A tomorrow", the feedback information can be fed back by calling the third-party weather query service: "The weather in place A tomorrow: sunny; gentle breeze,...". And it can be extended to the weather, travel advice, and clothing advice in the next N days in place A.
[0137] Way 3: When the feedback information is an answer statement to a user's question and answer and the third-party service requested to be called in the input statement;
[0138] Specifically, combining Way 1 and Way 2, for example, if the input statement is: "Hello, the weather in place A tomorrow", the feedback information is fed back by calling the third-party weather query server as: "Do you want to know the weather in place A tomorrow? The weather in place A tomorrow: sunny; gentle breeze,...".
[0139] It should be noted that the feedback information in the embodiments of the present application is only illustrated by the above examples, and is subject to the interactive method provided by the embodiments of the present application, and is not specifically limited.
[0140] In an embodiment of the present invention, input statements input by multiple objects are collected, and semantic entities in the input statements are obtained, where the semantic entities are used to represent the semantics in the input statements; based on the semantic entities of different objects, the language features of different objects are determined; based on the language features of different objects, different dialogue models are called; the input statements input by different objects are processed using the corresponding dialogue models to generate feedback information corresponding to different objects, where the feedback information includes: feedback information on the input statement, and / or, a third-party service requested to be called in the input statement, achieving the purpose of directly outputting feedback information through the dialogue model, thereby realizing the technical effect of avoiding the logical code writing steps in the traditional method, and further solving the technical problem of large code requirements in the process of dialogue decision-making and implementation in the prior art.
[0141] Embodiment 3
[0142] This application provides an Figure 4 interactive method as shown. Figure 4 is a flowchart of an interactive method according to Embodiment 3 of the present invention, as Figure 4 shown, the method includes the following steps:
[0143] Step S402, obtain the language features of the sample data, and obtain the corresponding dialogue vector according to the language features;
[0144] The sample data can be obtained by receiving the user's voice through a voice receiving device, and then converting the user's voice into the user's input text by means of voice conversion. For example, the above voice is recognized as text using ASR technology. The above sample data can be the user's voice or the text corresponding to the user's voice.
[0145] The above language features may include state features, entity features, semantic features, and word features. The above state features may be the user's personalized tags and behavior tags. For example, whether the user is a new user, and whether the user's gender is male or female. The above state features can be obtained according to the user's information and can be recorded in the form of 0 / 1. The above entity features may be entities mentioned by the user in the current interaction.
[0146] For example, the user says: I like to listen to "XXX" by singer A.
[0147] Among them, singer A is a singer, and "XXX" is the name of a song. The above semantic features may be the vectorized representation obtained by the model for the user's current input. The above word features may be the specific words in the voice or text input by the current user.
[0148] The above language features may also include other features such as BiGram and word vector, so as to enrich the language features and enable the dialogue model to more accurately determine the corresponding feedback information.
[0149] After obtaining the sample data, language features are determined from the sample data. Since the language features include multiple contents, different contents are obtained in different ways. The above state features can be obtained through user information.
[0150] The above entity features can be determined according to the sample data. Specifically, the above sample data can be first passed through an entity parsing module to identify the entities in the sample data and associate the entities with the entity nodes in the knowledge graph to determine the entity features.
[0151] The above semantic features can be determined by a bidirectional LSTM (Long Short-Term Memory) sentence encoder. The sample data is input into the above LSTM sentence encoder, and the semantic features of the sample data are output by the above LSTM sentence encoder.
[0152] As an alternative embodiment, to determine the semantic features from the sample data, a semantic recognition model can also be established first. The above semantic recognition model can be a deep learning model or a machine learning model, which is trained with the sample data and the corresponding semantic features, and then the sample data is input into the semantic recognition model, and the semantic features are output by the semantic recognition model.
[0153] The above word features can directly perform text recognition on the text of the above sample data to determine the word features in the text.
[0154] Obtaining the corresponding dialogue vector according to the language features can be achieved by associating the language features with the previous historical dialogue data to generate the dialogue vector under the above language features of the current dialogue. Specifically, the result of the GRU obtained from the language features and the previous dialogue interaction can be input into the GRU to obtain the dialogue vector of the current dialogue interaction.
[0155] The above historical dialogue can be the dialogue interaction between the system and the user in the previous round, including the user's question and the system's answer to the user's question. The question and the answer form the dialogue interaction in the previous round. In the previous round of system dialogue, the result of the GRU corresponding to the question is determined according to the language features of the user's question. In the current round of dialogue interaction, the input of the GRU is the above language features and the result of the GRU of the previous dialogue interaction.
[0156] As an alternative embodiment, obtaining the language features of the sample data includes: respectively performing entity recognition and semantic recognition on the sample data to obtain the language features.
[0157] The above language features may include status features, entity features, semantic features, and word features. Among them, the status features and word features can be directly determined and obtained through user information and the text of sample data respectively. The above entity features and semantic features need to identify the sample data. Specifically, the entity features need to obtain through entity recognition of the sample data; the semantic features need to obtain through semantic recognition of the sample. Therefore, by performing entity recognition and semantic recognition on the sample data respectively, the above language features can be obtained.
[0158] In this embodiment, for entity recognition of the sample data, the sample data can be first passed through an entity parsing module to identify the entities in the sample data. For example, when the user says "I like to listen to the song 《XXX》 by singer A", the entities included are: singer A is a singer, and 《XXX》 is the name of a song.
[0159] For semantic recognition of the sample data, it can be recognized through a semantic recognition model. The above speech recognition model can be a deep learning model, a machine learning model, a sentence encoder, etc. In this embodiment, the above sample data can be semantically recognized through a bidirectional LSTM (Long Short-Term Memory) sentence encoder.
[0160] As an alternative embodiment, performing entity recognition and semantic recognition on the sample data respectively, the obtained language features include: performing entity recognition on the sample data to obtain semantic entities; matching the semantic entities with preset data to obtain a first recognition result; performing semantic recognition on the sample data to obtain a second recognition result; obtaining language features based on the first recognition result, the second recognition result, and the status of the input sample data. Among them, the language features include: status features, entity features, semantic features, and word features.
[0161] For entity recognition of the sample data to obtain semantic entities, the sample data can be first passed through an entity parsing module to identify the semantic entities in the sample data. Specifically, the text of the above sample data is input into the entity parsing module, and the corresponding semantic entities are output by the entity parsing module.
[0162] The above preset data can be a knowledge graph. The knowledge graph can be created in advance by crawling online data, such as music information, including singers, album names, composers, song names, etc. The nodes in the knowledge graph represent entities, including singers, songs, etc. The edges in the knowledge graph represent the relationships between different nodes or different entities. For example, the edge between the A singer node and the B song node represents that singer A "sings" song B.
[0163] The above first recognition result may be the above entity features, including entities and the relationships between entities. By matching the semantic entities with the preset data to obtain the first recognition result, it may be to associate the semantic entities of the sample data with the entity nodes in the knowledge graph to obtain entity features, which are recorded in the knowledge graph, including entities and the relationships between entities.
[0164] Perform semantic recognition on the sample data to obtain a second recognition result. The above second recognition result may be semantic features. In this embodiment, it may be determined by a bidirectional LSTM (Long Short-Term Memory) sentence encoder. Input the sample data into the above LSTM sentence encoder, and the semantic features of the sample data are output by the above LSTM sentence encoder.
[0165] Based on the first recognition result, the second recognition result, and the state of the input sample data, the language features are obtained. The above first recognition result is the entity features obtained by entity recognition, the second recognition result is the semantic features obtained by semantic recognition, and from the state of the input sample data, the state features of the sample can be obtained, and word features can be obtained according to the text of the above sample data.
[0166] The above language features may also include features obtained by other methods for determining feedback information, such as binary word segmentation, word vectors, and other features.
[0167] As an optional embodiment, obtaining the corresponding dialogue vector according to the language features includes: associating the language features with at least one historical dialogue data to obtain the dialogue vector.
[0168] Obtaining the corresponding dialogue vector according to the language features can be achieved by associating the language features with the previous historical dialogue data to generate the dialogue vector under the above language features of the current dialogue. Specifically, the result of the GRU obtained from the language features and the previous dialogue interaction can be input into the GRU to obtain the dialogue vector of the current dialogue interaction.
[0169] The above historical dialogue may be the dialogue interaction between the system and the user in the previous round, including the user's question and the system's answer to the user's question. The question and the answer form the previous round of dialogue interaction. In the previous round of system dialogue, the result of the GRU corresponding to the question is determined according to the language features of the user's question. In the current round of dialogue interaction, the input of the GRU is the above language features and the result of the GRU of the previous dialogue interaction.
[0170] As an alternative embodiment, associating language features with at least one historical dialogue data to obtain a dialogue vector includes: associating state features, entity features, semantic features, and word features in the language features with at least one historical dialogue data to obtain a dialogue vector that includes language features and at least one historical dialogue data.
[0171] The above historical dialogue can include the user's historical questions and the corresponding system responses. One question and one response constitute one historical dialogue. The above dialogue vector includes the above language features and the above historical dialogue data.
[0172] The above historical dialogue data can be the previous dialogue of the current historical dialogue. The closer the distance to the current dialogue, the closer the association between the two dialogues may be, and the greater the influence on each other. Therefore, the above historical dialogue can be the historical dialogue closest to the current dialogue.
[0173] Step S404, obtaining the category label and action feature of the sample data according to the language features and the dialogue vector;
[0174] The above category label can be a dialogue action category label, for example, asking, conversing, querying, etc. The above category label can also be a dialogue language category label. The above action feature can be answering questions, responding to conversations, querying information, etc.
[0175] As an alternative embodiment, obtaining the category label and action feature of the sample data according to the language features and the dialogue vector includes: obtaining the category label of the sample data according to the dialogue vector; obtaining the action feature of the sample data according to the language features and the dialogue vector.
[0176] In this embodiment, the above category label can be a dialogue action category label. Obtaining the category label of the sample data according to the dialogue vector can be achieved through a classifier. Specifically, taking the dialogue vector as input and inputting it into the classifier to obtain the category label corresponding to the dialogue vector. For example, decoding the above dialogue vector through the classifier to obtain the scores of the class labels, and determining the category label corresponding to the dialogue vector according to the scores of different category labels.
[0177] Obtaining the action feature of the sample data according to the language features and the dialogue vector can be achieved through an action decision module. Specifically, taking the above language features and the dialogue vector as input and inputting them into the action decision module to obtain the action feature. According to this action feature, the final output of the system is obtained, that is, the system's response to the user or the external service requested by the user is provided.
[0178] As an alternative embodiment, obtaining the category label of the sample data according to the dialogue vector includes: classifying the semantic labels in the dialogue vector to obtain the category label of the sample data.
[0179] Step S406: Generate feedback information for the corresponding sample data based on the category label and action feature;
[0180] Obtain the feedback information of the system on the sample data input by the user finally according to the category label and action feature, including the answer statement to the user's question and answer, or the reply to the user's conversation, or the response to the user's request for an external service.
[0181] Step S408: Train the process of obtaining the language feature, dialogue vector, category label and action feature of the sample data according to the historical conversation containing the feedback information to obtain a dialogue model, where the dialogue model is used to obtain the feedback information corresponding to the input statement, and the feedback information includes at least one of the following: the feedback information on the input statement, the third-party service requested to be called in the input statement.
[0182] Train according to the above historical conversation, language feature, dialogue vector, category label and action feature to obtain a dialogue model. Specifically, the model part for obtaining the language feature can be trained through the sample data and language feature, the model part for obtaining the dialogue vector can be trained through the language feature and dialogue vector, the classifier can be trained through the language feature and category label, and the action decision module can be trained through the language feature, dialogue vector and action feature, so as to realize the distributed training of the model and establish a dialogue model.
[0183] As an optional embodiment, the above dialogue model can also be optimized through a dialogue tree. The dialogue tree is the data source, and what is optimized through the dialogue tree is the whole of the dialogue model. That is, the interaction of each dialogue can be obtained from the dialogue tree, including the user's input question and the system's action. The dialogue data of the dialogue tree is used as the training sample to optimize the above overall end-to-end dialogue model.
[0184] The manifestation form of the feedback information in the embodiment of the present application includes the following ways:
[0185] Way 1: When the feedback information is the answer statement to the user's question and answer;
[0186] For example, if the input statement is: "Hello, are you there?" The feedback information can be: "Yes, please go ahead."
[0187] Way 2: When the feedback information is the third-party service requested to be called in the input statement;
[0188] For example, if the input statement is: "Weather in Area A tomorrow", the feedback information can be to call the third-party weather query service and feedback: "Weather in Area A tomorrow: sunny; gentle breeze,...". And it can be extended to the weather, travel advice and dressing advice in Area A in the next N days.
[0189] Method 3: when the feedback information is the answer statement of the user's question and answer and the third-party service requested to be called in the input statement;
[0190] Specifically, combining Method 1 and Method 2, for example, the input statement is: "Hello, what's the weather in place A tomorrow", and the feedback information is obtained by calling a third-party weather query server and is feedback as: "Do you want to know the weather in place A tomorrow? The weather in place A tomorrow: sunny; gentle breeze, ……".
[0191] It should be noted that the feedback information in the embodiments of the present application is only illustrated by the above examples, and is subject to the interactive method provided by the embodiments of the present application, and is not specifically limited.
[0192] According to the above steps, by obtaining the language features of the sample data and obtaining the corresponding dialogue vectors based on the language features; based on the language features and dialogue vectors, obtaining the category labels and action features of the sample data; based on the category labels and action features, generating the feedback information corresponding to the sample data; based on the historical dialogue pairs containing the feedback information, training the process of obtaining the language features, dialogue vectors, category labels and action features of the sample data, and obtaining the dialogue model. By using the language features, dialogue vectors, category labels and action features of the sample data, and the corresponding feedback information to train the model, the purpose of directly outputting the feedback information through the dialogue model is achieved, thus realizing the technical effect of avoiding the logical code writing steps in the traditional method, and further solving the technical problem of the large demand for code in the process of dialogue decision-making and implementation in the prior art.
[0193] As an optional embodiment, the interactive method is applied to an end-to-end interactive system, where the end-to-end interactive system includes: an online interactive system, an intelligent interactive device, and a device embedded in the intelligent interactive device.
[0194] Specifically, the input statement can be received by a voice receiving device to receive the user's voice, and then the user's voice is converted into the user's input text by means of voice conversion. For example, the above voice is recognized as text by using ASR technology. The above input information can be the user's voice or the text corresponding to the user's voice. The input statement can also be received by an interactive device to receive the text input by the user.
[0195] The above language features can include status features, entity features, semantic features and word features. The above status features can be the user's personalized label and behavior label. For example, whether the user is a new user, and whether the user's gender is male or female. The above status features can be obtained according to the user's information. The above entity features can be the entity mentioned by the user in the current interaction. The above semantic features can be the vectorized representation obtained by the model for the user's current input. The above word features can be the specific vocabulary in the voice or text input by the current user.
[0196] After obtaining the input statement, language features are determined from the input statement. Since language features include multiple contents and different contents are obtained in different ways. The above state features can be obtained through user information. The above entity features can be determined according to the input statement. Specifically, the above input statement can be first passed through an entity parsing module to identify the entities in the input statement and associate the entities with the entity nodes in the knowledge graph to determine the entity features. The above semantic features can be determined by a bidirectional LSTM sentence encoder. The input statement is input into the above LSTM sentence encoder, and the semantic features of the input statement are output by the above LSTM sentence encoder. The above word features can directly perform text recognition on the text of the above input statement to determine the word features in the text.
[0197] Obtaining the corresponding dialogue vector according to the language features can be achieved by associating the language features with the previous historical dialogue data to generate the dialogue vector under the above language features of the current dialogue. Specifically, the result of the GRU obtained through the interaction between the language features and the previous dialogue can be input into the GRU to obtain the dialogue vector of the current dialogue interaction. The above historical dialogue can be the dialogue interaction between the system and the user in the previous round, including the user's question and the system's answer to the user's question. The question and the answer form the dialogue interaction in the previous round. In the previous system dialogue, the result of the GRU corresponding to the question is determined according to the language features of the user's question. In the current round of dialogue interaction, the input of the GRU is the above language features and the result of the GRU of the previous dialogue interaction.
[0198] Obtaining the class label and action features of the sample data based on the language features and the dialogue vector includes: obtaining the class label of the sample data based on the dialogue vector; obtaining the action features of the sample data based on the language features and the dialogue vector. Obtaining the class label of the sample data based on the dialogue vector can be achieved through a classifier. Specifically, the dialogue vector is used as the input and input into the classifier to obtain the class label corresponding to the dialogue vector. For example, the score of the class label is decoded from the above dialogue vector through the classifier, and the class label corresponding to the dialogue vector is determined according to the scores of different class labels.
[0199] Obtaining the action features of the sample data based on the language features and the dialogue vector can be achieved through an action decision module. Specifically, the above language features and the dialogue vector are used as the input and input into the action decision module to obtain the action features. According to the action features, the final output of the system is obtained, that is, the system's reply to the user or the external service requested by the user is provided.
[0200] Based on the category label and action feature, obtain the final feedback information of the system for the sample data input by the user, including the answer statement to the user's question and answer, or the reply to the user's conversation, or the response to the user's external service requirement.
[0201] Train a dialogue model according to the above historical conversation, language feature, dialogue vector, category label and action feature. Specifically, the model part for obtaining language features can be trained through sample data and language features, the model part for obtaining dialogue vectors can be trained through language features and dialogue vectors, the classifier can be trained through language features and category labels, and the action decision module can be trained through language features, dialogue vectors and action features, so as to realize distributed training of the model and establish a dialogue model.
[0202] According to the above steps, by receiving the input statement; obtaining the feedback information corresponding to the input statement through the dialogue model; wherein, the dialogue model obtains the language feature of the input statement and obtains the corresponding dialogue vector according to the language feature; according to the language feature and dialogue vector, obtain the category label and action feature of the input statement; according to the category label and action feature, generate the feedback information corresponding to the input statement; according to the historical conversation containing the feedback information, train the process of obtaining the language feature, dialogue vector, category label and action feature of the input statement, and train the model through the language feature, dialogue vector, category label and action feature of the sample data, as well as the corresponding feedback information, achieving the purpose of directly outputting the feedback information through the dialogue model, thus realizing the technical effect of avoiding the logical code writing steps in the traditional method, and further solving the technical problem of large code demand in the process of dialogue decision-making and implementation in the prior art.
[0203] It should be noted that this embodiment also provides an optional implementation manner, which will be described in detail below.
[0204] In this implementation manner, a dialogue interaction model is learned from the dialogue tree through an end-to-end sequence classification model, where end-to-end means that the required output is directly obtained from the input without passing through the pipeline processing. This system only includes two algorithm modules. Specifically, 1. The entity recognition module, which recognizes the entity in the user's question (i.e., the input statement), for example, knowing that singer A in "I want to listen to singer A's 'XXX'" is a singer and 'XXX' is a song; 2. The dialogue interaction model, which directly obtains the action executed by the system to the user according to the user's input statement and the recognized entity.
[0205] This embodiment can reduce the problem of error propagation caused by the module splitting of traditional methods, and directly obtain the response from the system to the user or the external service to be requested through an end-to-end dialogue interaction algorithm, avoiding the logical code writing of the dialogue decision module in traditional methods.
[0206] Figure 5 It is a schematic diagram of a dialogue system in an interaction method according to Embodiment 3 of the present invention. As Figure 5 shown, the main process of this embodiment is as follows:
[0207] 1). The input text of the user first passes through the entity parsing module to identify the entity and associate the entity with the entity node in the knowledge graph. Among them, the construction method of the knowledge graph can be through crawling online data, such as music information, including singers, album names, composers, song names, etc. The nodes in the knowledge graph represent entities, such as singers and songs, and the edges in the knowledge graph represent relationships, such as a singer "performing" a song.
[0208] 2). The input text of the user can obtain the semantic representation of the text through semantic modeling, that is, the semantic feature. In this system, the semantic representation of the input text is output through a bidirectional LSTM sentence encoder.
[0209] 3). The input state of the user, together with the results of the above two steps, can obtain the features of the entire input of the user, including: state features, entity features, semantic features, word features, etc.; the input state includes the personalized and behavioral features of the user, such as: whether the user is a new user, whether it is the first time to enter, male or female, etc.
[0210] The above state features are the personalized labels and behavioral labels of the user, such as whether the user is a new user, male or female, etc., and are input in the form of a 0 / 1 vector;
[0211] The above entity features represent the entities mentioned by the user in the current interaction, such as what singer and what song the user mentioned, and are also input in the form of a 0 / 1 vector;
[0212] The above semantic features are the vectorized representations obtained by the model for the current input of the user;
[0213] The above word features are which words are said in the current input question of the user and are input in the form of a 0 / 1 vector.
[0214] 4). The features of the third step and the results of the GRU obtained from the previous interaction can obtain the vector representation of the current interaction through the GRU (Gated Recurrent Unit), which is the dialogue vector. The input of the GRU is the four features in the third step: the state feature, the entity feature, the semantic feature, and the word feature. The output is the hidden layer vector, that is, the vector representation. The previous interaction is the interaction between the system and the user in the previous round, including the user's input question and the system's reply, and these two constitute a round of dialogue.
[0215] 5). The vector representation of the current interaction obtained in the fourth step passes through the classification module and the action pre-decision module that takes the aforementioned features as input to obtain the output of the final system, that is, the system's reply or the external service required by the request. The classification module is a classifier, which can be considered as the category label score obtained by decoding from the aforementioned vector. The input is the aforementioned interaction vector, and the output is the score of the dialogue action category label. The input of the action decision module is the four types of features obtained in the aforementioned third step, and the vector representation of the GRU obtained in the previous round; the output represents the action.
[0216] 6). Finally, optimize the above modules according to the overall dialogue tree data, in order to make the actions of the dialogue data consistent with the actions predicted by the above system. The dialogue tree refers to a tree-structured dialogue, where each node represents what the user said and how the system replied, and the next layer of nodes also corresponds to what the user said and how the system replied. The dialogue tree is the data source, and what is optimized through the dialogue tree is the overall dialogue model. That is, from the dialogue tree, the interactions of each dialogue, including the user's input question and the system's actions, can be obtained, and this data is used as a training sample to optimize the above overall end-to-end dialogue model.
[0217] For the feature part in the above step 3), more features can also be adopted, such as features like BiGram (binary word segmentation) and wordvector (word vector). In addition to the GRU (Gated Recurrent Unit) in the above step 4), units commonly used in recurrent neural networks such as LSTM can also be adopted.
[0218] The processing flow of the original Pipeline composed of multiple modules led to error propagation and could not be optimized as a whole. In this embodiment, the original three modules (actually more) are reduced to two modules, and overall optimization can be carried out; in the traditional method, code writing is required for dialogue interaction decision-making. In this embodiment, the dialogue decision model is directly learned from the dialogue data, avoiding code writing in dialogue decision-making, and at the same time ensuring the logical controllability of this method through the dialogue pre-decision module.
[0219] The key point of this embodiment is to model the system of the Pipeline with at least three modules in the traditional method into an end-to-end dialog interaction system that can be optimized as a whole; use the sequence learning algorithm to directly learn the decision-making actions given by the system in the dialog interaction, avoiding the need to write code for dialog decision-making, and greatly reducing the threshold.
[0220] This embodiment can be applied to the tensorflow basic development tool.
[0221] Example 4
[0222] This application provides an interactive method as Figure 6 shown. Figure 6 is a flowchart of an interactive method according to Embodiment 4 of the present invention, as Figure 6 shown, the method includes the following steps:
[0223] Step A, collect the input statement of the current input;
[0224] Among them, the method of collecting the input statement of the current input may include: collecting the voice of the user speaking through the voice collection device (for example, microphone) inside the terminal device, and converting the collected voice of the user speaking into the input statement of the input terminal device; and / or, collecting the input statement input by the user through a touch screen / external input device (such as: keyboard and / or mouse, or, a display screen with input, selection, and control functions).
[0225] Step B, call the dialog model to analyze the input statement of the current input, and display the feedback information with a semantic relationship with the input statement of the current input on the touch screen;
[0226] Among them, by calling the dialog model to analyze the input statement collected in Step A, the feedback information with a semantic relationship with the input statement is displayed on the touch screen. In the embodiment of this application, the touch screen can be used for display in addition to being an input device.
[0227] In Step B of this application above, by calling the dialog model, the collected input statement is semantically analyzed, and the semantic analysis result is matched with the corresponding feedback information. Among them, the feedback information may include: a reply to the input statement and / or a third-party service requested to be called in the input statement.
[0228] For example, user A says: "The weather in place A tomorrow." The terminal device collects the "The weather in place A tomorrow." for semantic analysis, and obtains the semantic analysis result that user A wants to know: tomorrow, place A, weather. The terminal device calls the third-party weather query service, feeds back the weather in place A tomorrow to user A, and can expand to display the weather in place A for the next N days.
[0229] Step C: Collect the input statement that is re - input based on the feedback information;
[0230] In step C of the present application, still taking the example in step B for illustration, if user A learns about the weather in place A tomorrow, the terminal device collects the input statement re - input by user A again. For example, if user A wants to know more information about place A, that is, the way to take a vehicle from the airport in place A to the hotel booked by user A, it can be expressed as "The vehicle - taking guide from the departure place, the airport in place A, to the hotel A booked by user A".
[0231] Step D: Invoke the dialogue model to analyze the re - input input statement, and display the feedback information semantically related to the re - input input statement on the touch screen;
[0232] In step D of the present application, based on the input statement in step C, by invoking the dialogue model again, the input statement is further analyzed. The terminal device learns through invoking the dialogue model that: the departure place of user A is the airport in place A, the destination is the hotel A booked by user A, and the semantic analysis result of the vehicle - taking guide from the airport in place A to hotel A. Then, according to the semantic analysis result, feedback information is generated. That is, the terminal device can provide feedback: bus route (including boarding and alighting points, travel time by bus, and ticket price), driving route, taxi route (including online taxi - hailing, total travel time, estimated price).
[0233] Step E: Loop and execute the above - mentioned steps C and D until an end instruction is received.
[0234] In step E of the present application, by looping and executing steps C and D, until the user completes the interaction with the terminal device and considers that this round of human - machine interaction ends.
[0235] In addition, the interaction record between user A and the terminal device is saved, and this interaction record is returned to the neural network as a sample for further learning, so as to enrich the number and type of sample learning of the neural network, provide more suitable services for the future interaction of user A, or provide reference interaction results for other users based on the interaction record of user A.
[0236] In the embodiments of the present application, the presentation forms of the feedback information include the following ways:
[0237] Way 1: When the feedback information is the answer statement to the user's question and answer;
[0238] For example, if the input statement is: "Hello, are you there?" The feedback information can be: "Yes, please go ahead."
[0239] Way 2: When the feedback information is the third - party service requested to be invoked in the input statement;
[0240] For example, if the input statement is: "Weather in Area A tomorrow", the feedback information can be obtained by calling a third-party weather query service and feedback: "Weather in Area A tomorrow: sunny; gentle breeze,...". It can be extended to the weather, travel advice, and clothing advice for the next N days in Area A.
[0241] Method 3: When the feedback information is the answer statement to the user's question and the third-party service requested to be called in the input statement;
[0242] Specifically, combining Method 1 and Method 2, for example, if the input statement is: "Hello, what's the weather in Area A tomorrow", the feedback information is obtained by calling a third-party weather query server and feedback: "Do you want to know the weather in Area A tomorrow? Weather in Area A tomorrow: sunny; gentle breeze,...".
[0243] It should be noted that the feedback information in the embodiments of the present application is only illustrated by the above examples, and is subject to the interactive method provided by the embodiments of the present application, and is not specifically limited.
[0244] It should be noted that the terminal device applied in the interactive method provided by the embodiments of the present application can be a smart speaker with a touch screen, or a smart mobile device installed with an APP containing a dialogue model. Among them, the smart mobile device can include: smart phones, tablets, laptops, desktop computers, smart wearable devices (smart watches, smart glasses, VR devices or AR devices).
[0245] The interactive method provided by the embodiments of the present application is illustrated by taking weather query and traffic travel query as examples. In addition, it can also include the control of indoor smart home, which is not specifically limited and is subject to the interactive method provided by the embodiments of the present application.
[0246] Embodiment 5
[0247] The present application provides an Figure 7 interactive method as shown. Figure 7 It is a flowchart of an interactive method according to Embodiment 5 of the present invention. As Figure 7 shown, on the server side, the method includes the following steps:
[0248] Step S702, the server receives the input statement collected by the smart device and obtains the semantic entity in the input statement, where the semantic entity is used to represent the semantics in the input statement;
[0249] In step S702 of the present application, the server can be a cloud server or a server cluster, and the cloud server cluster can include a distributed server cluster; the smart device can include: smart mobile terminals (such as smart phones), smart (small) appliances (such as smart speakers);
[0250] Specifically, taking the smart device as an example of a smart speaker, the server receives the input statement of user A collected by the smart speaker, and obtains the semantic entity in the input statement by analyzing the input statement of user A.
[0251] Step S704, the server obtains the feedback information matching the input statement through the dialogue model according to the semantic entity, where the feedback information includes at least one of the following: a reply to the input statement, a third-party service requested to be invoked in the input statement;
[0252] In step S704 of the present application, after obtaining the semantic entity, the server obtains the feedback information matching the input statement by invoking the dialogue model according to the semantic entity.
[0253] For example, the smart device collects that user A says: "The weather in place A tomorrow." The server obtains the semantic entities: time (tomorrow), location (place A), weather by obtaining the input statement "The weather in place A tomorrow";
[0254] Based on the obtained semantic entities above, determine the action corresponding to the input statement (asking about the weather) through the dialogue model, and thus obtain the feedback information corresponding to the input statement, such as: The weather in place A tomorrow is sunny, the ultraviolet index is X, gentle breeze, recommended clothing: #¥%……, and the weather of place A in the next N days can be displayed expandably.
[0255] Step S706, the server sends the feedback information to the smart device.
[0256] In step S706 of the present application, based on the feedback information obtained in step S704, it is returned to the smart device.
[0257] In the embodiments of the present application, the presentation form of the feedback information includes the following ways:
[0258] Way 1: When the feedback information is a reply statement to the user's question and answer;
[0259] For example, the input statement is: "Hello, are you there?" The feedback information can be: "Yes, please go ahead."
[0260] Way 2: When the feedback information is a third-party service requested to be invoked in the input statement;
[0261] For example, the input statement is: "The weather in place A tomorrow", the feedback information can be to call the third-party weather query service and feedback: "The weather in place A tomorrow: sunny; gentle breeze,...". And it can be expanded to the weather, travel opinions and clothing opinions of place A in the next N days.
[0262] Way 3: When the feedback information is a reply statement to the user's question and answer and a third-party service requested to be invoked in the input statement;
[0263] Specifically, combining Method 1 and Method 2, for example, the input statement is: "Hello, what's the weather like in Place A tomorrow", and the feedback information is obtained by calling a third-party weather query server, and the feedback is: "Do you want to know the weather in Place A tomorrow? The weather in Place A tomorrow: sunny; gentle breeze, ……".
[0264] It should be noted that the feedback information in the embodiments of the present application is only illustrated by the above examples, and is subject to the interactive method provided by the embodiments of the present application, and is not specifically limited.
[0265] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0266] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. The computer software product is stored in a non-volatile storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods of the various embodiments of the present invention.
[0267] Embodiment 6
[0268] According to an embodiment of the present invention, there is also provided an intelligent device, Figure 8a is a schematic diagram of an intelligent device according to Embodiment 6 of the present invention, as Figure 8a shown, including: a voice collection device 82a for collecting voice information; a processor 84a connected to the voice collection device 82a for identifying an input statement from the voice information and processing the semantic entity of the input statement through a dialogue model to generate corresponding feedback information, wherein the semantic entity of the input statement represents the semantics in the input statement.
[0269] Specifically, the intelligent device in the embodiments of the present application may include: a smart speaker, or an intelligent home appliance with a sound collection function, such as a TV, a refrigerator, an air conditioner, a washing machine, an electric pressure cooker;
[0270] When the intelligent device is an intelligent speaker, the intelligent speaker collects the user's voice information through a built-in voice collection device (or a connected mobile phone or tablet computer), identifies the input statement from the voice information through the voice collection device, analyzes the input statement through a dialogue model to obtain the semantic entity of the input statement, and obtains the corresponding action of the input statement based on the semantic entity (taking the query of weather as an example, the action corresponding to the input statement pair is query at this time), and finally generates the corresponding feedback information, for example, the future weather in Area A is #¥%.
[0271] Similarly, for intelligent home appliances with a sound collection function, the parsing method of voice information is the same as the above.
[0272] It should be noted that the display form of the feedback information in the embodiments of the present application is the same as that in any one of the examples in Embodiments 1 to 5, and will not be elaborated here.
[0273] Embodiment 7
[0274] According to an embodiment of the present invention, there is also provided an intelligent device. Figure 8b It is a schematic diagram of an intelligent device according to Embodiment 7 of the present invention, as Figure 8b shown, including: a touch screen 82b, configured to collect an input statement input through a touch interaction function, wherein the semantic entity of the input statement represents the semantics of the input statement; a processor 84b, connected to the touch screen 82b, configured to process the semantic entity of the input statement through a dialogue model to generate corresponding feedback information; wherein, the touch screen is further configured to display the feedback information.
[0275] Specifically, the intelligent device in the embodiments of the present application may include: an intelligent speaker with a touch screen, or an intelligent home appliance with a touch screen, such as a TV, a refrigerator, an air conditioner, a washing machine, an electric pressure cooker; wherein the touch screen may be an LED screen with an input function.
[0276] When the intelligent device is an intelligent speaker with a touch screen having an input function, the intelligent speaker collects the user's input statement through the touch screen, analyzes the input statement through the dialogue model in the processor to obtain the semantic entity of the input statement, and obtains the corresponding action of the input statement based on the semantic entity (taking the query of weather as an example, the action corresponding to the input statement pair is query at this time), and finally generates the corresponding feedback information, for example, the future weather in Area A is #¥%.
[0277] Similarly, for intelligent home appliances, the parsing method of the input statement is the same as the above.
[0278] It should be noted that the display form of the feedback information in the embodiments of the present application is the same as that in any one of the examples in Embodiments 1 to 5, and will not be elaborated here.
[0279] Embodiment 8
[0280] According to an embodiment of the present invention, there is also provided an intelligent device. Figure 8c It is a schematic diagram of an intelligent device according to Embodiment 8 of the present invention, as Figure 8c shown, including: a voice collection device 82c for collecting voice information; a touch screen 84c for displaying at least one statement generated by recognizing the voice information and receiving an input statement selected from the at least one statement through a touch interaction function, wherein the semantic entity of the input statement represents the semantics of the input statement; a processor 86c connected to the voice collection device and the touch screen for processing the semantic entity of the input statement through a dialogue model to generate corresponding feedback information; wherein, the touch screen is further used for displaying the feedback information.
[0281] Different from those in Embodiments 6 and 7, the intelligent device in the embodiment of the present application is an intelligent device that simultaneously has a voice collection device and a touch screen. In implementation, the voice collection device collects the voice information of the user, and the touch screen displays at least one statement corresponding to the voice information, and the dialogue model in the processor is used to identify and analyze the selected statement to generate corresponding feedback information.
[0282] It should be noted that the display form of the feedback information in the embodiment of the present application is the same as that in any one of Examples 1 to 5, and will not be elaborated here.
[0283] Embodiment 9
[0284] According to an embodiment of the present invention, there is also provided a device for implementing the interaction method of the above Embodiment 4. The device includes: a first collection module for collecting the currently input input statement; a second call module for calling a dialogue model to analyze the currently input input statement and displaying feedback information having a semantic relationship with the currently input input statement on the touch screen; a second collection module for collecting an input statement input again based on the feedback information; a second call module for calling the dialogue model to analyze the input statement input again and displaying feedback information having a semantic relationship with the input statement input again on the touch screen; an execution module for repeatedly executing the above second collection module and second call module until an end instruction is received.
[0285] Embodiment 10
[0286] According to an embodiment of the present invention, there is also provided an apparatus for implementing the interaction method in the above Embodiment 5. On the server side, the apparatus includes: a receiving module, configured to receive an input statement collected by an intelligent device and obtain a semantic entity in the input statement, where the semantic entity is used to represent the semantics in the input statement; a matching module, configured to obtain, according to the semantic entity, feedback information matching the input statement through a dialogue model, where the feedback information includes at least one of the following: a reply to the input statement, and a third-party service requested to be invoked in the input statement; a sending module, configured to send the feedback information to the intelligent device.
[0287] Embodiment 11
[0288] According to an embodiment of the present invention, there is also provided an apparatus for implementing the interaction method in the above Embodiment 1. Figure 9 is a schematic diagram of an interaction apparatus according to Embodiment XI of the present invention, as Figure 9 shown, the apparatus includes: an obtaining module 92, configured to obtain a semantic entity in the input statement according to the collected input statement, where the semantic entity is used to identify the semantics in the input statement; an interaction module 94, configured to obtain, according to the input statement and the semantic entity, feedback information corresponding to the input statement through a dialogue model, where the feedback information includes at least one of the following: a reply to the input statement, and a third-party service requested to be invoked corresponding to the input statement.
[0289] According to the above apparatus, through the obtaining module, which is configured to obtain a semantic entity in the input statement according to the obtained input statement, where the semantic entity is used to identify the semantics in the input statement; and the interaction module, which is configured to obtain, according to the input statement and the semantic entity, feedback information corresponding to the input statement through a dialogue model, the purpose of directly outputting feedback information through the dialogue model is achieved, thereby achieving the technical effect of avoiding the logical code writing steps in the traditional method, and further solving the technical problem of large code requirements in the process of dialogue decision-making and implementation in the prior art.
[0290] It should be noted here that the above obtaining module 92 and interaction module 94 correspond to steps S202 to S204 in Embodiment 1. The functions of the two modules and the corresponding steps are the same in terms of the implemented examples and application scenarios, but are not limited to the content disclosed in the above Embodiment 1.
[0291] Embodiment 12
[0292] According to an embodiment of the present invention, there is also provided an apparatus for implementing the interaction method in the above Embodiment 2. Figure 10 is a schematic diagram of an interaction apparatus according to Embodiment XII of the present invention, as Figure 10As shown in the figure, the device includes: an acquisition module 102, configured to acquire input statements input by multiple objects and obtain semantic entities in the input statements, where the semantic entities are used to represent the semantics in the input statements; a feature module 104, configured to determine language features of different objects based on the semantic entities of different objects; a call module 106, configured to call different dialogue models based on the language features of different objects; and an interaction module 108, configured to process the input statements input by different objects using corresponding dialogue models to generate feedback information corresponding to different objects, where the feedback information includes at least one of the following: feedback information on the input statement, and a third-party service requested to be called in the input statement.
[0293] According to the above device, through the acquisition module, which is configured to acquire input statements input by multiple objects and obtain semantic entities in the input statements, where the semantic entities are used to represent the semantics in the input statements; the feature module, which is configured to determine language features of different objects based on the semantic entities of different objects; the call module, which is configured to call different dialogue models based on the language features of different objects; and the interaction module, which is configured to process the input statements input by different objects using corresponding dialogue models to generate feedback information corresponding to different objects, where the feedback information includes: feedback information on the input statement, and / or, a third-party service requested to be called in the input statement, the purpose of directly outputting feedback information through the dialogue model is achieved, thereby realizing the technical effect of avoiding the logical code writing steps in the traditional method, and further solving the technical problem of large code requirements in the prior art during the process of dialogue decision-making and implementation.
[0294] It should be noted here that the above acquisition module 102, feature module 104, call module 106, and interaction module 108 correspond to steps S302 to S308 in Embodiment 2. The examples and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 2.
[0295] Embodiment 13
[0296] According to an embodiment of the present invention, there is also provided a device for implementing the interaction method of the above Embodiment 3. Figure 11 It is a schematic diagram of an interaction device according to Embodiment 13 of the present invention. As Figure 11 shown, the device includes: a first acquisition module 112, a second acquisition module 114, a statement generation module 116, and a model training module 118. The device will be described in detail below.
[0297] A first acquisition module 112, configured to acquire the language feature of the sample data and acquire the corresponding dialogue vector according to the language feature; a second acquisition module 114, connected to the first acquisition module 112, configured to acquire the category label and action feature of the sample data according to the language feature and the dialogue vector; a statement generation module 116, connected to the second acquisition module 114, configured to generate feedback information corresponding to the sample data according to the category label and the action feature; a model training module 118, connected to the statement generation module 116, configured to train the process of acquiring the language feature, dialogue vector, category label and action feature of the sample data according to the historical dialogue pair including the feedback information, to obtain a dialogue model, wherein the dialogue model is configured to acquire the feedback information corresponding to the input statement, and the feedback information includes at least one of the following: feedback information on the input statement, and a third-party service requested to be invoked in the input statement.
[0298] According to the above device, the first acquisition module acquires the language feature of the sample data and acquires the corresponding dialogue vector according to the language feature; the second acquisition module acquires the category label and action feature of the sample data according to the language feature and the dialogue vector; the statement generation module generates feedback information corresponding to the sample data according to the category label and the action feature; the model training module trains the process of acquiring the language feature, dialogue vector, category label and action feature of the sample data according to the historical dialogue pair including the feedback information, to obtain a dialogue model. In this way, by using the language feature, dialogue vector, category label and action feature of the sample data, and the corresponding feedback information to train the model, the purpose of directly outputting the feedback information through the dialogue model is achieved, thereby realizing the technical effect of avoiding the logical code writing steps in the traditional method, and further solving the technical problem of large code demand in the process of dialogue decision-making and implementation in the prior art.
[0299] It should be noted here that the above first acquisition module 112, second acquisition module 114, statement generation module 116 and model training module 118 correspond to steps S402 to S408 in Embodiment 3. The instances and application scenarios realized by the four modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 3.
[0300] Embodiment 14
[0301] According to an embodiment of the present invention, there is also provided an interaction system. Figure 12 FIG. 14 is a schematic diagram of an interaction system according to Embodiment 14 of the present invention. The system includes: an entity recognition module 1202, a semantic modeling module 1204, a feature acquisition module 1206, a gated recurrent module 1208, an action decision module 1210, a classification module 1212 and a dialogue reply module 1214. The system will be described in detail below.
[0302] Entity recognition module 1202, semantic modeling module 1204, feature acquisition module 1206, gated recurrent module 1208, action decision module 1210, classification module 1212, and dialogue response module 1214. Among them, the entity recognition module 1202 is used to perform entity recognition on the input statement to obtain semantic entities, and match the semantic entities with preset data to obtain a first recognition result; the semantic modeling module 1204 is used to perform semantic recognition on the input statement to obtain a second recognition result; the feature acquisition module 1206 is connected to the entity recognition module 1202 and the semantic modeling module 1204 respectively, and is used to obtain language features according to the first recognition result, the second recognition result, and the state of the input statement, where the language features include: state features, entity features, semantic features, and word features; the gated recurrent module 1208 is connected to the feature acquisition module 1206, and is used to associate with at least one historical dialogue data according to the state features, entity features, semantic features, and word features in the language features to obtain a dialogue vector including language features and at least one historical dialogue data; the action decision module 1210 is connected to the feature acquisition module 1206 and the gated recurrent module 1208, and is used to obtain the action features of the input statement according to the language features and the dialogue vector; the classification module 1212 is connected to the gated recurrent module 1208, and is used to obtain the category label of the input statement according to the dialogue vector; the dialogue response module 1214 generates feedback information corresponding to the input statement according to the category label and the action features.
[0303] According to the above system, the entity recognition module 1202 is used to perform entity recognition on the input statement to obtain semantic entities, and match the semantic entities with preset data to obtain a first recognition result; the semantic modeling module 1204 is used to perform semantic recognition on the input statement to obtain a second recognition result; the feature acquisition module 1206 is connected to the entity recognition module 1202 and the semantic modeling module 1204 respectively, and is used to obtain language features according to the first recognition result, the second recognition result, and the state of the input statement, where the language features include: state features, entity features, semantic features, and word features; the gated recurrent module 1208 is connected to the feature acquisition module 1206, and is used to associate with at least one historical conversation data according to the state features, entity features, semantic features, and word features in the language features to obtain a conversation vector including the language features and at least one historical conversation data; the action decision module 1210 is connected to the feature acquisition module 1206 and the gated recurrent module 1208, and is used to obtain the action features of the input statement according to the language features and the conversation vector; the classification module 1212 is connected to the gated recurrent module 1208, and is used to obtain the category label of the input statement according to the conversation vector; the conversation reply module 1214 generates feedback information corresponding to the input statement according to the category label and the action features, achieving the purpose of directly outputting feedback information through the conversation model, thereby realizing the technical effect of avoiding the logical code writing steps in the traditional method, and further solving the technical problem of large code demand in the process of dialogue decision-making and implementation by the prior art.
[0304] It should be noted here that the above modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.
[0305] Embodiment 15
[0306] An embodiment of the present invention may provide a computer terminal, and the computer terminal may be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the above computer terminal may also be replaced with a terminal device such as a mobile terminal.
[0307] Optionally, in this embodiment, the above computer terminal may be located in at least one of multiple network devices in a computer network.
[0308] In this embodiment, the above computer terminal may execute the program code of the following steps in the vulnerability detection method of the application program: obtaining the language feature of the sample data, and obtaining the corresponding dialogue vector according to the language feature; obtaining the category label and action feature of the sample data according to the language feature and the dialogue vector; generating the feedback information corresponding to the sample data according to the category label and the action feature; training the process of obtaining the language feature, dialogue vector, category label and action feature of the sample data according to the historical dialogue containing the feedback information, to obtain a dialogue model.
[0309] Optionally, Figure 13 is a structural block diagram of a computer terminal according to an embodiment of the present invention. As Figure 13 shown, the computer terminal 130 may include: one or more (only one is shown in the figure) processors 132, a memory 134, and a peripheral interface 136.
[0310] Among them, the memory may be used to store software programs and modules, such as the program instructions / modules corresponding to the security vulnerability detection method and device in the embodiment of the present invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the above interaction method. The memory may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, a flash memory, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely disposed relative to the processor, and these remote memories may be connected to the terminal 130 through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0311] The processor may call the information and application program stored in the memory through a transmission device to execute the following steps: obtaining the language feature of the sample data, and obtaining the corresponding dialogue vector according to the language feature; obtaining the category label and action feature of the sample data according to the language feature and the dialogue vector; generating the feedback information corresponding to the sample data according to the category label and the action feature; training the process of obtaining the language feature, dialogue vector, category label and action feature of the sample data according to the historical dialogue containing the feedback information, to obtain a dialogue model, where the dialogue model is used to obtain the feedback information corresponding to the input statement, and the feedback information includes: the feedback information on the input statement, and / or, the third-party service requested to be called in the input statement.
[0312] Optionally, the above processor may further execute the program code of the following steps: obtaining the language feature of the sample data includes: respectively performing entity recognition and semantic recognition on the sample data to obtain the language feature.
[0313] Optionally, the above-mentioned processor may also execute the program code of the following steps: respectively perform entity recognition and semantic recognition on the sample data to obtain language features, including: performing entity recognition on the sample data to obtain semantic entities; matching the semantic entities with preset data to obtain a first recognition result; performing semantic recognition on the sample data to obtain a second recognition result; obtaining language features based on the first recognition result, the second recognition result, and the state of the input sample data, where the language features include: state features, entity features, semantic features, and word features.
[0314] Optionally, the above-mentioned processor may also execute the program code of the following steps: obtaining a corresponding dialogue vector based on the language features, including: associating the language features with at least one historical dialogue data to obtain a dialogue vector.
[0315] Optionally, the above-mentioned processor may also execute the program code of the following steps: associating the language features with at least one historical dialogue data to obtain a dialogue vector, including: associating the state features, entity features, semantic features, and word features in the language features with at least one historical dialogue data to obtain a dialogue vector including the language features and at least one historical dialogue data.
[0316] Optionally, the above-mentioned processor may also execute the program code of the following steps: obtaining the category label and action features of the sample data based on the language features and the dialogue vector, including: obtaining the category label of the sample data based on the dialogue vector; obtaining the action features of the sample data based on the language features and the dialogue vector.
[0317] Optionally, the above-mentioned processor may also execute the program code of the following steps: obtaining the category label of the sample data based on the dialogue vector, including: classifying the semantic labels in the dialogue vector to obtain the category label of the sample data.
[0318] Optionally, the above-mentioned processor may also execute the program code of the following steps: the interactive method is applied to an end-to-end interactive system, where the end-to-end interactive system includes: an online interactive system, an intelligent interactive device, and an embedded intelligent interactive device.
[0319] The processor may call the information and application programs stored in the memory through a transmission device to execute the following steps: receiving an input statement; obtaining feedback information corresponding to the input statement through a dialogue model; where the dialogue model obtains the language features of the input statement and obtains a corresponding dialogue vector based on the language features; obtaining the category label and action features of the input statement based on the language features and the dialogue vector; generating feedback information corresponding to the input statement based on the category label and the action features; and training the process of obtaining the language features, dialogue vector, category label, and action features of the input statement based on the historical dialogue pair containing the feedback information.
[0320] An embodiment of the present invention provides a solution for an interaction method. By obtaining the language features of sample data and obtaining corresponding dialogue vectors based on the language features; obtaining the category labels and action features of the sample data based on the language features and dialogue vectors; generating feedback information corresponding to the sample data based on the category labels and action features; and training the process of obtaining the language features, dialogue vectors, category labels and action features of the sample data based on the historical dialogue containing the feedback information to obtain a dialogue model, the model is trained through the language features, dialogue vectors, category labels and action features of the sample data, as well as the corresponding feedback information, achieving the purpose of directly outputting feedback information through the dialogue model, thus realizing the technical effect of avoiding the logical code writing steps in the traditional method, and further solving the technical problem of the large demand for code in the process of dialogue decision-making and implementation in the prior art.
[0321] Those of ordinary skill in the art can understand that Figure 13 The structure shown is only for illustration. The computer terminal can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and terminal devices such as Mobile Internet Devices (MID), PAD, etc. Figure 13 It does not limit the structure of the above electronic device. For example, the computer terminal 130 may further include more or fewer components (such as a network interface, a display device, etc.) than those shown in Figure 13 or have a different configuration from that shown in Figure 13 shown.
[0322] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program, and the program can be stored in a computer-readable non-volatile storage medium. The non-volatile storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disc, etc.
[0323] Embodiment 16
[0324] An embodiment of the present invention further provides a non-volatile storage medium. Optionally, in this embodiment, the above non-volatile storage medium can be used to store the program code executed by the interaction method provided in the above Embodiment 1.
[0325] Optionally, in this embodiment, the above non-volatile storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.
[0326] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: obtaining the language features of the sample data, and obtaining the corresponding dialogue vector according to the language features; obtaining the category label and action features of the sample data according to the language features and the dialogue vector; generating feedback information corresponding to the sample data according to the category label and the action features; training the process of obtaining the language features, dialogue vector, category label, and action features of the sample data according to the historical dialogue pair containing the feedback information to obtain a dialogue model, where the dialogue model is used to obtain the feedback information corresponding to the input statement, and the feedback information includes: the feedback information for the input statement, and / or, the third-party service requested to be called in the input statement.
[0327] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: obtaining the language features of the sample data includes: respectively performing entity recognition and semantic recognition on the sample data to obtain the language features.
[0328] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: respectively performing entity recognition and semantic recognition on the sample data to obtain the language features includes: performing entity recognition on the sample data to obtain semantic entities; matching the semantic entities with preset data to obtain a first recognition result; performing semantic recognition on the sample data to obtain a second recognition result; obtaining the language features according to the first recognition result, the second recognition result, and the state of the input sample data, where the language features include: state features, entity features, semantic features, and word features.
[0329] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: obtaining the corresponding dialogue vector according to the language features includes: associating the language features with at least one historical dialogue data to obtain the dialogue vector.
[0330] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: associating the language features with at least one historical dialogue data to obtain the dialogue vector includes: associating the state features, entity features, semantic features, and word features in the language features with at least one historical dialogue data to obtain a dialogue vector including the language features and at least one historical dialogue data.
[0331] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: obtaining the category label and action features of the sample data according to the language features and the dialogue vector includes: obtaining the category label of the sample data according to the dialogue vector; obtaining the action features of the sample data according to the language features and the dialogue vector.
[0332] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: obtaining the class label of the sample data according to the dialogue vector, including: classifying the semantic labels in the dialogue vector to obtain the class label of the sample data.
[0333] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: the interaction method is applied to an end-to-end interaction system, where the end-to-end interaction system includes: an online interaction system, an intelligent interaction device, and a device embedded in the intelligent interaction device.
[0334] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: receiving an input statement; obtaining feedback information corresponding to the input statement through a dialogue model; where the dialogue model obtains the language features of the input statement and obtains the corresponding dialogue vector according to the language features; obtaining the class label and action features of the input statement according to the language features and the dialogue vector; generating feedback information corresponding to the input statement according to the class label and the action features; and training the process of obtaining the language features, dialogue vector, class label, and action features of the input statement according to the historical dialogue pair containing the feedback information.
[0335] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0336] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0337] In the several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0338] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0339] In addition, in each embodiment of the present invention, each functional unit may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0340] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable non-volatile storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a non-volatile storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned non-volatile storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs.
[0341] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An interactive method, characterized in that, it includes: According to the collected input statement, obtain the semantic entity in the input statement, where the semantic entity is used to represent the semantics in the input statement; According to the semantic entity, obtain the feedback information corresponding to the input statement through a dialogue model, where the feedback information includes at least one of the following: a reply to the input statement, a third-party service requested to be called in the input statement, The dialogue model is trained according to the language features, dialogue vectors, category labels, and action features of the sample data. The language features, dialogue vectors, category labels, and action features of the sample data are obtained from historical dialogue pairs containing the feedback information. The category label of the sample data is obtained according to the dialogue vector, and the action feature of the sample data is obtained according to the language feature and the dialogue vector. The category label includes dialogue action category labels, including: inquiry, dialogue, and query; the action feature includes: answering questions, responding to dialogue, and querying information; during the training process of the dialogue model, the dialogue model generates feedback information corresponding to the input statement according to the category label and the action feature, including: an answer statement to the user's question and answer, or a reply to the user's dialogue, or a response to the user's external service requirements; The method further includes: during the training process of the dialogue model, use the following data obtained from the dialogue tree to optimize the whole dialogue model to learn a dialogue interaction model from the dialogue tree: the interaction of each dialogue, including: the user's input question and the system's action, where the dialogue tree is a tree-structured dialogue, and each node of each layer of the dialogue tree is used to record the following information: what the user said and how the system answered, and the dialogue interaction model is used to directly obtain the action executed by the system to the user according to the user's input statement and the recognized entity.
2. The method according to claim 1, characterized in that, The obtaining of the semantic entity in the input statement includes: Segment the input statement to obtain at least one word segment that makes up the input statement; Obtain the category to which the at least one word segment belongs, and determine the semantic entity in the input statement according to the category.
3. The method according to claim 1 or 2, characterized in that, The obtaining of the feedback information corresponding to the input statement through the dialogue model according to the input statement and the semantic entity includes: Obtain the historical interaction result; According to the input statement, the semantic entity, and the historical interaction result, obtain the dialogue vector corresponding to the input statement; Obtain the feedback information corresponding to the dialogue vector through the dialogue model.
4. The method according to claim 3, characterized in that, The method further includes: After obtaining the feedback information, input the feedback information and the input statement into the dialogue model as updated data to obtain the optimized dialogue model.
5. The method according to claim 4, characterized in that, The dialogue model is obtained by acquiring the language features of the input statement, and obtaining the corresponding dialogue vector according to the language features; obtaining the category label and action features of the input statement according to the language features and the dialogue vector; generating the feedback information corresponding to the input statement according to the category label and the action features; and training according to the process of obtaining the language features, dialogue vector, category label and action features of the input statement based on the historical dialogue pair containing the feedback information.
6. An interaction method, characterized in that, it includes: collecting input statements input by multiple objects, and obtaining semantic entities in the input statements, where the semantic entities are used to represent the semantics in the input statements; determining the language features of different objects based on the semantic entities of different objects; invoking different dialogue models based on the language features of different objects; processing the input statements input by different objects with corresponding dialogue models to generate the feedback information corresponding to different objects, where the feedback information includes at least one of the following: feedback information on the input statement, third-party services requested to be invoked in the input statement; wherein, the dialogue model is obtained by training according to the process of the language features, dialogue vector, category label and action features of sample data, the language features, dialogue vector, category label and action features of the sample data are obtained from the historical dialogue pair containing the feedback information, the category label of the sample data is obtained according to the dialogue vector, the action features of the sample data are obtained according to the language features and the dialogue vector, the category label includes dialogue action category labels, including: inquiry, dialogue and query; the action features include: answering questions, responding to dialogue and querying information; during the training process of the dialogue model, the dialogue model generates the feedback information corresponding to the input statement according to the category label and the action features, including: answer statements to user questions and answers, or responses to user dialogues, or responses to user requests for external services; The method further includes: during the training process of the dialogue model, optimizing the whole of the dialogue model with the following data obtained from the dialogue tree to learn the dialogue interaction model from the dialogue tree: the interaction of each dialogue, including: the input questions of the user and the actions of the system, where the dialogue tree is a tree-structured dialogue, and each node of each layer of the dialogue tree is used to record the following information: what the user said and how the system answered, and the dialogue interaction model is used to directly obtain the actions executed by the system to the user according to the input statement of the user and the recognized entity.
7. An interaction method, characterized in that, it includes: acquiring the language features of sample data, and obtaining the corresponding dialogue vector according to the language features; Obtain the class label and action feature of the sample data according to the language feature and the dialogue vector, including: obtain the class label of the sample data according to the dialogue vector; obtain the action feature of the sample data according to the language feature and the dialogue vector, where the class label includes dialogue action class labels, including: inquiry, dialogue, and query; the action feature includes: answering questions, responding to dialogue, and querying information; during the training process of the dialogue model, the dialogue model generates feedback information corresponding to the input statement according to the class label and the action feature, including: an answer statement to the user's question and answer, or a reply to the user's dialogue, or a response to the user's request for an external service; Generate feedback information corresponding to the sample data according to the class label and the action feature; Train the process of obtaining the language feature, dialogue vector, class label, and action feature of the sample data according to the historical dialogue pair containing the feedback information to obtain a dialogue model, where the dialogue model is used to obtain feedback information corresponding to the input statement, and the feedback information includes at least one of the following: feedback information for the input statement, a third-party service requested to be called in the input statement; The method further includes: during the training process of the dialogue model, optimize the whole of the dialogue model using the following data obtained from the dialogue tree to learn a dialogue interaction model from the dialogue tree: the interaction of each dialogue, including: the user's input question and the system's action, where the dialogue tree is a tree-structured dialogue, and each node of each layer of the dialogue tree is used to record the following information: what the user said and how the system answered, and the dialogue interaction model is used to directly obtain the action executed by the system to the user according to the user's input statement and the recognized entity.
8. The method according to claim 7, wherein, Obtaining the language feature of the sample data includes: Perform entity recognition and semantic recognition on the sample data respectively to obtain the language feature.
9. The method according to claim 8, wherein, Performing entity recognition and semantic recognition on the sample data respectively to obtain the language feature includes: Perform entity recognition on the sample data to obtain semantic entities; Match the semantic entity with preset data to obtain a first recognition result; Perform semantic recognition on the sample data to obtain a second recognition result; Obtain the language feature according to the first recognition result, the second recognition result, and the state of the input sample data, where the language feature includes: state feature, entity feature, semantic feature, and word feature.
10. The method according to claim 7 or 8, wherein, Obtaining the corresponding dialogue vector according to the language feature includes: Associate the language feature with at least one historical dialogue data to obtain a dialogue vector.
11. The method according to claim 10, wherein, Associating the language feature with at least one historical dialogue data to obtain a dialogue vector includes: Associate with the at least one piece of historical conversation data according to the status feature, entity feature, semantic feature, and word feature in the language feature to obtain a conversation vector including the language feature and the at least one piece of historical conversation data.
12. The method according to claim 7, wherein, obtaining the class label of the sample data according to the conversation vector includes: classifying the semantic labels in the conversation vector to obtain the class label of the sample data.
13. The method according to claim 7, wherein, the interactive method is applied to an end-to-end interactive system, where the end-to-end interactive system includes: an online interactive system, an intelligent interactive device, and a device embedded in the intelligent interactive device.
14. An interactive method, wherein, includes: Step A, collect the currently input input statement; Step B, call the conversation model to analyze the currently input input statement, and display feedback information semantically related to the currently input input statement on the touch screen; Step C, collect the input statement input again based on the feedback information; Step D, call the conversation model to analyze the input statement input again, and display feedback information semantically related to the input statement input again on the touch screen; Step E, loop and execute the above Step C and Step D until an end instruction is received; wherein, the conversation model in Step B is trained according to the process of the language feature, conversation vector, class label, and action feature of the sample data, the language feature, conversation vector, class label, and action feature of the sample data are obtained from the historical conversation pairs containing the feedback information, the class label of the sample data is obtained according to the conversation vector, the action feature of the sample data is obtained according to the language feature and the conversation vector, the class label includes conversation action class labels, including: inquiry, conversation, and query; the action features include: answering questions, responding to conversations, and querying information; during the training process of the conversation model, the conversation model generates feedback information corresponding to the input statement according to the class label and the action feature, including: an answer statement to the user's question and answer, or a reply to the user's conversation, or a response to the user's need for external services; The method further includes: during the training process of the conversation model, optimizing the whole of the conversation model with the following data obtained from the conversation tree to learn the conversation interaction model from the conversation tree: the interaction of each conversation, including: the user's input question and the system's action, where the conversation tree is a tree-structured conversation, and each node of each layer of the conversation tree is used to record the following information: what the user said and how the system answered, and the conversation interaction model is used to directly obtain the action executed by the system to the user according to the user's input statement and the recognized entity.
15. An interactive method, wherein, includes: The server receives the input statement collected by the intelligent device and obtains the semantic entity in the input statement, where the semantic entity is used to represent the semantics in the input statement. Based on the semantic entity, the server obtains feedback information matching the input statement through a dialogue model. The feedback information includes at least one of the following: a reply to the input statement, and a third-party service requested to be invoked in the input statement. The dialogue model is trained according to the language features, dialogue vectors, category labels, and action features of sample data. The language features, dialogue vectors, category labels, and action features of the sample data are obtained from historical dialogue pairs containing the feedback information. The category labels of the sample data are obtained based on the dialogue vectors, and the action features of the sample data are obtained based on the language features and the dialogue vectors. The category labels include dialogue action category labels, including: inquiry, dialogue, and query; the action features include: answering questions, responding to dialogue, and querying information. During the training process of the dialogue model, the dialogue model generates feedback information corresponding to the input statement based on the category labels and the action features, including: an answer statement to the user's question and answer, or a reply to the user's dialogue, or a response to the user's external service requirements. During the training process of the dialogue model, the following data obtained from the dialogue tree is also used to optimize the overall dialogue model to learn the dialogue interaction model from the dialogue tree: the interaction of each dialogue, including: the user's input question and the system's action. The dialogue tree is a tree-structured dialogue, and each node in each layer of the dialogue tree is used to record the following information: what the user said and how the system answered. The dialogue interaction model is used to directly obtain the action executed by the system to the user according to the user's input statement and the recognized entity. The server sends the feedback information to the intelligent device.
16. An intelligent device Characterized in that It includes: A voice collection device for collecting voice information. A processor, connected to the voice acquisition device, is configured to identify an input statement from the voice information, process the semantic entity of the input statement through a dialogue model, and generate a corresponding feedback message. Wherein, the semantic entity of the input statement represents the semantics in the input statement. The dialogue model is trained according to the language features, dialogue vectors, category labels, and action features of sample data. The language features, dialogue vectors, category labels, and action features of the sample data are obtained from historical dialogue pairs containing the feedback message. The category label of the sample data is obtained according to the dialogue vector. The action feature of the sample data is obtained according to the language feature and the dialogue vector. The category label includes dialogue action category labels, including: inquiry, conversation, and query; the action features include: answering questions, responding to conversations, and querying information. During the training process of the dialogue model, the dialogue model generates a feedback message corresponding to the input statement according to the category label and the action feature, including: An answer statement to the user's question and answer, or a reply to the user's conversation, or a response to the user's request for an external service; During the training process of the dialogue model, the following data obtained from the dialogue tree is also used to optimize the overall dialogue model to learn a dialogue interaction model from the dialogue tree: the interaction of each dialogue, including: the user's input question and the system's action. The dialogue tree is a tree-structured dialogue. Each node of each layer of the dialogue tree is used to record the following information: what the user said and how the system answered. The dialogue interaction model is used to directly obtain the action executed by the system to the user according to the user's input statement and the recognized entity.
17. An intelligent device Characterized in that Comprising A touch screen for collecting an input statement input through a touch interaction function, wherein the semantic entity of the input statement represents the semantics of the input statement; A processor, connected to the touch screen, is configured to process the semantic entities of the input statement through a dialogue model and generate corresponding feedback information. The dialogue model is trained according to the language features, dialogue vectors, category labels, and action features of sample data. The language features, dialogue vectors, category labels, and action features of the sample data are obtained from historical dialogue pairs containing the feedback information. The category labels of the sample data are obtained according to the dialogue vectors, and the action features of the sample data are obtained according to the language features and the dialogue vectors. The category labels include dialogue action category labels, including: inquiry, dialogue, and query; the action features include: answering questions, responding to dialogue, and querying information. During the training process of the dialogue model, the dialogue model generates feedback information corresponding to the input statement according to the category labels and the action features, including: an answer statement to the user's question and answer, or a reply to the user's dialogue, or a response to the user's demand for external services. During the training process of the dialogue model, the following data obtained from the dialogue tree is also used to optimize the whole dialogue model to learn the dialogue interaction model from the dialogue tree: the interaction of each dialogue, including: the user's input question and the system's action. The dialogue tree is a tree-structured dialogue. Each node of each layer of the dialogue tree is used to record the following information: what the user said and how the system answered. The dialogue interaction model is used to directly obtain the action executed by the system to the user according to the user's input statement and the recognized entity. Wherein, the touch screen is further configured to display the feedback information.
18. An intelligent device Characterized in that Comprising: A voice collection device, configured to collect voice information; A touch screen, configured to display at least one statement generated by recognizing the voice information and receive an input statement selected from the at least one statement through a touch interaction function, wherein the semantic entity of the input statement represents the semantics of the input statement. A processor, connected to the voice acquisition device and the touch screen, is configured to process the semantic entity of the input statement through a dialogue model and generate corresponding feedback information. The dialogue model is trained according to the language features, dialogue vectors, category labels, and action features of sample data. The language features, dialogue vectors, category labels, and action features of the sample data are obtained from historical dialogue pairs containing the feedback information. The category labels of the sample data are obtained based on the dialogue vectors, and the action features of the sample data are obtained based on the language features and the dialogue vectors. The category labels include dialogue action category labels, including: inquiry, dialogue, and query; the action features include: answering questions, responding to dialogues, and querying information. During the training process of the dialogue model, the dialogue model generates feedback information corresponding to the input statement according to the category labels and the action features, including: an answer statement to the user's question and answer, or a reply to the user's dialogue, or a response to the user's request for an external service. During the training process of the dialogue model, the following data obtained from the dialogue tree is also used to optimize the overall dialogue model to learn the dialogue interaction model from the dialogue tree: the interaction of each dialogue, including: the user's input question and the system's action. The dialogue tree is a tree-structured dialogue. Each node of each layer of the dialogue tree is used to record the following information: what the user said and how the system answered. The dialogue interaction model is used to directly obtain the action executed by the system to the user according to the user's input statement and the recognized entity. The touch screen is further configured to display the feedback information.
19. An interaction device Characterized in that It includes: An acquisition module, configured to obtain the semantic entity in the input statement according to the collected input statement, where the semantic entity is used to identify the semantics in the input statement. An interaction module, configured to obtain feedback information corresponding to the input statement through a dialogue model according to the input statement and the semantic entity, where the feedback information includes at least one of the following: a reply to the input statement, a third-party service requested to be invoked in the input statement. The dialogue model is trained according to the language features, dialogue vectors, category labels, and action features of sample data. The language features, dialogue vectors, category labels, and action features of the sample data are obtained from historical dialogue pairs containing the feedback information. The category labels of the sample data are obtained according to the dialogue vectors, and the action features of the sample data are obtained according to the language features and the dialogue vectors. The category labels include dialogue action category labels, including: inquiry, dialogue, and query; the action features include: answering questions, responding to dialogue, and querying information. During the training process of the dialogue model, the dialogue model generates feedback information corresponding to the input statement according to the category labels and the action features, including: an answer statement to the user's question and answer, or a reply to the user's dialogue, or a response to the user's external service requirement. During the training process of the dialogue model, the following data obtained from the dialogue tree is also used to optimize the whole dialogue model to learn a dialogue interaction model from the dialogue tree: the interaction of each dialogue, including: the user's input question and the system's action. The dialogue tree is a tree-structured dialogue. Each node of each layer of the dialogue tree is used to record the following information: what the user said and how the system answered. The dialogue interaction model is used to directly obtain the action executed by the system to the user according to the user's input statement and the recognized entity.
20. An interaction device Characterized in that It includes: A collection module, configured to collect input statements input by multiple objects and obtain semantic entities in the input statements, where the semantic entities are used to represent the semantics in the input statements; A feature module, configured to determine the language features of different objects based on the semantic entities of different objects; An invocation module, configured to invoke different dialogue models based on the language features of different objects; An interaction module, configured to process the input statements input by different objects using corresponding dialogue models to generate feedback information corresponding to different objects, where the feedback information includes at least one of the following: feedback information on the input statement, a third-party service requested to be invoked in the input statement; Among them, the dialogue model is trained according to the process of the language features, dialogue vectors, category labels, and action features of the sample data. The language features, dialogue vectors, category labels, and action features of the sample data are obtained from the historical dialogue pairs containing the feedback information. The category labels of the sample data are obtained according to the dialogue vectors, and the action features of the sample data are obtained according to the language features and the dialogue vectors. The category labels include dialogue action category labels, including: inquiry, dialogue, and query; the action features include: answering questions, responding to dialogues, and querying information. During the training process of the dialogue model, the dialogue model generates feedback information corresponding to the input statement according to the category labels and the action features, including: an answer statement to the user's question and answer, or a reply to the user's dialogue, or a response to the user's need for an external service. During the training process of the dialogue model, the following data obtained from the dialogue tree is also used to optimize the whole dialogue model to learn the dialogue interaction model from the dialogue tree: the interaction of each dialogue, including: the user's input question and the system's action. Among them, the dialogue tree is a tree-structured dialogue, and each node of each layer of the dialogue tree is used to record the following information: what the user said and how the system answered. The dialogue interaction model is used to directly obtain the action executed by the system to the user according to the user's input statement and the recognized entity.
21. An interactive device, Characterized in that, Comprising: A first acquisition module, configured to acquire the language features of the sample data and acquire the corresponding dialogue vectors according to the language features; A second acquisition module, configured to acquire the category labels and action features of the sample data according to the language features and the dialogue vectors, including: acquiring the category labels of the sample data according to the dialogue vectors; acquiring the action features of the sample data according to the language features and the dialogue vectors. The category labels include dialogue action category labels, including: inquiry, dialogue, and query; the action features include: answering questions, responding to dialogues, and querying information; A statement generation module, configured to generate feedback information corresponding to the sample data according to the category labels and the action features; The model training module is used to train the process of obtaining the language features, dialogue vectors, category labels, and action features of the sample data based on the historical conversations containing the feedback information, so as to obtain a dialogue model. Among them, the dialogue model is used to obtain the feedback information corresponding to the input statement. The feedback information includes at least one of the following: the feedback information on the input statement, the third-party service requested to be called in the input statement. During the training process of the dialogue model, the dialogue model generates the feedback information corresponding to the input statement according to the category label and the action feature, including: the answer statement to the user's question and answer, or the reply to the user's conversation, or the response to the user's demand for an external service. During the training process of the dialogue model, the following data obtained from the dialogue tree is also used to optimize the whole dialogue model to learn the dialogue interaction model from the dialogue tree: the interaction of each dialogue, including: the user's input question and the system's action. The dialogue tree is a tree-structured dialogue. Each node of each layer of the dialogue tree is used to record the following information: what the user said and how the system answered. The dialogue interaction model is used to directly obtain the action executed by the system to the user according to the user's input statement and the recognized entity.
22. An interactive system Characterized in that It includes: An entity recognition module, a semantic modeling module, a feature acquisition module, a gated recurrent module, an action decision module, a classification module, and a dialogue reply module, where The entity recognition module is used to perform entity recognition on the input statement to obtain semantic entities, and match the semantic entities with preset data to obtain a first recognition result; The semantic modeling module is used to perform semantic recognition on the input statement to obtain a second recognition result; The feature acquisition module is respectively connected to the entity recognition module and the semantic modeling module, and is used to obtain language features according to the first recognition result, the second recognition result, and the state of the input statement. The language features include: state features, entity features, semantic features, and word features; The gated recurrent module is connected to the feature acquisition module, and is used to associate the state features, entity features, semantic features, and word features in the language features with at least one historical conversation data to obtain a dialogue vector containing the language features and the at least one historical conversation data; The action decision module is connected to the feature acquisition module and the gated recurrent module, and is used to obtain the action feature of the input statement according to the language features and the dialogue vector; The classification module is connected to the gated recurrent module, and is used to obtain the category label of the input statement according to the dialogue vector; The dialogue reply module generates feedback information corresponding to the input statement according to the category label and the action feature. Among them, the category label includes dialogue action category labels, including: inquiry, conversation, and query; the action feature includes: answering questions, responding to conversations, and querying information; during the training process of the dialogue model, the dialogue model generates feedback information corresponding to the input statement according to the category label and the action feature, including: an answer statement to the user's question and answer, or a reply to the user's conversation, or a response to the user's need for an external service; during the training process of the dialogue model, the following data obtained from the dialogue tree is also used to optimize the overall dialogue model to learn the dialogue interaction model from the dialogue tree: the interaction of each conversation, including: the user's input question and the system's action. Among them, the dialogue tree is a tree-structured dialogue, and each node of each layer of the dialogue tree is used to record the following information: what the user said and how the system answered. The dialogue interaction model is used to directly obtain the action executed by the system to the user according to the user's input statement and the identified entity.
23. A non-volatile storage medium, wherein, the non-volatile storage medium includes a stored program, wherein, when the program runs, it controls the device where the non-volatile storage medium is located to execute the interaction method according to any one of claims 1 to 15.
24. A processor, wherein, the processor is used to run a program, wherein, when the program runs, it executes the interaction method according to any one of claims 1 to 15.
Citation Information
Patent Citations
Deep learning-based dialogue method, device and equipment
CN107885756A