Method, apparatus, system, electronic device and storage medium for voice interaction
By storing historical information locally on the client, the problem of cloud storage dependency is solved, faster response time and higher voice reply accuracy are achieved, while protecting user privacy.
Patent Information
- Application Number
- CN202111339114.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-12
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-11-12
AI Technical Summary
In the prior art, the human-computer voice interaction system relies on the cloud to store the above information, resulting in additional economic costs, increased human-computer interaction response time and privacy risks, and the unreliability of the network connection affects the accuracy of the reply.
Store historical information locally on the client, and voice recognition and reply generation are performed through the client's interaction with the server, avoid cloud storage and network connections, and ensure local updates of the above information.
It reduces cloud storage and maintenance costs, reduces human-computer interaction response time, improves the correctness of voice reply results, and protects user privacy.
Smart Images

Figure CN114187903B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and particularly to artificial intelligence technology fields such as speech technology and natural language processing, and more particularly to methods, devices, systems, electronic devices, and storage media for speech interaction. Background Art
[0002] With the progress of artificial intelligence technology, human-computer speech interaction (hereinafter referred to as speech interaction) has also been rapidly developed and widely applied. For example, it can be widely applied to intelligent devices such as smart TVs, smart speakers, virtual reality (VR) glasses, and various speech assistant applications (APPs).
[0003] The existing human-computer interaction process mainly includes four parts: speech recognition, semantic parsing, dialogue service, and speech synthesis. Among them, both semantic parsing and dialogue service rely on the above context information (session information). For the same user request, if the above context information is different, the speech interaction system will give different responses. For example, for the current round of user request "eight o'clock tomorrow", if the previous round of user request is "set an alarm for me", and the device in the previous round of speech interaction asks back "Okay, what time do you want to set the alarm", then for the current round of user request "eight o'clock tomorrow", the device will reply "Okay, the alarm for eight o'clock tomorrow has been set for you"; if the previous round of user request is "remind me to have a meeting", and the device in the previous round asks back "Okay, what time do you want me to remind you", then for the current round of user request "eight o'clock tomorrow", the device will reply "Okay, I will remind you to have a meeting at eight o'clock tomorrow".
[0004] In the prior art, storing the above context information in a dedicated cloud-based above context information storage device has at least the following problems: it is necessary to maintain additional above context information storage device resources in the cloud, which incurs certain economic costs; the connection between the central control module in the speech interaction system and the above context information storage device depends on network connection, resulting in an increase in the human-computer interaction response time perceived by the end user; in addition, if the cloud central control fails to obtain the above context information, it will lead to incorrect responses to users, which does not meet user expectations; the above context information of users usually includes a large amount of user behaviors, and storing them in the cloud poses certain privacy risks. Summary of the Invention
[0005] The present disclosure provides a method, device, system, electronic device, and storage medium for speech interaction.
[0006] According to one aspect of the present disclosure, there is provided a method for speech interaction, which is applied to a client, and the method includes:
[0007] In response to receiving a request statement sent by a user, obtain historical information from a local storage module, and send the request statement and the historical information to a server;
[0008] Receive the reply statement returned by the server and the current context information corresponding to the request statement;
[0009] Play the reply statement and store the current context information as historical information in the storage module.
[0010] According to another aspect of the present disclosure, another method for voice interaction is provided, which is applied to a server. The method includes:
[0011] Receive a request statement and historical information sent by a client;
[0012] Perform speech recognition on the request statement to obtain a speech recognition result;
[0013] Determine a reply statement based on the speech recognition result and the historical information, and generate the current context information corresponding to the request statement based on the speech recognition result and the reply statement;
[0014] Send the reply statement and the current context information to the client so that the client can play the reply statement and store the current context information as historical information in the storage module local to the client.
[0015] According to still another aspect of the present disclosure, a client is provided, including:
[0016] A storage module for storing historical information;
[0017] A first receiving unit for obtaining historical information from the storage module in response to receiving a request statement sent by a user;
[0018] A transceiver unit for sending the request statement and the historical information to a server; and receiving the reply statement returned by the server and the current context information corresponding to the request statement;
[0019] A playback unit for playing the reply statement;
[0020] A storage processing unit for storing the current context information as historical information in the storage module.
[0021] According to still another aspect of the present disclosure, a voice interaction device is provided, which is applied to a server. The device includes:
[0022] A second receiving unit for receiving a request statement and historical information sent by a client;
[0023] A speech recognition unit for performing speech recognition on the request statement to obtain a speech recognition result;
[0024] A determination unit, configured to determine a reply statement based on the speech recognition result and the historical information;
[0025] A generation unit, configured to generate current context information corresponding to the request statement based on the speech recognition result and the reply statement;
[0026] A sending unit, configured to send the reply statement and the current context information to the client, so that the client plays the reply statement and stores the current context information as historical information in a storage module local to the client.
[0027] According to another aspect of the present disclosure, there is provided a voice interaction system, including any possible client as described above and any possible voice interaction device as described above.
[0028] According to another aspect of the present disclosure, there is provided an electronic device, including:
[0029] At least one processor; and
[0030] A memory communicatively connected to the at least one processor; wherein,
[0031] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the methods of the aspects and any possible implementation manners as described above.
[0032] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause a computer to execute the methods of the aspects and any possible implementation manners as described above.
[0033] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, and the computer program implements the methods of the aspects and any possible implementation manners as described above when executed by a processor.
[0034] According to another aspect of the present disclosure, there is provided an artificial intelligence device, including the electronic device as described above.
[0035] As can be seen from the above technical solution, when the client receives an interaction request initiated by the user, the received request statement and the historical information stored this time are sent to the server. The server performs speech recognition on the request statement to obtain a speech recognition result. Then, a reply statement is determined based on the speech recognition result and the historical information, and the current context information corresponding to the request statement is generated based on the speech recognition result and the reply statement. Furthermore, the reply statement and the current context information are sent to the client, so that the client can play the reply statement and update and store the current context information as historical information in the storage module local to the client.
[0036] In this way, by locally storing historical information on the client, the cloud server no longer needs to maintain a dedicated context information storage, which can reduce the storage and maintenance costs of the cloud. There is no longer a need for the central control module in the voice interaction system to send and receive data by establishing a network connection with the context information storage and perform read / write operations on the context information storage, which can save time, reduce the human-computer interaction response time perceived by the user, and improve the user experience.
[0037] In addition, by binding the user's upstream request statement and historical information and sending them to the server, and binding the reply statement and the current context information and returning them to the client, the consistency of the speech recognition result and the response of the voice interaction system can be ensured. That is, as long as the user's request statement is successfully recognized and the reply statement is successfully broadcast, the content of the context information is successfully accessed and stored, which improves the correctness of the voice reply result.
[0038] In addition, by locally storing the user's historical information on the client, user privacy can be better protected.
[0039] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0041] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure;
[0042] Figure 2 is a schematic diagram according to the second embodiment of the present disclosure;
[0043] Figure 3 is a schematic diagram according to the third embodiment of the present disclosure;
[0044] Figure 4 is a schematic diagram according to the fourth embodiment of the present disclosure;
[0045] Figure 5 It is a schematic diagram of a voice interaction process according to an embodiment of the present disclosure;
[0046] Figure 6 It is a schematic diagram according to the fifth embodiment of the present disclosure;
[0047] Figure 7 It is a schematic diagram according to the sixth embodiment of the present disclosure;
[0048] Figure 8 It is a schematic diagram according to the seventh embodiment of the present disclosure;
[0049] Figure 9 It is a block diagram of an electronic device for implementing the method of voice interaction according to the embodiments of the present disclosure. Detailed implementation manners
[0050] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0051] Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts fall within the scope of protection of the present disclosure.
[0052] It should be noted that the terminal devices involved in the embodiments of the present disclosure may include, but are not limited to, intelligent devices such as mobile phones, personal digital assistants (PDAs), wireless handheld devices, tablet computers, and computing devices in vehicles; display devices may include, but are not limited to, devices with display functions such as personal computers, televisions, and displays coupled in vehicles.
[0053] In addition, the term "and / or" herein is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.
[0054] The way of human-computer interaction has been gradually improving in terms of efficiency and convenience, from the keyboard and mouse interaction in the initial PC era, to the touch screen interaction in the mobile era, and then to the voice interaction in the artificial intelligence era. The effect of voice interaction needs to be further improved to bring a better human-computer interaction experience to users.
[0055] The inventors of the present disclosure found through research that in the prior art, storing the above information in a dedicated cloud-based above information memory has at least the following problems: It is necessary to maintain additional above information memory resources in the cloud, which incurs certain economic costs; the connection between the central control module in the voice interaction system and the above information memory depends on network connection. Since establishing a connection between the two, sending and receiving data, and the above information memory maintaining the reading and writing of the above information of all users all consume a certain amount of time, it will lead to an increase in the human-computer interaction response time perceived by the end user; in addition, the network connection and the reading / writing of the above information memory cannot guarantee a 100% success rate. If the cloud central control fails to obtain the above information, it will result in incorrect replies to users, which does not meet user expectations; the above information of users usually includes a large amount of user behavior, and storing it in the cloud poses certain privacy risks.
[0056] Therefore, there is an urgent need to provide a voice interaction processing method to reduce the maintenance cost of the cloud, reduce the human-computer interaction response time, improve the correctness of voice reply results, and protect user privacy.
[0057] Figure 1 It is a schematic diagram according to the first embodiment of the present disclosure, and this embodiment can be applied to a client, such as Figure 1 as shown.
[0058] 101. In response to receiving a request statement sent by a user, obtain historical information (session) from the local storage module, and send the request statement and the historical information to the server.
[0059] Among them, the request statement is a statement sent by the user that wants to get a reply. For example, it can be words with different requirements such as "Play children's songs" or "I want to listen to Dao Xiang". The embodiments of the present disclosure do not limit this.
[0060] Among them, the historical information is the above information that the voice interaction device needs to refer to when making a reply to the current user request statement. It can include information related to the request statement initiated by the user and the information related to the reply made by the voice interaction device in a complete human-computer interaction service, such as in a complete human-computer interaction service like user booking an alarm, booking a meeting reminder, querying the weather, etc. For example, the domain in the semantic parsing result corresponding to the request statement initiated by the user above, the broadcast script of the reply statement made by the device, and the slot information in the reply statement, etc. The embodiments of the present disclosure do not limit the specific content of the historical information.
[0061] 102. Receive the reply statement returned by the server and the current context information corresponding to the request statement.
[0062] Among them, the reply statement is that the server calls the dialogue service according to the speech recognition result of the request statement and the historical information. The dialogue service determines the user's intention according to the speech recognition result and the dialogue state, performs the calculation of the dialogue model and resource acquisition, and generates the reply statement to be broadcast to the user.
[0063] Among them, the current context information is the context information of the current request statement, which is different from the historical information before the current request statement.
[0064] 103. Play the reply statement and store the current context information as historical information in the local storage module.
[0065] It should be noted that part or all of the execution subjects of 101 to 103 can be applications located in the local terminal, that is, the terminal device of the service provider, or can also be functional units such as plugins or software development kits (SDKs) set in the applications located in the local terminal. This embodiment does not make special limitations on this.
[0066] It can be understood that the application can be a native program (nativeApp) installed on the terminal, or can also be a web program (webApp) of a browser on the terminal. This embodiment does not make special limitations on this.
[0067] In this way, by locally storing historical information on the client side, the cloud server no longer needs to maintain a dedicated context information memory, which can reduce the storage and maintenance costs of the cloud; it no longer needs the central control module in the voice interaction system to send and receive data with the context information memory through a network connection and perform read / write operations on the context information memory, which can save time and reduce the human-computer interaction response time perceived by users, improving the user experience. In addition, binding the user's upstream request statement and historical information and sending them to the server, and receiving the reply statement and the current context information returned by the server together can ensure the consistency of the speech recognition result and the reply of the voice interaction system, that is, as long as the user's request statement is successfully recognized and the reply statement is successfully broadcast, the content of the context information is successfully stored and retrieved, improving the correctness of the voice reply result. In addition, by locally storing the user's historical information on the client side, the user's privacy can be better protected.
[0068] Optionally, in a possible implementation of this embodiment, in 101, historical information of the most recent preset number of rounds may be obtained from the storage module based on a preset rule. For example, according to the preset rule, obtain the historical information of the most recent round from the storage module; or, according to the preset rule, obtain the historical information of the most recent five rounds from the storage module; or, according to the preset rule, obtain all the historical information from the storage module; and so on. Specifically, how many rounds of historical information need to be obtained can be preset according to the actual application requirements and can be updated as needed. The embodiments of the present disclosure do not limit this.
[0069] Based on this embodiment, the preset rule can be set according to specific application requirements, and the historical information of the most recent preset number of rounds can be obtained from the storage module, so that the historical information sent to the server can not only meet the requirements for determining the reply statement, but also avoid occupying network resources and computing resources by sending and analyzing redundant information, improving the resource utilization rate of the server and the accuracy of the voice interaction result.
[0070] Optionally, in a possible implementation of this embodiment, in 103, the current upstream information may be stored as historical information in the local storage module based on the reception time sequence.
[0071] Based on this embodiment, the current upstream information can be sequentially stored as historical information in the local storage module based on the reception time sequence, which helps the subsequent client to obtain the historical information of the most recent preset number of rounds according to the information storage sequence, improving the accuracy and efficiency of obtaining historical information.
[0072] Optionally, in a possible implementation of this embodiment, the current upstream information may include, for example, but is not limited to: the domain in the semantic parsing result corresponding to the request statement, the broadcast script of the reply statement, and the slot information in the reply statement, and so on.
[0073] Based on this embodiment, the domain in the semantic parsing result corresponding to the request statement, the broadcast script of the reply statement, and the slot information in the reply statement are used as the current upstream information, so as to be sent to the server as the historical information of the next round of user request statement, so that the server can correctly determine the reply result to the user based on this.
[0074] Figure 2 It is a schematic diagram according to the second embodiment of the present disclosure. This embodiment can be applied to a server, such as Figure 2 shown.
[0075] 201. Receive a request statement and historical information sent by the client.
[0076] Among them, the request statement is the statement sent by the user that wants to get a reply. For example, it can be words with different requirements such as "Play children's songs" and "I want to listen to Dao Xiang". The embodiments of the present disclosure do not make any limitations on this.
[0077] 202. Perform speech recognition on the request statement to obtain a speech recognition result.
[0078] 203. Determine a reply statement based on the speech recognition result and the historical information, and generate the current context information corresponding to the request statement based on the speech recognition result and the reply statement.
[0079] Among them, the reply statement is that the server calls a dialogue service according to the speech recognition result and historical information of the request statement. The dialogue service determines the user's intention according to the speech recognition result and the dialogue state, and performs calculations and resource acquisition of the dialogue model to generate a reply statement to be broadcast to the user.
[0080] 204. Send the reply statement and the current context information to the client, so that the client plays the reply statement and stores the current context information as historical information in the storage module local to the client.
[0081] It should be noted that part or all of the execution entities of 201 to 204 can be an application located on the server, or can also be a plug-in or software development kit (SDK) and other functional units set in the application located on the server, or can also be a processing engine in the network-side server, or can also be a distributed system on the network side. The present embodiment does not make any special limitations on this.
[0082] It can be understood that the application can be a native app installed on the server, or can also be a web app of a browser on the server. The present embodiment does not make any special limitations on this.
[0083] In this way, by locally storing historical information on the client side, the cloud server no longer needs to maintain a dedicated previous context information storage, which can reduce the storage and maintenance costs of the cloud. It also no longer requires the central control module in the voice interaction system to send and receive data through a network connection with the previous context information storage and perform read / write operations on the previous context information storage, which can save time, reduce the human-computer interaction response time perceived by the user, and improve the user experience. Additionally, by receiving the user's upstream request statement and historical information sent together by the client and binding the reply statement to the current previous context information and returning it to the client, the consistency of the speech recognition result and the reply of the voice interaction system can be ensured. That is, as long as the user's request statement is successfully recognized and the reply statement is successfully broadcast, the content of the previous context information is successfully stored and retrieved, improving the correctness of the voice reply result. Moreover, by locally storing the user's historical information on the client side, user privacy can be better protected.
[0084] Optionally, in a possible implementation manner of this embodiment, the historical information may include, for example, but is not limited to: the client obtains the historical information of the most recent preset number of rounds from the storage module in the client based on a preset rule.
[0085] Based on this embodiment, the historical information sent by the client is obtained from the storage module according to a preset rule set according to specific application requirements, for the most recent preset number of rounds. This enables the historical information to not only meet the requirements for determining the reply statement but also avoid occupying network resources and computing resources by sending and analyzing redundant information, improving the resource utilization rate of the server and the accuracy of the voice interaction result.
[0086] Figure 3 is a schematic diagram according to the third embodiment of the present disclosure, as Figure 3 shown. On the basis of the embodiment shown in Figure 2 , operation 202 may include:
[0087] 301. Perform speech recognition on the request statement, and within a first preset time after receiving the request statement, obtain at least one intermediate recognition result.
[0088] Among them, the end moment of the first preset time is earlier than the tail point moment (VAD end) of the voice activity detection of the request statement. The first preset time refers to a period of time that starts counting after the tail sound of the request statement sent by the user has dropped, that is, the first preset time is the duration of the silent state detected by the voice interaction device. The length of the first preset time can be set as needed. In actual use, in order to reduce the waiting time of the user, the length of the first preset time can be set to 5 ms, 10 ms, 20 ms, etc. Here, this is only an example and cannot be used as a limitation on the length of the first preset time in the present disclosure.
[0089] Among them, the intermediate recognition result is the result obtained by performing speech recognition on the query statement within the first preset time after the end sound of the request statement sent by the user has dropped.
[0090] 302. In response to obtaining the first intermediate recognition result among the at least one intermediate recognition result, sequentially recognize whether the semantics of the at least one intermediate recognition result is complete according to the time sequence of obtaining the at least one intermediate recognition result, that is, whether the intermediate recognition result has complete semantics and whether the expressed meaning is complete.
[0091] Among them, the completeness of the semantics of the intermediate recognition result means that the intermediate recognition result has complete semantics and can completely express the user's meaning.
[0092] 303. In response to recognizing the first intermediate recognition result with complete semantics from the at least one intermediate recognition result, use this first intermediate recognition result with complete semantics as the speech recognition result.
[0093] Based on this embodiment, by performing streaming speech recognition in advance, when the first intermediate recognition result with complete semantics is recognized, use this first intermediate recognition result with complete semantics as the speech recognition result to determine the response statement. This can not only reduce the response time of the speech interaction by pulling the dialogue resources in advance, but also avoid a large number of calls to the dialogue service for the calculation and resource pulling of the dialogue model based on the intermediate recognition results with incomplete semantics, which can greatly reduce the ineffective calls to the dialogue service, reduce the request volume of the dialogue service, thereby saving the computing resources, storage resources, and charging resource services of the dialogue service, and thus reducing the cost.
[0094] In the embodiment of the present disclosure, in 302, various methods can be used to recognize whether the semantics of each intermediate recognition result is complete.
[0095] For example, in a possible implementation manner of this embodiment, the following method can be used to recognize whether the semantics of each intermediate recognition result is complete:
[0096] Sequentially for each intermediate recognition result among the at least one intermediate recognition result, use the semantic integrity model to obtain the first probability of each intermediate recognition result as the prefix in the historical final recognition result, and the second probability of each intermediate recognition result as the historical final recognition result. Then, based on the first probability and the second probability, determine the third probability of the semantics of each intermediate recognition result being complete; furthermore, according to whether the third probability is greater than the preset threshold, determine whether the semantics of each intermediate recognition result is complete.
[0097] Among them, the semantic integrity model is obtained through statistical calculation based on the final recognition results in the user logs. For example, in the large-scale online user logs, it is counted how many times the final recognition result "I want to listen to Dao Xiang" appears in total, how many times the final recognition result "I want to set the alarm for 6:50 tomorrow morning" appears in total, and so on. The total number of occurrences of each final recognition result in the large-scale online user logs is statistically calculated respectively to obtain the semantic integrity model. The semantic integrity model can be updated based on the update of the online user logs according to a certain update period, such as one week. The embodiments of the present disclosure do not limit whether to update the semantic integrity model and the update period.
[0098] In the embodiments of the present disclosure, the historical final recognition results are the final recognition results in the large-scale online user logs. Each time a user makes a request statement during a voice interaction online, at least one intermediate recognition result and one final recognition result will be obtained. The server for implementing the embodiments of the present disclosure will generate user logs for each user, recording at least one intermediate recognition result and one final recognition result in each voice interaction of the user.
[0099] In the embodiments of the present disclosure, the first probability of the intermediate recognition result as the prefix in the historical final recognition result, that is, the probability that the intermediate recognition result appears only as a part of the content in the front of the historical final recognition result, rather than the complete historical final recognition result. The second probability of the intermediate recognition result as the historical final recognition result, that is, the probability that the intermediate recognition result appears alone as the historical final recognition result (that is, as a complete historical final recognition result).
[0100] In the embodiments of the present disclosure, the higher the first probability and the lower the second probability, the lower the third probability that the intermediate recognition result is semantically complete. For example, for the intermediate result "I want to listen", in the user's expression, it rarely appears alone as a request statement and almost always appears as the prefix of the request statement. Then the third probability of the intermediate result "I want to listen" is relatively low, so as to determine that the intermediate result "I want to listen" is semantically incomplete. In the embodiments of the present disclosure, when the third probability is greater than a preset threshold, it can be determined that the intermediate recognition result is semantically complete, the meaning expressed is relatively complete, and it is a complete sentence. Thus, the intermediate recognition result can be used as a complete sentence, and then the first reply statement can be determined according to the intermediate recognition result; otherwise, when the third probability is less than or equal to the preset threshold, it is determined that the semantic of the intermediate recognition result is incomplete, the meaning expressed is not complete enough, and it may not be a complete sentence, and the intermediate recognition result is unreliable. The preset threshold can be set according to actual needs. For example, it can be set to 0.5 and can be adjusted according to actual needs.
[0101] Based on this embodiment, a semantic integrity model can be obtained in advance by statistically calculating the final recognition results in a large-scale online user log. Then, using the semantic integrity model, the first probability of each intermediate recognition result as a prefix in the historical final recognition results and the second probability of each intermediate recognition result as the historical final recognition result can be quickly obtained. The third probability of the semantic integrity of each intermediate recognition result can be accurately determined from the first probability and the second probability, so as to quickly and objectively determine whether the semantics of each intermediate recognition result is complete by using the method of historical statistical calculation.
[0102] Alternatively, in another possible implementation manner of this embodiment, the following method can also be used to identify whether the semantics of each intermediate recognition result is complete:
[0103] For each intermediate recognition result in the at least one intermediate recognition result in turn, obtain the word vector of each intermediate recognition result. For example, through the Word to the vector method, the intermediate recognition result can be converted from text to a word vector;
[0104] Obtain the first probability of each intermediate recognition result as a prefix in the historical final recognition results and the second probability of each intermediate recognition result as the historical final recognition result. For example, using the above semantic integrity model, obtain the first probability of each intermediate recognition result as a prefix in the historical final recognition results and the second probability of each intermediate recognition result as the historical final recognition result;
[0105] Obtain the popularity of each intermediate recognition result. The popularity of the intermediate recognition result, that is, within a certain time period or all historical times in the past, the usage degree of the user sending a request statement containing the intermediate recognition result through the server for implementing the embodiments of the present disclosure, can be statistically calculated from the intermediate recognition results and the final recognition results in the large-scale online user log;
[0106] Input the word vector, the first probability, the second probability, and the popularity into a pre-trained neural network model, and the fourth probability of the semantic integrity of each intermediate recognition result is output by the neural network model.
[0107] Determine whether the semantics of each intermediate recognition result is complete according to whether the fourth probability is greater than a preset threshold.
[0108] In the embodiments of the present disclosure, when the fourth probability is greater than a preset threshold, it can be determined that the semantics of the intermediate recognition result is complete, the expressed meaning is relatively complete, and it is a complete sentence. Thus, the intermediate recognition result can be used as a complete sentence, and then a first reply sentence can be determined according to the intermediate recognition result; otherwise, when the fourth probability is less than or equal to the preset threshold, it is determined that the semantics of the intermediate recognition result is incomplete, the expressed meaning is not complete enough, and it may not be a complete sentence, and the intermediate recognition result is unreliable. The preset threshold can be set according to actual needs. For example, it can be set to 0.5 and can be adjusted according to actual needs.
[0109] The neural network model in the embodiments of the present disclosure can be any neural network model based on deep learning methods, such as Deep Neural Network (DNN), Long Short Term Memory (LSTM), a model based on the multi-head attention mechanism (transformer), etc. The embodiments of the present disclosure do not limit this.
[0110] Based on this embodiment, based on the word vector of the intermediate recognition result, the first probability that the intermediate recognition result is used as the prefix in the historical final recognition result and the second probability as the historical final recognition result, and the popularity, through the neural network model, in a deep learning manner, to predict the fourth probability that the semantics of the intermediate recognition result is complete, can avoid the problem that it is impossible to determine whether the semantics of the intermediate recognition result is complete or the determination result is inaccurate due to data sparsity in the statistical calculation method, can solve the generalization problem of each intermediate recognition result, realize the determination of whether the semantics of the intermediate recognition result is complete, and improve the accuracy of the determination result of whether the semantics of the intermediate recognition result is complete.
[0111] Optionally, in a possible implementation manner of this embodiment, in 202, it may further include:
[0112] Within a second preset time after receiving the request sentence, obtain the final recognition result of the request sentence, where the end time of the second preset time is later than the end point time of the voice activity detection of the request sentence. The second preset time is a period of time that starts counting after the tail sound of the request sentence sent by the user has dropped, and the length of this counting time period is greater than the length of the first preset time. It can be understood that the second preset time is also the duration of the silent state detected by the voice interaction device;
[0113] In response to not recognizing an intermediate recognition result with complete semantics from the at least one intermediate recognition result, use the final recognition result as the speech recognition result.
[0114] Based on this embodiment, when no semantically complete intermediate recognition result is recognized from the at least one intermediate recognition result, the final recognition result of the request statement can be used as the speech recognition result to determine the response statement, so as to ensure the accuracy of the speech interaction result and avoid the influence of incorrect speech interaction results on the user experience.
[0115] Figure 4 is a schematic diagram according to the fourth embodiment of the present disclosure, as Figure 4 shown. In Figure 2 or Figure 3 Based on the embodiment shown, in operation 203, determining the response statement based on the speech recognition result and the historical information may include:
[0116] 401. Based on the speech recognition result and the historical information, perform semantic parsing on the speech recognition result to obtain a semantic parsing result.
[0117] Optionally, in a possible implementation manner of this embodiment, the semantic parsing result may include: domain, intent, and slot information. Then in 105, when the domain, intent, and slot information in the first semantic parsing result respectively correspond and are consistent with the domain, intent, and slot information in the second semantic parsing result, it can be considered that the first semantic parsing result is consistent with the second semantic parsing result. Otherwise, if the first semantic parsing result and the second semantic parsing result are inconsistent in any one or more of the domain, intent, and slot information, it is considered that the first semantic parsing result is inconsistent with the second semantic parsing result, and the first response statement is not played.
[0118] Among them, the domain is the domain to which the request statement belongs, such as alarm clock, weather, music, etc. The intent is the specific intent of the request statement in the current domain. For example, in the alarm clock domain, there are intents such as setting an alarm clock and deleting an alarm clock. The slot is the specific slot information of the request statement in the current domain and intent. For example, for the request statement "What's the weather in Beijing", the domain is weather, the intent is to query the weather, and the slot information includes: slot "city", slot value is "Beijing"; for the request statement "I want to listen to Jay Chou's songs", the domain is music, the intent is music search, and the slot information is: slot "singer", slot value is "Jay Chou".
[0119] Optionally, in a possible implementation manner of this embodiment, a pre-trained semantic parsing model may be used to perform semantic parsing on the speech recognition result to obtain a semantic parsing result. The semantic parsing model in the embodiments of the present disclosure may be implemented based on a neural network model in a deep learning manner, such as DNN, LSTM, LSTM+CRF model, transformer, etc., and the embodiments of the present disclosure do not limit this.
[0120] 402. Obtain a broadcast script and a resource result based on the semantic parsing result and the historical information.
[0121] 403. Perform speech synthesis on the broadcast script and the resource result to obtain the reply statement.
[0122] Based on this embodiment, the speech recognition result can be semantically parsed based on the speech recognition result and the historical information. According to the semantic parsing result and the historical information, a dialogue service is called to perform the calculation of the dialogue model and resource acquisition. The obtained broadcast script and resource result are subjected to speech synthesis to obtain a reply statement for the user, so as to improve the accuracy of the user's reply result.
[0123] Optionally, in a possible implementation manner of this embodiment, in 203, when generating the current context information corresponding to the request statement based on the speech recognition result and the reply statement, the current context information can be generated based on the semantic parsing result, the broadcast script, and the resource result. The current context information can include, for example, but is not limited to: the domain in the semantic parsing result, and the slot information in the broadcast script and the resource result.
[0124] Based on this embodiment, the current context information can be generated based on the semantic parsing result, the broadcast script, and the resource result to provide the necessary context information for the server to reply to the next round of user request statements, and improve the accuracy of the user's reply result.
[0125] Figure 5 is a schematic diagram of a voice interaction process according to an embodiment of the present disclosure, as Figure 5 shown. Taking a specific voice interaction process as an example, the embodiment of the present disclosure will be further described. In a complete voice interaction process, the physical link corresponding to the process from the user sending a request statement to the server returning a reply statement is as follows:
[0126] The client performs functions such as wake-up, noise cancellation, and VAD start and end point detection of the audio of the user request statement. The user sends a request statement such as "8 o'clock tomorrow morning";
[0127] After receiving the audio, the client obtains historical information from the local storage module, and transmits the audio and the historical information to the central control module in the server through link 501. Taking the historical information as a specific round as an example, it can include the domain to which the previous round of user request belongs, the script (i.e., the broadcast script) asked by the server in the previous round (for example, for the previous round of user request statement "Set an alarm for me", the script asked by the server in the previous round is "Okay, what time do you want to set the alarm"), the slot information expected to be obtained in the previous round (i.e., the slot information in the previous round of reply statement), etc.;
[0128] The central control module sends the recognized audio to the speech recognizer via link 502 for speech recognition, and obtains the speech recognition result via link 503;
[0129] The central control module sends the device ID (used to uniquely identify a client) or user ID (used to uniquely identify a user) carried by the audio, the speech recognition result, and the historical information to the semantic parser via link 504. The semantic parser performs semantic parsing on the speech recognition result, and obtains the semantic parsing result of the speech recognition result via link 505, including domain, intent, and slot information. For example, for the speech recognition result "Beijing weather", the semantic parsing result is obtained: domain: weather, intent: get_weather, slots: {city: Beijing};
[0130] The central control module sends the semantic parsing result and the historical information to the dialogue service via link 506. The dialogue service obtains the broadcast script and the resource result. For example, the broadcast script and the resource result corresponding to "Beijing weather" are "The weather in Beijing is sunny today, the highest temperature is 28 degrees, and the lowest temperature is 20 degrees", and obtains the broadcast script and the resource result via link 507;
[0131] The central control module sends the broadcast script and the resource result to the speech synthesizer via link 508 for speech synthesis, and obtains the reply statement to be broadcast to the user via link 509;
[0132] The central control module generates the current context information based on the semantic parsing result, the broadcast script, and the resource result, and sends the reply statement and the current context information to the client together via link 510;
[0133] The client plays the reply statement, and stores the current context information as historical information in the storage module local to the client, so that when the user initiates the next round of voice interaction, the historical information of the most recent preset number of rounds is extracted from the storage module and sent to the central control module via 501.
[0134] Thus, the process of one round of voice interaction is completed.
[0135] In the embodiment of the present disclosure, when the client receives an interaction request initiated by the user, the received request statement and the historical information stored this time are sent to the server. The server performs speech recognition on the request statement to obtain a speech recognition result. Then, based on the speech recognition result and the historical information, a reply statement is determined, and based on the speech recognition result and the reply statement, the current context information corresponding to the request statement is generated. Furthermore, the reply statement and the current context information are sent to the client, so that the client plays the reply statement and updates and stores the current context information as historical information in the storage module local to the client.
[0136] In this way, by locally storing historical information on the client side, the cloud server no longer needs to maintain a dedicated previous context information storage, which can reduce the storage and maintenance costs of the cloud. There is no longer a need for the central control module in the voice interaction system to send and receive data by establishing a network connection with the previous context information storage and perform read / write operations on the previous context information storage, which can save time, reduce the human-computer interaction response time perceived by the user, and improve the user experience.
[0137] In addition, binding the user's uplink request statement with the historical information and sending it to the server, and binding the reply statement with the current previous context information and returning it to the client can ensure the consistency of the voice recognition result and the reply of the voice interaction system. That is, as long as the user's request statement is successfully recognized and the reply statement is successfully broadcast, the content of the previous context information is successfully stored and retrieved, which improves the correctness of the voice reply result.
[0138] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present disclosure is not limited by the described action sequence, because according to the present disclosure, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present disclosure.
[0139] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0140] Figure 6 is a schematic diagram according to the fifth embodiment of the present disclosure, as Figure 6 shown. The client 600 of this embodiment may include a storage module 601, a first receiving unit 602, a transceiver unit 603, a playback unit 604, and a storage processing unit 605. Among them, the storage module 601 is used to store historical information; the first receiving unit 602 is used to obtain historical information from the storage module 601 in response to receiving a request statement sent by the user; the transceiver unit 603 is used to send the request statement and the historical information to the server; and receive the reply statement returned by the server and the current previous context information corresponding to the request statement; the playback unit 604 is used to play the reply statement; the storage processing unit 605 is used to store the current previous context information as historical information in the storage module 601.
[0141] It should be noted that part or all of the client 600 in this embodiment may be an application located on the local terminal, i.e., the terminal device of the service provider, or may also be a functional unit such as a plug-in or a software development kit (SDK) set in the application located on the local terminal. This embodiment does not make special limitations in this regard.
[0142] It can be understood that the application may be a native program installed on the terminal or a web program of the browser on the terminal. This embodiment does not make special limitations in this regard.
[0143] Optionally, in a possible implementation manner of this embodiment, the first receiving unit 602 is specifically configured to: obtain the historical information of the most recent preset rounds from the storage module based on a preset rule.
[0144] Optionally, in a possible implementation manner of this embodiment, the storage processing unit 605 is specifically configured to: store the current upstream information as historical information in the storage module 601 based on the reception time sequence.
[0145] Optionally, in a possible implementation manner of this embodiment, the current upstream information may include, for example, but is not limited to: the domain in the semantic parsing result corresponding to the request statement, the broadcast script of the reply statement, the slot information in the reply statement, and so on.
[0146] Figure 7 is a schematic diagram according to the sixth embodiment of the present disclosure, as Figure 7 shown. The voice interaction device 700 in this embodiment is applied to a server and may include a second receiving unit 701, a voice recognition unit 702, a determination unit 703, a generation unit 704, and a sending unit 705. Among them, the second receiving unit 701 is configured to receive a request statement and historical information sent by the client; the voice recognition unit 702 is configured to perform voice recognition on the request statement to obtain a voice recognition result; the determination unit 703 is configured to determine a reply statement based on the voice recognition result and the historical information; the generation unit 704 is configured to generate the current upstream information corresponding to the request statement based on the voice recognition result and the reply statement; the sending unit 705 is configured to send the reply statement and the current upstream information to the client so that the client can play the reply statement and store the current upstream information as historical information in the storage module local to the client.
[0147] It should be noted that part or all of the voice interaction device 700 in this embodiment can be an application located on the local terminal, i.e., the terminal device of the service provider, or can also be a functional unit such as a plug-in or a software development kit (SDK) set in the application located on the local terminal. This embodiment does not make special limitations in this regard.
[0148] It can be understood that the application can be a native app installed on the terminal, or can also be a web app of a browser on the terminal. This embodiment does not make special limitations in this regard.
[0149] Optionally, in a possible implementation manner of this embodiment, the historical information may include, for example, but is not limited to: the client obtains the historical information of the most recent preset number of rounds from the storage module based on a preset rule.
[0150] Figure 8 is a schematic diagram according to the seventh embodiment of the present disclosure, as Figure 8 shown, on the basis of the embodiment Figure 7 shown, the voice interaction device 800 in this embodiment may further include a semantic integrity recognition unit 801. In this embodiment, the voice recognition unit 702 is specifically configured to: perform voice recognition on the request statement and obtain at least one intermediate recognition result within a first preset time after receiving the request statement; wherein, the end time of the first preset time is earlier than the end point time of the voice activity detection of the request statement. The semantic integrity recognition unit 801 is configured to, in response to obtaining the first intermediate recognition result among the at least one intermediate recognition results, sequentially recognize whether the semantics of the at least one intermediate recognition result is complete in the time order of obtaining the at least one intermediate recognition result; in response to recognizing the first semantically complete intermediate recognition result from the at least one intermediate recognition results, use the first semantically complete intermediate recognition result as the voice recognition result.
[0151] Optionally, in a possible implementation manner of this embodiment, the semantic integrity recognition unit 801 is specifically configured to: sequentially for each intermediate recognition result among the at least one intermediate recognition results, use a semantic integrity model to obtain a first probability that each intermediate recognition result is a prefix in the historical final recognition result, and a second probability that each intermediate recognition result is the historical final recognition result; the semantic integrity model is obtained by statistical calculation based on the final recognition results in the user log; determine a third probability that the semantics of each intermediate recognition result is complete based on the first probability and the second probability; and determine whether the semantics of each intermediate recognition result is complete according to whether the third probability is greater than a preset threshold.
[0152] Optionally, in another possible implementation of this embodiment, the semantic integrity recognition unit 801 is specifically configured to: sequentially obtain the word vectors of each intermediate recognition result in the at least one intermediate recognition result; obtain the first probability of each intermediate recognition result as a prefix in the historical final recognition result, and the second probability of each intermediate recognition result as the historical final recognition result; obtain the popularity of each intermediate recognition result; input the word vector, the first probability, the second probability, and the popularity into a neural network model, and output the fourth probability that each intermediate recognition result is semantically complete through the neural network model; determine whether the semantics of the current text is complete according to whether the fourth probability is greater than a preset threshold.
[0153] Optionally, in a possible implementation of this embodiment, the speech recognition unit 702 is specifically configured to: obtain the final recognition result of the request statement within a second preset time after receiving the request statement, where the end time of the second preset time is later than the tail point time of the voice activity detection of the request statement; in response to not recognizing a semantically complete intermediate recognition result from the at least one intermediate recognition result, use the final recognition result as the speech recognition result.
[0154] Optionally, in a possible implementation of this embodiment, the determination unit 703 may include (not shown in the figure): a semantic parsing module, configured to perform semantic parsing on the speech recognition result based on the speech recognition result and the historical information to obtain a semantic parsing result; a dialogue service module, configured to obtain a broadcast script and a resource result based on the semantic parsing result and the historical information; a speech synthesis module, configured to perform speech synthesis on the broadcast script and the resource result to obtain the reply statement.
[0155] Optionally, in a possible implementation of this embodiment, the generation unit 704 is specifically configured to: generate the current previous context information based on the semantic parsing result, the broadcast script, and the resource result; the current previous context information includes: the domain in the semantic parsing result, the slot information in the broadcast script and the resource result.
[0156] When the client in the embodiments of the present disclosure receives an interaction request initiated by a user, the received request statement and the historical information stored this time are sent to the server. The server performs speech recognition on the request statement to obtain a speech recognition result. Then, a reply statement is determined based on the speech recognition result and the historical information, and current context information corresponding to the request statement is generated based on the speech recognition result and the reply statement. Furthermore, the reply statement and the current context information are sent to the client, so that the client can play the reply statement and update and store the current context information as historical information in the storage module local to the client.
[0157] In this way, by locally storing historical information on the client, the cloud server no longer needs to maintain a dedicated context information storage, which can reduce the storage and maintenance costs of the cloud. There is no longer a need for the central control module in the voice interaction system to send and receive data by establishing a network connection with the context information storage and perform read / write operations on the context information storage, which can save time, reduce the human-computer interaction response time perceived by the user, and improve the user experience.
[0158] In addition, by binding the user's uplink request statement and historical information and sending them to the server, and binding the reply statement and the current context information and returning them to the client, the consistency of the speech recognition result and the reply of the voice interaction system can be ensured. That is, as long as the user's request statement is successfully recognized and the reply statement is successfully broadcast, the content of the context information is successfully stored and retrieved, which improves the correctness of the voice reply result.
[0159] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information and the like all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0160] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product. Further, an autonomous vehicle including the provided electronic device is also provided.
[0161] Figure 9 FIG. shows a schematic block diagram of an exemplary electronic device 900 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0162] As Figure 9As shown, device 900 includes a computing unit 901 which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0163] Multiple components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disc, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0164] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as the method of voice interaction. For example, in some embodiments, the method of voice interaction can be implemented as a computer software program which is tangibly contained in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the method of voice interaction described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the method of voice interaction in any other appropriate way (e.g., by means of firmware).
[0165] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0166] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0167] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0168] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0169] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0170] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0171] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0172] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for voice interaction, the method comprising: In response to receiving a request statement sent by a user, obtaining historical information from a local storage module, and sending the request statement and the historical information to a server; Receiving a reply statement returned by the server and current context information corresponding to the request statement, wherein the reply statement is determined by the server based on a speech recognition result of the request statement and the historical information, and the server obtains the speech recognition result of the request statement in the following manner: performing speech recognition on the request statement, and within a first preset time after receiving the request statement, obtaining at least one intermediate recognition result; wherein an end time of the first preset time is earlier than a tail point time of voice activity detection of the request statement; in response to obtaining a first intermediate recognition result among the at least one intermediate recognition results, sequentially recognizing whether semantics of the at least one intermediate recognition result is complete according to a time order of obtaining the at least one intermediate recognition result, based on a first probability that each intermediate recognition result is a prefix in a historical final recognition result and a second probability that each intermediate recognition result is the historical final recognition result; in response to recognizing a first intermediate recognition result with complete semantics from the at least one intermediate recognition results, using the first intermediate recognition result with complete semantics as the speech recognition result; Playing the reply statement, and storing the current context information as historical information in the storage module.
2. The method according to claim 1, wherein, The obtaining historical information from the local storage module includes: Obtaining historical information of a most recent preset number of turns from the storage module based on a preset rule.
3. The method according to claim 1 or 2, wherein, The storing the current context information as historical information in the storage module includes: Storing the current context information as historical information in the storage module based on a reception time order.
4. The method according to claim 1, wherein, The current context information includes: a domain in a semantic parsing result corresponding to the request statement, a broadcast script of the reply statement, and slot information in the reply statement.
5. A method for voice interaction, the method comprising: Receiving a request statement and historical information sent by a client; Performing speech recognition on the request statement to obtain a speech recognition result; Determining a reply statement based on the speech recognition result and the historical information, and generating current context information corresponding to the request statement based on the speech recognition result and the reply statement; Sending the reply statement and the current context information to the client, so that the client plays the reply statement and stores the current context information as historical information in a storage module local to the client; Wherein, the performing speech recognition on the request statement to obtain a speech recognition result includes: Performing speech recognition on the request statement, and within a first preset time after receiving the request statement, obtaining at least one intermediate recognition result; wherein an end time of the first preset time is earlier than a tail point time of voice activity detection of the request statement; In response to obtaining the first intermediate recognition result among the at least one intermediate recognition result, in the chronological order of obtaining the at least one intermediate recognition result, based on the first probability of each intermediate recognition result as a prefix in the historical final recognition result and the second probability of each intermediate recognition result as the historical final recognition result, sequentially identify whether the semantics of the at least one intermediate recognition result is complete; In response to identifying the first intermediate recognition result with complete semantics from the at least one intermediate recognition result, use the first intermediate recognition result with complete semantics as the speech recognition result.
6. The method according to claim 5, wherein The historical information includes: the client obtains the historical information of the most recent preset rounds from the storage module based on a preset rule.
7. The method according to claim 5, wherein The sequentially identifying whether the semantics of the at least one intermediate recognition result is complete based on the first probability of each intermediate recognition result as a prefix in the historical final recognition result and the second probability of each intermediate recognition result as the historical final recognition result includes: Sequentially for each intermediate recognition result among the at least one intermediate recognition result, use a semantic integrity model to obtain the first probability of each intermediate recognition result as a prefix in the historical final recognition result, and the second probability of each intermediate recognition result as the historical final recognition result; the semantic integrity model is obtained by statistical calculation based on the final recognition results in the user logs; Determine the third probability of the semantics of each intermediate recognition result being complete based on the first probability and the second probability; Determine whether the semantics of each intermediate recognition result is complete according to whether the third probability is greater than a preset threshold.
8. The method according to claim 5, wherein The sequentially identifying whether the semantics of the at least one intermediate recognition result is complete based on the first probability of each intermediate recognition result as a prefix in the historical final recognition result and the second probability of each intermediate recognition result as the historical final recognition result includes: Sequentially for each intermediate recognition result among the at least one intermediate recognition result, obtain the word vector of each intermediate recognition result; Obtain the first probability of each intermediate recognition result as a prefix in the historical final recognition result, and the second probability of each intermediate recognition result as the historical final recognition result; Obtain the popularity of each intermediate recognition result; Input the word vector, the first probability, the second probability, and the popularity into a neural network model, and output the fourth probability of the semantics of each intermediate recognition result being complete through the neural network model; Determine whether the semantics of each intermediate recognition result is complete according to whether the fourth probability is greater than a preset threshold.
9. The method according to claim 5, wherein The performing speech recognition on the request statement to obtain a speech recognition result further includes: Obtain the final recognition result of the request statement within a second preset time after receiving the request statement, where the end moment of the second preset time is later than the tail point moment of the voice activity detection of the request statement; In response to not identifying an intermediate recognition result with complete semantics from the at least one intermediate recognition result, use the final recognition result as the speech recognition result.
10. The method according to claim 5, wherein, The determining a reply statement based on the speech recognition result and the historical information includes: Based on the speech recognition result and the historical information, perform semantic parsing on the speech recognition result to obtain a semantic parsing result; Based on the semantic parsing result and the historical information, obtain a broadcast script and a resource result; Perform speech synthesis on the broadcast script and the resource result to obtain the reply statement.
11. The method according to claim 10, wherein, The generating the current context information corresponding to the request statement based on the speech recognition result and the reply statement includes: Generate the current context information based on the semantic parsing result, the broadcast script, and the resource result; the current context information includes: the domain in the semantic parsing result, the slot information in the broadcast script and the resource result.
12. A client, comprising: A storage module, configured to store historical information; A first receiving unit, configured to obtain historical information from the storage module in response to receiving a request statement sent by a user; A transceiver unit, configured to send the request statement and the historical information to a server; And receive a reply statement returned by the server and the current context information corresponding to the request statement, wherein the reply statement is determined by the server based on the speech recognition result of the request statement and the historical information, and the server obtains the speech recognition result of the request statement in the following manner: perform speech recognition on the request statement, and within a first preset time after receiving the request statement, obtain at least one intermediate recognition result; wherein, the end moment of the first preset time is earlier than the end moment of the voice activity detection of the request statement; in response to obtaining the first intermediate recognition result in the at least one intermediate recognition result, in the time order of obtaining the at least one intermediate recognition result, based on the first probability of each intermediate recognition result as a prefix in the historical final recognition result and the second probability of each intermediate recognition result as the historical final recognition result, sequentially identify whether the semantics of the at least one intermediate recognition result is complete; in response to identifying the first semantically complete intermediate recognition result from the at least one intermediate recognition result, use the first semantically complete intermediate recognition result as the speech recognition result; A playback unit, configured to play the reply statement; A storage processing unit, configured to store the current context information as historical information in the storage module.
13. The client according to claim 12, wherein, The first receiving unit is specifically configured to: Obtain the historical information of the most recent preset number of rounds from the storage module based on a preset rule.
14. The client according to claim 12 or 13, wherein, The storage processing unit is specifically configured to: Store the current context information as historical information in the storage module based on the reception time order.
15. The client according to claim 12, wherein, The current context information includes: the domain in the semantic parsing result corresponding to the request statement, the broadcast script of the reply statement, and the slot information in the reply statement.
16. A voice interaction device, applied to a server, the device comprising: A second receiving unit, configured to receive a request statement and historical information sent by a client; A speech recognition unit, configured to perform speech recognition on the request statement to obtain a speech recognition result; A determination unit, configured to determine a reply statement based on the speech recognition result and the historical information; A generation unit, configured to generate current context information corresponding to the request statement based on the speech recognition result and the reply statement; A sending unit, configured to send the reply statement and the current context information to the client, so that the client plays the reply statement and stores the current context information as historical information in a storage module local to the client; Wherein, the speech recognition unit is specifically configured to: perform speech recognition on the request statement, and obtain at least one intermediate recognition result within a first preset time after receiving the request statement; wherein, the end time of the first preset time is earlier than the tail point time of the voice activity detection of the request statement; The apparatus further includes: A semantic integrity recognition unit, configured to, in response to obtaining the first intermediate recognition result among the at least one intermediate recognition results, sequentially recognize whether the semantics of the at least one intermediate recognition result is complete according to the time sequence of obtaining the at least one intermediate recognition result, based on a first probability that each intermediate recognition result is a prefix in a historical final recognition result and a second probability that each intermediate recognition result is the historical final recognition result; in response to recognizing the first intermediate recognition result with complete semantics among the at least one intermediate recognition results, using the first intermediate recognition result with complete semantics as the speech recognition result.
17. The apparatus according to claim 16, wherein, The historical information includes: the client obtains historical information of the most recent preset number of rounds from the storage module based on a preset rule.
18. The apparatus according to claim 16, wherein, The semantic integrity recognition unit is specifically configured to: Sequentially for each intermediate recognition result among the at least one intermediate recognition results, use a semantic integrity model to obtain a first probability that each intermediate recognition result is a prefix in a historical final recognition result, and a second probability that each intermediate recognition result is the historical final recognition result; the semantic integrity model is obtained by statistical calculation based on final recognition results in user logs; Determine a third probability that the semantics of each intermediate recognition result is complete based on the first probability and the second probability; Determine whether the semantics of each intermediate recognition result is complete according to whether the third probability is greater than a preset threshold.
19. The device according to claim 16, wherein, The semantic integrity recognition unit is specifically configured to: Sequentially for each intermediate recognition result among the at least one intermediate recognition results, obtain a word vector of each intermediate recognition result; Obtain a first probability that each intermediate recognition result is a prefix in a historical final recognition result, and a second probability that each intermediate recognition result is the historical final recognition result; Obtain the popularity of each intermediate recognition result; Input the word vector, the first probability, the second probability, and the popularity into a neural network model, and output a fourth probability that the semantics of each intermediate recognition result is complete through the neural network model; Determine whether the semantics of each intermediate recognition result is complete according to whether the fourth probability is greater than a preset threshold.
20. The apparatus according to claim 16, wherein The speech recognition unit is specifically configured to: Within a second preset time after receiving the request statement, obtain a final recognition result of the request statement, where the end time of the second preset time is later than the tail point time of the voice activity detection of the request statement; In response to not recognizing a semantically complete intermediate recognition result from the at least one intermediate recognition result, use the final recognition result as the speech recognition result.
21. The apparatus according to claim 16, wherein, The determining unit includes: A semantic parsing module, configured to perform semantic parsing on the speech recognition result based on the speech recognition result and the historical information to obtain a semantic parsing result; A dialogue service module, configured to obtain a broadcast speech script and a resource result based on the semantic parsing result and the historical information; A speech synthesis module, configured to perform speech synthesis on the broadcast speech script and the resource result to obtain the reply statement.
22. The device according to claim 21, wherein, The generating unit is specifically configured to: Generate the current previous context information based on the semantic parsing result, the broadcast speech script, and the resource result; the current previous context information includes: the domain in the semantic parsing result, the broadcast speech script, and the slot information in the resource result.
23. A voice interaction system, including the client according to any one of claims 12-15 and the voice interaction device according to any one of claims 16-22.
24. An electronic device, including: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-11.
25. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-11.
26. A computer program product, including a computer program, where the computer program, when executed by a processor, implements the method according to any one of claims 1-11.
Citation Information
Patent Citations
Multi-round interaction semantic understanding method and device, and computer storage medium
CN111429895A
Voice interaction method and device, electronic equipment and storage medium
CN112466302A