Method and device for improving voice interaction response speed and storage medium

By calling the intent parsing program in parallel and querying the preset intent cache library, the problem of the response speed of the voice interaction system is too slow, and fast and accurate intent parsing is achieved, improving the efficiency and user experience of voice interaction.

CN119964562AActive Publication Date: 2025-05-09HAIER YOUJIA INTELLIGENT TECH (BEIJING) CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202411883735.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-05-09
Estimated Expiration
2044-12-19

AI Technical Summary

Technical Problem

The existing voice interaction system has efficiency problems when processing voice commands, resulting in significant time delays between speech recognition and intention resolution, and the response speed is too slow, which affects user experience and processing efficiency.

Method used

By calling the intent parsing program in parallel and querying the preset intent cache library, text data is obtained for intent parsing. If the first intent data and the second intent data are inconsistent, compensation processing is performed to determine the target intent data.

Benefits of technology

It effectively improves the accuracy and response speed of voice interaction command processing, improves the efficiency and accuracy of voice interaction, enhances the flexibility and intelligence of intention recognition, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964562A_ABST
    Figure CN119964562A_ABST
Patent Text Reader

Abstract

The invention discloses a method and device for improving the voice interaction response speed and a storage medium, and relates to the technical field of smart home, and the method comprises the steps: obtaining text data obtained after the recognition of a voice interaction instruction of a target object; performing intention analysis on the text data to obtain first intention data, and querying second intention data corresponding to the text data from a preset intention cache library; and when it is determined that the first intention data and the second intention data are inconsistent, performing compensation processing on the second intention data based on the first intention data to obtain third intention data, and determining target intention data according to the third intention data. The technical problem that the response speed of voice interaction is too slow is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of smart home technology, and more specifically, to a method, device and storage medium for improving voice interaction response speed. Background Art

[0002] With the rapid development of information technology, voice interaction technology has become an important part of the field of human-computer interaction. This technology realizes natural communication between people and machines by converting users' voice instructions into commands that machines can understand. The application scope of voice interaction technology is expanding, from smart homes to mobile devices to car systems, covering almost every aspect of daily life.

[0003] In a typical voice interaction system, users use voice to issue voice interaction commands such as "turn on the TV" or "check tomorrow's weather". The process of processing these commands usually includes two key stages: speech recognition and intent analysis. The speech recognition stage is responsible for converting the user's voice signal into text information, while the intent analysis stage analyzes this text information to determine the user's intention and perform the corresponding operation.

[0004] However, existing voice interaction systems have efficiency issues when processing voice commands. Specifically, the voice recognition process can usually be completed quickly, while intent analysis takes longer to process due to its complex logic and real-time requirements. This speed mismatch leads to a significant time delay between voice recognition and intent analysis, which increases the overall interaction time. This delay causes users to wait longer for a response from the system, that is, the response speed of voice interaction is too slow, which not only has a negative impact on the user experience, but also reduces the efficiency of voice interaction command processing.

[0005] Therefore, in the related art, there is a technical problem that the response speed of voice interaction is too slow.

[0006] With regard to the technical problem in related technologies that the response speed of voice interaction is too slow, no effective solution has been proposed yet. Summary of the invention

[0007] The embodiments of the present application provide a method, device and storage medium for improving the response speed of voice interaction, so as to at least solve the technical problem of slow response speed of voice interaction in related technologies.

[0008] According to one embodiment of the embodiments of the present application, a method for improving the response speed of voice interaction is provided, including: obtaining text data obtained after recognizing the voice interaction instructions of the target object; performing intent analysis on the text data to obtain first intent data, and querying second intent data corresponding to the text data from a preset intent cache library; when it is determined that the first intent data and the second intent data are inconsistent, performing compensation processing on the second intent data based on the first intent data to obtain third intent data, and determining the target intent data based on the third intent data.

[0009] In an exemplary embodiment, intent parsing is performed on the text data to obtain first intent data, including: using an intent parsing program to perform natural language processing on the text data to obtain the first intent data; performing text parsing on the preprocessed text data to obtain entity words, inputting the entity words into an intent recognition model to obtain intent categories output by the intent recognition model, wherein the intent recognition model is used to map the input words to predefined intent categories; and generating the first intent data according to a default intent framework corresponding to the entity words and the intent categories.

[0010] In an exemplary embodiment, compensation processing is performed on the second intent data based on the first intent data to obtain the third intent data, including: sending the queried multiple second intent data to the target object, and receiving the fourth intent data selected by the target object from the multiple second intent data; respectively calculating the first intent confidence between the first intent data and the fourth intent data, and the second intent confidence between the second intent data and the fourth intent data; when it is determined that the first intent confidence satisfies a preset matching condition, updating the second intent data based on the first intent data to obtain the third intent data.

[0011] In an exemplary embodiment, after determining the target intent data based on the third intent data, the method also includes: sending the queried multiple second intent data to the target object, and receiving fourth intent data selected by the target object from the multiple second intent data; calculating the third intent confidence between the target intent data and the fourth intent data; when it is determined that the third intent confidence does not meet the preset matching conditions, generating a query statement based on the target intent data, and sending the query statement to the target object, wherein the query statement is used to inquire the target object whether it recognizes the target intent data; and determining whether to adjust the target intent data based on the target object's reply statement based on the query statement.

[0012] In an exemplary embodiment, before querying the second intent data corresponding to the text data from the preset intent cache, the method also includes: obtaining historical high-frequency corpus with an interaction frequency higher than a preset frequency, and historical intent data corresponding to the historical high-frequency corpus from historical interaction data; establishing the preset intent cache based on the historical high-frequency corpus and the historical intent data corresponding to the historical high-frequency corpus; classifying the historical high-frequency corpus according to the object behavior of the target object in the preset intent cache to obtain historical behavior corpora corresponding to different object behaviors; querying the second intent data corresponding to the text data from the preset intent cache, including: determining the target behavior corpus belonging to the target object from the historical behavior corpora corresponding to the different object behaviors, and querying the second intent data corresponding to the text data from the historical intent data corresponding to the target behavior corpus.

[0013] In an exemplary embodiment, querying the second intent data corresponding to the text data from a preset intent cache library includes: obtaining a first text vector corresponding to the text data, and obtaining a second text vector of cache data in the preset intent cache library; using a word vector model to calculate the cosine similarity between the first text vector and the second text vector; obtaining the intent cache data whose cosine similarity exceeds a preset threshold from the cache data; splitting the text data to obtain multiple keywords, and querying the second intent data corresponding to the multiple keywords from the intent cache data.

[0014] In an exemplary embodiment, the method also includes: determining the time when the first intention data is obtained and the time when the second intention data is queried; when it is determined that the time when the first intention data is obtained is earlier than the query time, determining the target intention data according to the first intention data; when it is determined that the time when the first intention data is obtained is later than the query time, determining the target intention data according to the second intention data.

[0015] According to another aspect of an embodiment of the present application, a device for improving the response speed of voice interaction is also provided, including: an obtaining module, used to obtain text data obtained after recognizing the voice interaction instructions of the target object; a query module, used to perform intent analysis on the text data to obtain first intent data, and query second intent data corresponding to the text data from a preset intent cache library; a determination module, used to perform compensation processing on the second intent data based on the first intent data to obtain third intent data when it is determined that the first intent data and the second intent data are inconsistent, and determine the target intent data based on the third intent data.

[0016] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned method for improving the voice interaction response speed when running.

[0017] According to another aspect of an embodiment of the present application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the method for improving the voice interaction response speed through the computer program.

[0018] In an embodiment of the present application, the text data obtained after the voice interaction instruction of the target object is recognized is obtained; the text data is analyzed for intent to obtain the first intent data, and the second intent data corresponding to the text data is queried from the preset intent cache library; when it is determined that the first intent data and the second intent data are inconsistent, the second intent data is compensated based on the first intent data to obtain the third intent data, and the target intent data is determined according to the third intent data; the above technical solution is adopted to solve the technical problem that the response speed of voice interaction is too slow, that is, by calling the intent analysis program in parallel and querying the preset intent cache library, the intention of the user's voice instruction can be quickly and accurately analyzed, which effectively improves the accuracy and response speed of voice interaction instruction processing, improves the response speed of voice interaction, and improves the efficiency and accuracy of voice interaction. In addition, when the first intent data and the second intent data are inconsistent, the flexibility and intelligence of intent recognition are improved through compensation processing, combined with the user's historical behavior and the current intent recognition results, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0020] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0021] Figure 1 It is a hardware environment diagram of a method for improving voice interaction response speed in an embodiment of the present application;

[0022] Figure 2 is a flow chart of a method for improving voice interaction response speed according to an embodiment of the present application;

[0023] Figure 3 is a schematic diagram of a process for improving voice interaction response speed according to an embodiment of the present application;

[0024] Figure 4 is a flowchart of a process for improving voice interaction response speed according to an embodiment of the present application (I);

[0025] Figure 5 is a flowchart of a process for improving voice interaction response speed according to an embodiment of the present application (II);

[0026] Figure 6 is a flowchart of a process for improving voice interaction response speed according to an embodiment of the present application (III);

[0027] Figure 7 It is a structural block diagram of a device for improving voice interaction response speed according to an embodiment of the present application. DETAILED DESCRIPTION

[0028] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0029] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0030] According to one aspect of an embodiment of the present application, a method for improving the response speed of voice interaction is provided. The method for improving the response speed of voice interaction is widely used in smart home (Smart Home), smart home, smart home device ecology, smart residential (Intelligence House) ecology and other whole-house intelligent digital control application scenarios. Optionally, in this embodiment, the above-mentioned method for improving the response speed of voice interaction can be applied to Figure 1In the hardware environment composed of the terminal device 102 and the server 104 shown in FIG. Figure 1 As shown, the server 104 is connected to the terminal device 102 via a network, and can be used to provide services (such as application services, etc.) for the terminal or a client installed on the terminal. A database can be set on the server or independently of the server to provide data storage services for the server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data computing services for the server 104.

[0031] The network may include but is not limited to at least one of the following: wired network, wireless network. The wired network may include but is not limited to at least one of the following: wide area network, metropolitan area network, local area network, and the wireless network may include but is not limited to at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal device 102 may be but is not limited to a PC, a mobile phone, a tablet computer, a smart air conditioner, a smart range hood, a smart refrigerator, a smart oven, a smart stove, a smart washing machine, a smart water heater, a smart washing device, a smart dishwasher, a smart projection device, a smart TV, a smart clothes drying rack, a smart curtain, a smart audio and video, a smart socket, a smart speaker, a smart fresh air device, a smart kitchen and bathroom device, a smart bathroom device, a smart sweeping robot, a smart window cleaning robot, a smart mopping robot, a smart air purification device, a smart steamer, a smart microwave oven, a smart kitchen treasure, a smart purifier, a smart water dispenser, a smart door lock, etc.

[0032] In this embodiment, a method for improving the voice interaction response speed is provided, which is applied to the above terminal device. Figure 2 : is a flowchart of a method for improving voice interaction response speed according to an embodiment of the present application, the process includes the following steps:

[0033] Step S202, obtaining text data obtained after recognizing the voice interaction instruction of the target object;

[0034] Step S204, performing intent analysis on the text data to obtain first intent data, and querying second intent data corresponding to the text data from a preset intent cache library;

[0035] Among them, the process of performing intent analysis on the text data to obtain the first intent data may include: responding to a parallel call instruction, calling the intent analysis program of the text data, and using the intent analysis program to perform intent analysis on the text data to obtain the first intent data.

[0036] Among them, the two processes of performing intent parsing on the text data and querying the second intent data corresponding to the text data from the preset intent cache library are performed simultaneously and in parallel.

[0037] For example Figure 3 As shown, during voice interaction, the intent parsing process of NLP semantic analysis and the cache query process can be called in parallel. If the NLP intent cache is hit, the cache result will be obtained first. At this time, the cache result can be used for subsequent processing first, that is, the obtained result is used first for subsequent other processing.

[0038] Step S206, when it is determined that the first intention data and the second intention data are inconsistent, compensation processing is performed on the second intention data based on the first intention data to obtain third intention data, and target intention data is determined based on the third intention data.

[0039] Through the above steps, the text data obtained after the voice interaction instruction of the target object is recognized is obtained; the text data is analyzed for intent to obtain the first intent data, and the second intent data corresponding to the text data is queried from the preset intent cache library; when it is determined that the first intent data and the second intent data are inconsistent, the second intent data is compensated based on the first intent data to obtain the third intent data, and the target intent data is determined according to the third intent data; the above technical solution solves the technical problem of the slow response speed of voice interaction, that is, by calling the intent analysis program in parallel and querying the preset intent cache library, the intention of the user's voice instruction can be quickly and accurately analyzed, which effectively improves the accuracy and response speed of voice interaction instruction processing, improves the response speed of voice interaction, and improves the efficiency and accuracy of voice interaction. In addition, when the first intent data and the second intent data are inconsistent, the flexibility and intelligence of intent recognition are improved through compensation processing, combined with the user's historical behavior and the current intent recognition results, thereby improving the user experience.

[0040] It can be understood that the present application optimizes the internal logic of voice interaction processing by calling the intent resolution program in parallel and querying the response mechanism of the preset intent cache library, which can improve the voice interaction response speed and reduce the interaction time. It is also applicable to the situation where there are multiple intent resolution programs, and improves the efficiency of improving the voice interaction response speed.

[0041] In an alternative embodiment, if Figure 4 As shown, a process for implementing voice interaction instructions is proposed, which specifically includes:

[0042] Step 1: The user initiates a voice command.

[0043] Step 2, use the device to record.

[0044] Step 3: Perform automatic speech recognition (ASR) on the audio stream sent by the device in the cloud.

[0045] Step 4: Is it a high-frequency corpus? If yes, go to step 5; otherwise, go to step 6.

[0046] Step 5: Does the intent cache match? If yes, go to step 7; otherwise, go to step 6.

[0047] Step 6: Perform natural language processing (NLP).

[0048] Based on the above steps, it is also possible to achieve Figure 5 The following steps are shown:

[0049] 1. Get the ASR recognition result.

[0050] 2. Perform intent parsing and query cache in parallel.

[0051] 3. Does the recognition result meet the cache conditions? If yes, go to step 4; otherwise, go to step 2.

[0052] 4. Get the cached results.

[0053] 5. Call the NLP interface and further perform the NLP analysis in 5.1.

[0054] 6. Write cache results.

[0055] 7. If no cached result is found, wait for the NLP parsing result of calling the NLP interface.

[0056] 8. Subsequent processing.

[0057] 9. If the cached result and the NLP parsing result are inconsistent, update the cached result.

[0058] If the processing result of the cached result is inconsistent with the processing result of the NLP parsing result, the NLP content of the subsequent processing is changed.

[0059] Step 7, use intent to cache the results.

[0060] Step 8, skill processing.

[0061] Step 9: Generate the text for the broadcast.

[0062] Step 10, speech synthesis, can be achieved using TTS (Text-to-Speech).

[0063] Step 11: The device plays audio according to the audio stream sent from the cloud.

[0064] Step 12: The user hears the voice announcement.

[0065] In an exemplary embodiment, the process of performing intent parsing on the text data to obtain the first intent data specifically includes: using an intent parsing program to perform natural language processing on the text data to obtain the first intent data; performing text parsing on the preprocessed text data to obtain entity words, inputting the entity words into an intent recognition model to obtain intent categories output by the intent recognition model, wherein the intent recognition model is used to map the input words to predefined intent categories; and generating the first intent data according to a default intent framework corresponding to the entity words and the intent categories.

[0066] In this embodiment, the process of performing natural language processing to obtain the first intention data can also be described in combination with the following content: the original text of the text data is cleaned, irrelevant characters, punctuation and stop words are removed, and text normalization is performed, such as case conversion, standardization of numbers and text, etc. Then the sentence is decomposed into words or phrases, and the part of speech of each word is marked, such as noun, verb, adjective, etc., which helps to understand the basic structure of the text. Then the entity words such as names of people, places, and organizations in the text are identified, and the structure of the sentence is analyzed to identify the subject, predicate, object and other components, and understand the basic semantics of the sentence, which is crucial for understanding and locating specific information. Then the text is converted into a structured semantic representation, such as a logical form or a semantic framework. A machine learning model (equivalent to an intention recognition model) is used to identify the user's intention on the data after the above processing. Then the entities related to the specific intention are identified, and the slots in the intention framework are filled, such as the specific song name or singer name needs to be identified in the "play music" intention. Finally, the result of the intention recognition can be verified to ensure its logic and rationality. Through natural language processing, clear intentions can be extracted to provide information for further dialogue management and action execution.

[0067] Optionally, the above-mentioned intent recognition model is trained with historical entity words as input samples and intent categories corresponding to the historical entity words as output samples. For example, existing learning models such as recurrent neural networks and long short-term memory networks can be trained so that the model learns how to extract key information from entity words and map them to intent categories. When the user inputs new corpus, the model will output an intent category, or a probability distribution of a series of intent categories, which improves the accuracy and robustness of the intent recognition model and can handle a variety of language expressions and context dependencies.

[0068] It can be understood that generating the first intent data according to the default intent framework corresponding to the entity words and the intent category means filling the entity words into the corresponding slots in the default intent framework.

[0069] In an exemplary embodiment, the implementation process of performing compensation processing on the second intention data based on the first intention data to obtain the third intention data may include: sending the queried multiple second intention data to the target object, and receiving the fourth intention data selected by the target object from the multiple second intention data; respectively calculating the first intention confidence between the first intention data and the fourth intention data, and the second intention confidence between the second intention data and the fourth intention data; when it is determined that the first intention confidence meets the preset matching condition, updating the second intention data based on the first intention data to obtain the third intention data. This embodiment can dynamically select intent data based on confidence. For example, when the first intention confidence is higher than the second intention confidence, the first intention data can also be updated based on the second intention data to obtain the third intention data.

[0070] The process of determining whether the first intention confidence satisfies a preset matching condition includes: determining that the first intention confidence is higher than the second intention confidence.

[0071] In an exemplary embodiment, after determining the target intent data according to the third intent data, further, the queried multiple second intent data can also be sent to the target object, and the fourth intent data selected by the target object from the multiple second intent data can be received; the third intent confidence between the target intent data and the fourth intent data is calculated; when it is determined that the third intent confidence does not meet the preset matching conditions, an inquiry statement is generated based on the target intent data, and the inquiry statement is sent to the target object, wherein the inquiry statement is used to inquire whether the target object recognizes the target intent data; and whether to adjust the target intent data is determined based on the target object's reply statement based on the inquiry statement. In addition, by asking the user whether to recognize the target intent data, the interactivity and intelligence of the system are further enhanced, and the accuracy of voice command processing is ensured.

[0072] The process of determining whether the first intention confidence satisfies a preset matching condition includes: determining that the third intention confidence is lower than a preset intention confidence.

[0073] In this embodiment, if the target intention data is adjusted according to the reply statement of the target object based on the query statement, the target intention data can be regenerated or deleted.

[0074] In an exemplary embodiment, before querying the second intent data corresponding to the text data from the preset intent cache, further, obtain historical high-frequency corpus with an interaction frequency higher than the preset frequency and the historical intent data corresponding to the historical high-frequency corpus from the historical interaction data; establish the preset intent cache according to the historical high-frequency corpus and the historical intent data corresponding to the historical high-frequency corpus; classify the historical high-frequency corpus according to the object behavior of the target object in the preset intent cache to obtain historical behavior corpus corresponding to different object behaviors; query the second intent data corresponding to the text data from the preset intent cache, including: determining the target behavior corpus belonging to the target object from the historical behavior corpus corresponding to the different object behaviors, and querying the second intent data corresponding to the text data from the historical intent data corresponding to the target behavior corpus. By caching the historical intent data corresponding to the historical high-frequency corpus, in the subsequent user voice interaction, the step of real-time intent parsing can be skipped in some cases, so as to achieve the purpose of reducing the duration, thereby improving the interaction efficiency.

[0075] Among them, Figure 3 As shown, before or during the interaction, you can use high-frequency corpus to build an NLP cache corpus (i.e., a preset intent cache library). By caching high-frequency corpus intents, when a high-frequency corpus intent is hit, there is no need to call NLP, which can reduce the interaction time.

[0076] Optionally, in one embodiment, the historical interaction data, for example, represents the historical records of voice interaction between the user and the device, and has attributes such as time, device, corpus, intent, etc. The historical records are shown in Table 1 below.

[0077] Table 1 History table

[0078]

[0079]

[0080] Optionally, in one embodiment, the proportion of a type of corpus in the total number of interactions = the number of interactions of a type of corpus / the total number of interactions * 100%. By comparing the proportion of a type of corpus in the total number of interactions with a set threshold, it can be determined whether this part of the corpus is high-frequency corpus, and the high-frequency corpus is stored in the cache. This embodiment proposes an automatic screening mechanism for high-frequency corpus, which can automatically determine whether the corpus has cache value, thereby improving the accuracy of querying data from the cache.

[0081] In an exemplary embodiment, the technical solution for querying the second intent data corresponding to the text data from the preset intent cache library may include: obtaining the first text vector corresponding to the text data, and obtaining the second text vector of the cache data in the preset intent cache library; using the word vector model to calculate the cosine similarity between the first text vector and the second text vector; obtaining the intent cache data whose cosine similarity exceeds the preset threshold from the cache data; splitting the text data to obtain multiple keywords, and querying the second intent data corresponding to the multiple keywords from the intent cache data. In this embodiment, the second intent data corresponding to all the multiple keywords can be queried from the intent cache data, or the second intent data corresponding to some keywords can be queried from the intent cache data. In this embodiment, the similarity between the text data and the cached corpus is calculated by the word vector model, which realizes the mechanism of data fuzzy matching and speeds up the data matching.

[0082] Optionally, in one embodiment, the process of querying the second intent data corresponding to the keyword can be to pre-define the keywords corresponding to the intent, and directly query the intent cache data containing these keywords. For example, the corpus "play children's songs" can be accurately matched with keywords such as "children's songs" and "play", so as to quickly locate the intent "play children's songs". This method has fast processing speed and high processing efficiency.

[0083] In an exemplary embodiment, the time when the first intention data is obtained and the time when the second intention data is queried can also be determined; when it is determined that the time when the first intention data is obtained is earlier than the query time, the target intention data is determined based on the first intention data; when it is determined that the time when the first intention data is obtained is later than the query time, the target intention data is determined based on the second intention data.

[0084] In order to better understand the process of the above-mentioned method for improving the voice interaction response speed, the implementation method process of the above-mentioned method for improving the voice interaction response speed is described below in combination with an optional embodiment, but it is not used to limit the technical solution of the embodiment of the present application.

[0085] Optionally, in one embodiment, Figure 6 As shown in FIG. 1 , a process for improving the voice interaction response speed is also proposed, which specifically includes:

[0086] Step 1: The user initiates a voice command.

[0087] Step 2, use the device to record.

[0088] Step 3: Perform speech recognition (ASR) on the audio stream sent by the device in the cloud.

[0089] Step 4, Natural Language Processing (NLP).

[0090] Step 7, use intent to cache the results.

[0091] Step 5, skill processing.

[0092] Step 6: Generate the text for the broadcast.

[0093] Step 7: Text-to-speech synthesis (TTS).

[0094] Step 8: The device plays audio according to the audio stream sent from the cloud.

[0095] Step 9: The user hears the voice announcement.

[0096] In this embodiment, the speech recognition program can convert the speech instruction into text data in the speech recognition stage, and the intent analysis program can convert the obtained text data into intent data in the intent analysis stage. The intent data represents the core concept data corresponding to the text data. For example, if the text data is "turn on the TV", the corresponding intent data can be "turn on" or "TV".

[0097] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.

[0098] Figure 7 is a structural block diagram of a device for improving voice interaction response speed according to an embodiment of the present application; Figure 7 As shown, including:

[0099] An obtaining module 72 is used to obtain text data obtained after recognizing the voice interaction instruction of the target object;

[0100] A query module 74, configured to perform intent analysis on the text data to obtain first intent data, and query second intent data corresponding to the text data from a preset intent cache library;

[0101] The determination module 76 is used to perform compensation processing on the second intention data based on the first intention data to obtain third intention data when it is determined that the first intention data and the second intention data are inconsistent, and determine the target intention data according to the third intention data.

[0102] Through the above device, the text data obtained after the voice interaction instruction of the target object is recognized is obtained; the text data is analyzed for intent to obtain the first intent data, and the second intent data corresponding to the text data is queried from the preset intent cache library; when it is determined that the first intent data and the second intent data are inconsistent, the second intent data is compensated based on the first intent data to obtain the third intent data, and the target intent data is determined according to the third intent data; the above technical solution solves the technical problem of the slow response speed of voice interaction, that is, by calling the intent analysis program in parallel and querying the preset intent cache library, the intention of the user's voice instruction can be quickly and accurately analyzed, which effectively improves the accuracy and response speed of voice interaction instruction processing, improves the response speed of voice interaction, and improves the efficiency and accuracy of voice interaction. In addition, when the first intent data and the second intent data are inconsistent, the flexibility and intelligence of intent recognition are improved through compensation processing, combined with the user's historical behavior and the current intent recognition results, thereby improving the user experience.

[0103] In an exemplary embodiment, the query module is also used to: use an intent parsing program to perform natural language processing on the text data to obtain the first intent data; perform text parsing on the preprocessed text data to obtain entity words, input the entity words into an intent recognition model, and obtain intent categories output by the intent recognition model, wherein the intent recognition model is used to map the input words to predefined intent categories; and generate the first intent data according to a default intent framework corresponding to the entity words and the intent categories.

[0104] In an exemplary embodiment, the determination module is also used to: send the queried multiple second intention data to the target object, and receive the fourth intention data selected by the target object from the multiple second intention data; respectively calculate the first intention confidence between the first intention data and the fourth intention data, and the second intention confidence between the second intention data and the fourth intention data; and when it is determined that the first intention confidence meets the preset matching conditions, update the second intention data based on the first intention data to obtain the third intention data.

[0105] In an exemplary embodiment, the determination module is also used to: after determining the target intent data based on the third intent data, further, send the queried multiple second intent data to the target object, and receive the fourth intent data selected by the target object from the multiple second intent data; calculate the third intent confidence between the target intent data and the fourth intent data; when it is determined that the third intent confidence does not meet the preset matching conditions, generate a query statement based on the target intent data, and send the query statement to the target object, wherein the query statement is used to inquire the target object whether it recognizes the target intent data; determine whether to adjust the target intent data according to the target object's reply statement based on the query statement.

[0106] In an exemplary embodiment, the query module is also used to: before querying the second intent data corresponding to the text data from the preset intent cache, further, obtain historical high-frequency corpus with an interaction frequency higher than a preset frequency, and historical intent data corresponding to the historical high-frequency corpus from the historical interaction data; establish the preset intent cache according to the historical high-frequency corpus and the historical intent data corresponding to the historical high-frequency corpus; classify the historical high-frequency corpus according to the object behavior of the target object in the preset intent cache to obtain historical behavior corpora corresponding to different object behaviors; query the second intent data corresponding to the text data from the preset intent cache, including: determining the target behavior corpus belonging to the target object from the historical behavior corpora corresponding to the different object behaviors, and querying the second intent data corresponding to the text data from the historical intent data corresponding to the target behavior corpus.

[0107] In an exemplary embodiment, the query module is also used to: obtain a first text vector corresponding to the text data, and obtain a second text vector of the cache data in the preset intention cache library; use a word vector model to calculate the cosine similarity between the first text vector and the second text vector; obtain the intention cache data whose cosine similarity exceeds a preset threshold from the cache data; split the text data to obtain multiple keywords, and query the second intention data corresponding to the multiple keywords from the intention cache data.

[0108] In an exemplary embodiment, the determination module is also used to: determine the acquisition time of the first intention data and the query time of the second intention data; when it is determined that the acquisition time is earlier than the query time, determine the target intention data according to the first intention data; when it is determined that the acquisition time is later than the query time, determine the target intention data according to the second intention data.

[0109] An embodiment of the present application further provides a storage medium, which includes a stored program, wherein the program executes any of the above methods when it is run.

[0110] Optionally, in this embodiment, the storage medium may be configured to store program codes for executing the following steps:

[0111] S1, obtaining text data obtained after recognizing the voice interaction command of the target object;

[0112] S2, performing intent analysis on the text data to obtain first intent data, and querying second intent data corresponding to the text data from a preset intent cache library;

[0113] S3, when it is determined that the first intention data and the second intention data are inconsistent, performing compensation processing on the second intention data based on the first intention data to obtain third intention data, and determining target intention data according to the third intention data.

[0114] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0115] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0116] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:

[0117] S1, obtaining text data obtained after recognizing the voice interaction command of the target object;

[0118] S2, performing intent analysis on the text data to obtain first intent data, and querying second intent data corresponding to the text data from a preset intent cache library;

[0119] S3, when it is determined that the first intention data and the second intention data are inconsistent, performing compensation processing on the second intention data based on the first intention data to obtain third intention data, and determining target intention data according to the third intention data.

[0120] Optionally, in this embodiment, the above-mentioned storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store program codes.

[0121] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.

[0122] Obviously, those skilled in the art should understand that the above modules or steps of the present application can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, and optionally, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order from that herein, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.

[0123] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for improving voice interaction response speed, characterized in that: include: Acquire text data obtained after recognizing the voice interaction command of the target object; Performing intent parsing on the text data to obtain first intent data, and querying second intent data corresponding to the text data from a preset intent cache library; In the case where it is determined that the first intention data and the second intention data are inconsistent, compensation processing is performed on the second intention data based on the first intention data to obtain third intention data, and target intention data is determined based on the third intention data.

2. The method for improving voice interaction response speed according to claim 1, characterized in that: Performing intent parsing on the text data to obtain first intent data includes: Using an intention parsing program to perform natural language processing on the text data to obtain the first intention data; Performing text parsing on the preprocessed text data to obtain entity words, inputting the entity words into an intent recognition model, and obtaining intent categories output by the intent recognition model, wherein the intent recognition model is used to map the input words to predefined intent categories; The first intent data is generated according to a default intent framework corresponding to the entity word and the intent category.

3. The method for improving voice interaction response speed according to claim 1, characterized in that: Performing compensation processing on the second intention data based on the first intention data to obtain third intention data includes: Sending the queried multiple second intent data to the target object, and receiving fourth intent data selected by the target object from the multiple second intent data; respectively calculating a first intention confidence between the first intention data and the fourth intention data, and a second intention confidence between the second intention data and the fourth intention data; When it is determined that the first intention confidence satisfies a preset matching condition, the second intention data is updated based on the first intention data to obtain the third intention data.

4. The method for improving voice interaction response speed according to claim 1, characterized in that: After determining the target intention data according to the third intention data, the method further includes: Sending the queried multiple second intent data to the target object, and receiving fourth intent data selected by the target object from the multiple second intent data; Calculating a third intention confidence between the target intention data and the fourth intention data; In the case where it is determined that the third intention confidence does not meet the preset matching condition, generating a query statement based on the target intention data, and sending the query statement to the target object, wherein the query statement is used to inquire whether the target object recognizes the target intention data; Determine whether to adjust the target intention data based on the response statement of the target object based on the inquiry statement.

5. The method for improving voice interaction response speed according to claim 1, characterized in that: Before querying the second intent data corresponding to the text data from the preset intent cache library, the method further includes: acquiring historical high-frequency corpus with an interaction frequency higher than a preset frequency and historical intent data corresponding to the historical high-frequency corpus from historical interaction data; Establishing the preset intent cache library according to the historical high-frequency corpus and the historical intent data corresponding to the historical high-frequency corpus; Classifying the historical high-frequency corpus according to the object behavior of the target object in the preset intention cache library to obtain historical behavior corpus corresponding to different object behaviors; Querying the second intent data corresponding to the text data from the preset intent cache library includes: The target behavior corpus belonging to the target object is determined from the historical behavior corpus corresponding to the different object behaviors, and the second intention data corresponding to the text data is queried from the historical intention data corresponding to the target behavior corpus.

6. The method for improving voice interaction response speed according to claim 1, characterized in that: Querying the second intent data corresponding to the text data from the preset intent cache library includes: Obtaining a first text vector corresponding to the text data, and obtaining a second text vector of cache data in the preset intent cache library; Calculate the cosine similarity between the first text vector and the second text vector using a word vector model; Acquire the intention cache data whose cosine similarity exceeds a preset threshold from the cache data; The text data is split to obtain a plurality of keywords, and second intent data corresponding to the plurality of keywords is searched from the intent cache data.

7. The method for improving voice interaction response speed according to claim 1, characterized in that: The method further comprises: Determine a time when the first intention data is obtained and a time when the second intention data is queried; In a case where it is determined that the obtaining time is earlier than the querying time, determining the target intention data according to the first intention data; When it is determined that the obtaining time is later than the querying time, the target intention data is determined according to the second intention data.

8. A device for improving voice interaction response speed, characterized in that: include: The obtaining module is used to obtain text data obtained after the voice interaction instruction of the target object is recognized; A query module, used to perform intent analysis on the text data to obtain first intent data, and query second intent data corresponding to the text data from a preset intent cache library; A determination module is used to, when it is determined that the first intention data and the second intention data are inconsistent, perform compensation processing on the second intention data based on the first intention data to obtain third intention data, and determine target intention data based on the third intention data.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the method described in any one of claims 1 to 7 when executed.

10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 7 through the computer program.

Citation Information

Patent Citations

  • Data validity determination method and device, storage medium and electronic device

    CN112714020A

  • Control instruction intention recognition method, storage medium and electronic device

    CN116090461A

  • Voice instruction response method and device, storage medium and electronic device

    CN116312518A

  • Voice interaction method and device, electronic equipment and computer readable storage medium

    CN116403577A

  • Interaction response method and device based on intention recognition, equipment and storage medium

    CN116541493A