Voice interaction method, apparatus, device, medium and product
By acquiring multi-round interactive voice messages in a voice interaction system and judging their completeness, and identifying related voice messages to generate response content, the problem of low voice recognition accuracy is solved, thereby improving the response accuracy of the interaction system and the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- IFLYTEK CO LTD
- Filing Date
- 2024-09-11
- Publication Date
- 2026-04-21
AI Technical Summary
In existing voice interaction systems, errors in speech recognition or user speaking habits often lead to sentence segmentation errors, resulting in low speech recognition accuracy and hindering the subsequent interaction system from correctly responding to user needs.
By acquiring multi-round interactive voice recordings, the integrity of the interaction between voices is determined, related voices are identified, and response content is generated based on the related voices, thereby improving the accuracy of voice recognition.
It improves the accuracy of speech recognition and enhances the smoothness of the voice interaction process and the user experience.
Smart Images

Figure CN119181361B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice interaction technology, and in particular to a voice interaction method, apparatus, device, medium and product. Background Technology
[0002] Voice interaction technology is a field that has rapidly emerged in recent years with the development of artificial intelligence. It allows humans to communicate with computer systems through natural language. This technology has been widely used in various scenarios such as intelligent assistants, smart home devices, virtual customer service, and in-vehicle information systems.
[0003] In existing voice interaction systems, errors in speech recognition or user speaking habits often lead to grammatical errors, resulting in low speech recognition accuracy and consequently affecting the ability of subsequent interaction systems to respond correctly to user needs. Summary of the Invention
[0004] Based on the aforementioned technological status, this application proposes a voice interaction method, apparatus, device, medium, and product that can improve the accuracy of voice recognition, thereby improving the accuracy of responding to user needs.
[0005] To achieve the above-mentioned technical objectives, this application proposes the following technical solution:
[0006] The first aspect of this application proposes a voice interaction method, comprising: acquiring multi-round interactive voice, wherein the multi-round interactive voice includes current interactive voice and historical interactive voice; acquiring an interaction integrity judgment result between the current interactive voice and the historical interactive voice, and determining the associated voice of the current interactive voice from the historical interactive voice based on the interaction integrity judgment result; and generating response content based on the current interactive voice and the associated voice.
[0007] In some implementations, the historical interactive voice includes the previous round of interactive voice and the previous two rounds of interactive voice; obtaining the interaction integrity judgment result between the current round of interactive voice and the historical interactive voice, and determining the associated voice of the current round of interactive voice from the historical interactive voice based on the interaction integrity judgment result, includes: merging the current round of interactive voice and the previous round of interactive voice into a new round of interactive voice, and obtaining the interaction integrity judgment result between the multiple rounds of interactive voice and the new round of interactive voice; if the interaction integrity judgment result between the multiple rounds of interactive voice and the new round of interactive voice satisfies a first preset condition, then the previous round of interactive voice is determined as the associated voice of the current round of interactive voice; the first preset condition indicates that the new round of interactive voice is independent of the previous two rounds of interactive voice.
[0008] In some implementations, merging the current round of interactive voice and the previous round of interactive voice into a new round of interactive voice includes: obtaining the interaction integrity judgment result between the current round of interactive voice and the previous round of interactive voice; if the interaction integrity judgment result between the current round of interactive voice and the previous round of interactive voice meets a second preset condition, then merging the current round of interactive voice and the previous round of interactive voice into a new round of interactive voice.
[0009] In some implementations, the interaction integrity judgment result includes judgment results under multiple interaction integrity dimensions, including at least two of contextual relevance, interaction meaning, interaction domain intent, and interaction sequence relationship; the interaction integrity judgment result between the current round of interaction voice and the previous round of interaction voice satisfies the second preset condition, including: there is no contextual relevance, no interaction meaning, inconsistent interaction domain intent, and no interaction sequence relationship between the previous round of interaction voice and the current round of interaction voice.
[0010] In some implementations, the interaction integrity judgment result includes judgment results under multiple interaction integrity dimensions. These multiple interaction integrity dimensions include at least two of the following: contextual relevance, interaction meaning, interaction domain intent, and interaction sequence relationship. The judgment result under each interaction integrity dimension includes whether it meets the evaluation requirements or does not meet them. The first preset condition includes: for at least one of the multiple interaction integrity dimensions, the judgment results between the previous two rounds of interactive voice and the previous round of interactive voice, and between the previous round of interactive voice and the current round of interactive voice, are both non-compliant with the evaluation requirements under the at least one interaction integrity dimension; and the judgment results between the previous two rounds of interactive voice and the new current round of interactive voice are compliant with the evaluation requirements under the at least one interaction integrity dimension; or, the judgment results between the previous two rounds of interactive voice and the previous round of interactive voice are both compliant with the evaluation requirements under the multiple interaction integrity dimensions; the judgment results between the previous round of interactive voice and the current round of interactive voice are both non-compliant with the evaluation requirements under the multiple interaction integrity dimensions; and the judgment results between the previous two rounds of interactive voice and the new current round of interactive voice are both compliant with the evaluation requirements under the multiple interaction integrity dimensions.
[0011] In some implementations, the at least one interaction integrity dimension includes contextual relevance, interaction meaning, and interaction sequence; the judgment result under the at least one interaction integrity dimension is not in compliance with the evaluation requirements, including: no contextual relevance, no interaction meaning, and no interaction sequence; the judgment result under the at least one interaction integrity dimension is in compliance with the evaluation requirements, including: contextual relevance, interaction meaning, or interaction sequence; the judgment result under multiple interaction integrity dimensions is in compliance with the evaluation requirements, including: contextual relevance, interaction meaning, and interaction sequence; and the interaction domain intent between the first two rounds of interactive voice and the new current round of interactive voice is the same as the interaction domain intent between the first two rounds of interactive voice and the previous round of interactive voice. Figure 1 The results of the judgment under the multiple interaction integrity dimensions are all inconsistent with the evaluation requirements, including: no contextual relevance, no interactive meaning, no interactive sequence relationship, and the interaction domain intention between the previous round of interactive voice and the current round of interactive voice is inconsistent with the interaction domain intention between the previous two rounds of interactive voice and the previous round of interactive voice, as well as the interaction domain intention between the previous two rounds of interactive voice and the new current round of interactive voice.
[0012] In some implementations, the historical interactive voice includes the previous round of interactive voice, and the current round of interactive voice and the previous round of interactive voice are short-term continuous interactions. The short-term continuous interaction between the current round of interactive voice and the previous round of interactive voice is determined by the following steps: determining the interval between the current round of interactive voice and the previous round of interactive voice; if the interval is less than a preset duration, then the short-term continuous interaction between the current round of interactive voice and the previous round of interactive voice is determined.
[0013] In some implementations, the preset duration includes personalized preset durations for different users; wherein, the personalized preset duration for any user is determined based on the duration of silence when the user speaks continuously.
[0014] A second aspect of this application proposes a voice interaction device, comprising: an acquisition unit for acquiring multi-turn interactive voice, the multi-turn interactive voice including current interactive voice and historical interactive voice; a determination unit for acquiring judgment results of the current interactive voice and the historical interactive voice under multiple interaction integrity dimensions, and determining associated voice of the current interactive voice from the historical interactive voice based on the judgment results; the multiple interaction integrity dimensions include at least two of contextual relevance, interaction meaning, interaction domain intent, and interaction sequence relationship; and a response unit for generating response content based on the current interactive voice and the associated voice.
[0015] A third aspect of this application provides an electronic device, including a memory and a processor; the memory is connected to the processor and is used to store a program; the processor is used to implement the voice interaction method described in the first aspect and any of the implementations of the first aspect by running the program in the memory.
[0016] The fourth aspect of this application provides a storage medium storing a computer program, which, when executed by a processor, implements the voice interaction method described in the first aspect and any of its implementations.
[0017] The fifth aspect of this application provides a computer program product, including computer program instructions, which, when executed by a processor, cause the processor to implement the voice interaction method described in the first aspect and any of the implementations of the first aspect.
[0018] The voice interaction method, apparatus, device, medium, and product proposed in this application acquire multi-turn interactive voice, including the current round of interactive voice and historical interactive voice. It then obtains the interaction integrity judgment result between the current round of interactive voice and historical interactive voice, determines the related voice of the current round of interactive voice from the historical interactive voice based on the interaction integrity judgment result, and generates response content based on the current round of interactive voice and related voice. By evaluating the interaction integrity of the current round of interactive voice and historical interactive voice, the context of the current dialogue and its connection with historical interactions can be understood, thereby accurately understanding the user's intent and generating more accurate and personalized responses. This not only improves the accuracy of speech recognition itself but also enhances the fluency of the entire voice interaction process and the user experience. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0020] Figure 1 A flowchart illustrating a voice interaction method provided in an embodiment of this application;
[0021] Figure 2 A flowchart illustrating the judgment process for short-term continuous interaction provided in an embodiment of this application;
[0022] Figure 3 A schematic diagram of multi-turn interactive voice provided in an embodiment of this application;
[0023] Figure 4 A flowchart illustrating the process of determining the associated speech in the current round of interactive speech, provided for an embodiment of this application;
[0024] Figure 5 This is a schematic diagram of the structure of a voice interaction device provided in an embodiment of this application;
[0025] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0026] The technical solutions proposed in this application are applicable to any voice interaction application scenario, such as smart assistants, smart homes, in-vehicle information systems, virtual customer service, healthcare, education, office automation, public services, and other scenarios.
[0027] Smart assistants include smartphone assistants and smartwatches. In the context of smartphone assistants, users can use voice commands to make calls, send messages, check the weather, and set alarms. In the context of smartwatches, users can use the voice assistant on their smartwatch to receive notifications, control music playback, and perform health monitoring, among other things.
[0028] Smart homes encompass smart speakers, smart lighting systems, smart security systems, and smart appliances. In a smart speaker scenario, users can control music playback, access information, and set reminders via voice commands. In a smart lighting system scenario, users can turn lights on and off, and adjust brightness and color using voice commands. In a smart security system scenario, users can activate or deactivate surveillance cameras and alarm systems using voice commands. In a smart appliance scenario, users can control devices such as smart refrigerators, air conditioners, or televisions using voice commands.
[0029] In-vehicle information systems include navigation, entertainment, and vehicle control. In navigation systems, users can use voice commands to set destinations, query routes, and obtain real-time traffic information. In entertainment systems, users can use voice commands to play music, radio, audiobooks, etc. In vehicle control systems, users can use voice commands to control in-vehicle features such as windows, air conditioning, and seat heating.
[0030] Virtual customer service includes online customer service and call centers. In online customer service scenarios, virtual customer service representatives on a company's website or application can answer customer questions and provide product information through voice recognition and natural language processing technologies. In call center scenarios, users can handle customer calls through automated voice interaction systems, providing self-service options and reducing the workload of human customer service representatives.
[0031] Healthcare encompasses health consultations and elder care. In health consultations, patients can communicate with a smart health assistant via voice to obtain health advice or schedule doctor appointments. In elder care, smart devices can remind seniors to take medications, monitor their health status, and issue alerts in emergencies.
[0032] Education encompasses language learning and interactive teaching. In language learning scenarios, learners can improve their language skills by practicing conversations with voice assistants. In interactive teaching scenarios, students can interact with educational software via voice, asking questions or receiving feedback.
[0033] Office automation includes meeting minutes and task management. In meeting minutes, speech recognition technology automatically records meeting content and generates meeting summaries. In task management, voice commands are used to create task lists, set reminders, and schedule appointments.
[0034] Entertainment encompasses both gaming and virtual reality / augmented reality (VR / AR) scenarios. In gaming scenarios, players can interact with in-game characters and control gameplay via voice. In VR / AR scenarios, users can interact with the virtual environment through voice commands, enhancing immersion.
[0035] Public services include city navigation and public service inquiries. In the city navigation scenario, tourists can use voice commands to search for location information and obtain travel advice. In the public service scenario, users can use voice commands to inquire about government service information and book public services.
[0036] For the above application scenarios, the technical solution of the embodiments of this application can improve the accuracy of speech recognition, thereby improving the accuracy of response to user needs and enhancing the user's voice interaction experience.
[0037] The technical solutions provided in this application can be applied, by way of example, to hardware devices such as processors, electronic devices, and servers (including cloud servers), or packaged into software programs for execution. When the hardware device executes the processing procedure of the technical solutions in this application, or when the aforementioned software program is run, the target task can be automatically split and the application programming interfaces required by the task can be automatically invoked to achieve the purpose of the target task. This application only provides illustrative descriptions of the specific processing procedure of the technical solutions in this application and does not limit the specific implementation form of the technical solutions in this application. Any technical implementation form that can execute the processing procedure of the technical solutions in this application can be adopted by this application.
[0038] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0039] Before introducing the solution proposed in this application, the relevant technologies will first be introduced:
[0040] In voice interaction, the first step is to recognize the user's continuous voice requests. During this recognition phase, the content is typically segmented based on empirical values for Voice Activity Detection (VAD), generating several clauses. These clauses are then passed to the subsequent interaction system for processing to produce the appropriate response. However, due to errors in speech recognition or the influence of user speaking habits, sentence segmentation errors frequently occur. For example, due to user speaking habits, there are often pauses in the middle of a sentence. VAD might recognize the speech before and after this pause as two separate clauses, causing a single sentence to be incorrectly recognized as two or more clauses. This can lead to the subsequent interaction system failing to correctly understand the user's intent and thus failing to provide a proper response.
[0041] Such phrasing errors not only reduce the accuracy and efficiency of the system's response but also affect the user's interactive experience. For example, in a smart car navigation scenario, if a user repeatedly issues a voice request like "Navigate to Huangshan...Taiping Lake," the system might incorrectly interpret it as two clauses, "Navigate to Huangshan" and "Taiping Lake," because the user pauses after "Huangshan." This would lead to a failure to correctly understand the user's intent and thus an incorrect response. Similar problems occur in other application scenarios, such as smart assistants and in-vehicle information systems, where phrasing errors can cause interaction failures or incompleteness.
[0042] In view of this, this application proposes a voice interaction method, device, equipment, medium, and product, which determines the associated voice of the current round of interaction from the historical interaction voice by judging the interaction integrity between the historical interaction voice and the current round of interaction voice, and generates response content based on the current round of interaction voice and the associated voice, thereby improving the accuracy of voice recognition and thus improving the accuracy of responding to user needs.
[0043] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0044] This application first proposes a voice interaction method, see [link to previous document]. Figure 1 As shown, the method includes steps S101 to S103:
[0045] S101. Obtain multi-round interactive voice, which includes the current round of interactive voice and historical interactive voice.
[0046] Multi-turn interactive voice refers to a user's voice requests in a human-computer dialogue system, involving multiple consecutive rounds of dialogue between the user and the system. Users can continue to ask questions or provide supplementary information based on previous conversations. The system also needs to remember previous conversations and understand the user's intent based on the context to provide an accurate response. Multi-turn interaction more closely resembles natural human-to-human conversation. For example, in a smart assistant scenario, users can engage in multiple rounds of dialogue with the assistant, gradually refining their needs. A user might first ask about the weather, then ask for clothing suggestions based on the weather. Similarly, in a smart home scenario, users can control multiple devices in their home through multi-turn interaction; for instance, a user might first ask about the temperature of a room, then adjust the temperature based on the response.
[0047] In some embodiments, historical interactive voice includes the previous round of interactive voice, and the current round of interactive voice and the previous round of interactive voice are short-term continuous interactions. Short-term continuous interaction refers to the user speaking continuously within a preset duration. By comparing the interval between the previous round of interactive voice and the current round of interactive voice in a multi-round interactive voice interaction with the preset duration, it can be determined whether the current round of interactive voice and the previous round of interactive voice are short-term continuous interactions.
[0048] Figure 2 This is a flowchart illustrating the judgment process for short-term continuous interaction provided in an embodiment of this application. Figure 2 As shown, the process for determining short-term continuous interaction includes the following steps S201 and S202:
[0049] S201. Determine the interval between the current round of interactive voice and the previous round of interactive voice.
[0050] Figure 3 This is a schematic diagram of multi-turn interactive voice provided in an embodiment of this application. For example... Figure 3As shown, B represents the previous round of interactive voice, and t1 and t2 represent the start and end times of the previous round of interactive voice B, respectively. C represents the current round of interactive voice, and t3 and t4 represent the start and end times of the current round of interactive voice C, respectively.
[0051] The VAD algorithm can detect the start time t1 and end time t2 of the previous round of interactive speech B, and the start time t3 and end time t4 of the current round of interactive speech C. Specifically, assuming the user starts speaking, VAD performs voice activity detection. When the user's voice energy value is detected to be greater than or equal to a preset energy threshold, the current time is recorded as the start time t1; when the user's voice energy value is detected to be less than the preset energy threshold, the current time is recorded as the end time t2.
[0052] Similarly, when the user's voice energy value is detected again to be greater than or equal to the preset energy threshold, the current time is recorded as start time t3; when the user's voice energy value is detected again to be less than the preset energy threshold, the current time is recorded as t4. Between t2 and t3, if the user's voice energy value is continuously detected to be less than the preset energy threshold, the voice during this period is considered a silent segment.
[0053] Determining the interval between the current round of interactive voice and the historical interactive voice includes: determining the time difference between the start time of the current round of interactive voice and the end time of the previous round of interactive voice as the interval between the current round of interactive voice and the historical interactive voice.
[0054] After determining the preset duration, the interval duration can be compared with the preset duration, and based on the comparison result, it can be determined whether the current round of voice interaction is a short-term continuous interaction with the previous round of voice interaction. See step S202 below for details.
[0055] S202. If the interval between the current round of interactive voice and the historical interactive voice is less than the preset duration, it is determined that the current round of interactive voice and the previous round of interactive voice are short-term continuous interactions.
[0056] The preset duration can be set according to the user's speaking habits. For example, based on a summary of the speaking habits of a large number of users, the duration of pauses or prolongations in the middle of a complete sentence is usually greater than 0 and less than 400 milliseconds, 500 milliseconds, 600 milliseconds, 700 milliseconds, or 1 second. Therefore, the preset duration can be set to any time between 400 milliseconds and 1 second.
[0057] In some embodiments, if the interval between the current round of interactive voice and the previous round of interactive voice is longer than a preset duration, then it is determined that the current round of interactive voice and the previous round of interactive voice are non-short-term continuous interactions.
[0058] In some embodiments, the preset duration can also be personalized according to the speaking habits of different users, that is, the preset duration includes personalized preset durations corresponding to each user. The personalized preset duration for any user is determined based on the duration of silence when that user speaks continuously. Different users correspond to different personalized preset durations.
[0059] Specifically, in step S202, based on the user's identifier, the user's corresponding personalized preset duration can be queried from different pre-set personalized preset durations. The interval between the current round of interactive voice and the previous round of interactive voice is compared with the user's corresponding personalized preset duration. Based on the comparison result, it is determined whether the current round of interactive voice and the previous round of interactive voice constitute a short-term continuous interaction. Specifically, if the interval between the current round of interactive voice and the previous round of interactive voice is less than or equal to the user's corresponding personalized preset duration, the current round of interactive voice is determined to be a short-term continuous interaction with the previous round of interactive voice; if the interval between the current round of interactive voice and the previous round of interactive voice is greater than the user's corresponding personalized preset duration, the current round of interactive voice is determined to be a non-short-term continuous interaction with the previous round of interactive voice.
[0060] Assume user A's personalized preset duration is 500 milliseconds, and user B's personalized preset duration is 700 milliseconds. For example, if the interval between user A's current interaction and the previous interaction is 400 milliseconds, which is less than 500 milliseconds, it is determined to be a short-term continuous interaction. Conversely, if the interval between user B's current interaction and the previous interaction is 800 milliseconds, which is greater than 700 milliseconds, it is determined to be a non-short-term continuous interaction.
[0061] By setting personalized preset durations for different users, it is possible to more accurately identify users' speaking habits, thereby improving the recognition accuracy of short, continuous interactions. This method can better adapt to the speaking rhythms and habits of different users, enhancing the performance and user experience of the voice interaction system.
[0062] Continue reading Figure 1 The voice interaction method of this application may further include step S102 after step S101.
[0063] S102. Obtain the interaction integrity judgment result between the current round of interactive speech and the historical interactive speech, and determine the related speech of the current round of interactive speech from the historical interactive speech based on the interaction integrity judgment result.
[0064] The historical interactive voice includes the previous round of interactive voice and the previous two rounds of interactive voice. For example, suppose C represents the current round of interactive voice. B represents the previous round of interactive voice, and B' represents the response content of the previous round. A represents the previous two rounds of interactive voice, and A' represents the response content of the previous two rounds. Then the time sequence of multi-round interactive voice can be as follows: AA'->BB'->C.
[0065] For example, a multi-turn interactive voice AA'->BB'->C can be as follows:
[0066] A: What's the weather like in Jingxian these past few days?
[0067] A': The weather in Jingxian County these past two days has been [unclear].
[0068] B: Help me navigate there.
[0069] B': xxxxxx.
[0070] C: Kwun Tong Science and Technology Island.
[0071] Figure 4 This is a flowchart illustrating the process of determining the associated speech in the current round of interactive speech, provided as an embodiment of this application. For example... Figure 4 As shown, step S102 specifically includes steps S401 to S403:
[0072] S401. Merge the current round of interactive voice and the previous round of interactive voice into a new round of interactive voice.
[0073] There are several ways to implement step S401, as follows:
[0074] In some embodiments, step S401 includes the following steps a1 and a2:
[0075] Step a1: Obtain the interaction integrity judgment result between the current round of interactive voice and the previous round of interactive voice.
[0076] Specifically, a large model can be used to obtain the interaction integrity judgment result between the current round of interactive voice and the previous round of interactive voice. For example, the interaction integrity judgment result between BB'->C can be obtained.
[0077] After obtaining the interaction integrity judgment result between the current round of interactive voice and the previous round of interactive voice through step a1, it is possible to determine whether the current round of interactive voice and the previous round of interactive voice need to be merged into a new round of interactive voice based on the interaction integrity judgment result. Specifically, as shown in step b2.
[0078] Step a2: If the interaction integrity judgment result between the current round of interactive voice and the previous round of interactive voice meets the second preset condition, merge the current round of interactive voice and the previous round of interactive voice into a new current round of interactive voice.
[0079] The second presupposition condition indicates that the current round of interactive speech is not independent. In other words, there is a correlation between the current round of interactive speech and the previous round of interactive speech.
[0080] Steps a1 and a2 determine the likelihood of merging the current and previous rounds of interactive speech by assessing the interaction integrity between them. Specifically, if the interaction integrity assessment result between the current and previous rounds of interactive speech meets the second preset condition, the likelihood of merging is considered high, and the current and previous rounds of interactive speech are merged into a new current round of interactive speech; otherwise, the likelihood of merging is considered low, and the current and previous rounds of interactive speech are not merged.
[0081] For example, concatenating B and C into BC: "Help me navigate to Kwun Tong Tech Island." The judgment results include: Contextual relevance: whether BC discusses the same topic as AA' (e.g., Kwun Tong Tech Island); Interaction meaning: whether BC has a clear intention with AA' (e.g., navigating to Kwun Tong Tech Island); Interaction domain intent: whether BC belongs to the same domain as AA'; and Interaction continuity: whether BC is coherent with AA' (e.g., continuing the discussion of Kwun Tong Tech Island).
[0082] In some embodiments, the interaction integrity judgment result includes judgment results under multiple interaction integrity dimensions. Interaction integrity dimensions are evaluation dimensions used to assess whether the contexts of multi-turn interactive speech constitute a complete interaction. Multiple interaction integrity dimensions include at least two of the following: contextual relevance, interaction meaning, interaction domain intent, and interaction sequence relationship. The judgment results of the current round of interactive speech and the previous round of interactive speech under multiple interaction integrity dimensions satisfy a second preset condition, including: there is no contextual relevance, no interaction meaning, inconsistent interaction domain intent, and no interaction sequence relationship between the previous round of interactive speech and the current round of interactive speech.
[0083] Contextual relevance involves checking for logical continuity between the new round of interactive speech and the previous two rounds. Specifically, this can be checked using explicit pronouns, such as the presence of a pronoun from the previous round that corresponds to a question or answer. For example, if the previous round mentioned "Kwun Tong Tech Island," and the current round uses "there" to refer to "Kwun Tong Tech Island," then logical continuity exists. Alternatively, implicit referencing can be used to check for logical continuity, such as whether the context is discussing the same topic, like the same person, event, or object. For instance, if the previous two rounds discussed information related to "Kwun Tong Tech Island," and the current round continues the discussion around it, then logical continuity exists. Another approach is to check whether the current round of interactive speech completes, negates, or truncates content from the previous two rounds. For example, if the previous round of interactive voice asked a question, and the current round of interactive voice answered or supplemented it, this also indicates that there is a logical continuity.
[0084] Interaction meaning is determined by checking whether there is a clear intent between the new, current round of interaction and the previous two rounds. Specifically, this can be determined by checking whether the intent of the new, current round of interaction is consistent with that of the previous two rounds. For example, if the previous two rounds of interaction were asking for specific information, and the current round of interaction continues to ask for related information, it indicates that there is a clear intent.
[0085] For example, the voice prompts for the first two rounds of interaction are as follows:
[0086] A: What's the weather like in Jingxian County?
[0087] A': The weather in Jing County is xxxx.
[0088] The new interactive voice for this round is:
[0089] BC: What has been the recent temperature in Jingxian County?
[0090] In this example, the new round of interactive voice BC continues to inquire about weather information for Jingxian County, thus having a clear meaning compared to the previous two rounds of interactive voice. Figure 1 To the point of being responsive.
[0091] The meaning of an interaction can also be determined by checking whether the actions in the current round of interaction are consistent with those in the previous two rounds. If the previous two rounds of interaction were performing a certain operation, and the current round of interaction continues to perform the same operation, it indicates that there is a clear intention.
[0092] For example, the voice prompts for the first two rounds of interaction are as follows:
[0093] A: Please set my alarm for 7 a.m. tomorrow.
[0094] A': I have set your alarm for 7:00 AM tomorrow.
[0095] The new interactive voice for this round is:
[0096] BC: I have a meeting tomorrow morning, so I'll set an alarm for 8 a.m.
[0097] In this example, the new round of interactive voice BC continues to request setting an alarm, thus exhibiting clear consistency in action with the previous two rounds of interactive voice.
[0098] The domain intent of an interaction is to check whether the new, current-round interaction belongs to the same domain as the previous two rounds. This can be determined by checking whether the topic of the new, current-round interaction remains consistent with that of the previous two rounds. For example, if the previous two rounds of interaction discussed weather information, and the current round continues to discuss weather-related information, then it indicates that they belong to the same domain.
[0099] For example, the voice prompts for the first two rounds of interaction are as follows:
[0100] A: What's the weather like in Jingxian these past few days?
[0101] A': The weather in Jingxian County these past two days has been [unclear].
[0102] The new interactive voice for this round is:
[0103] BC: Will it rain tomorrow?
[0104] In this example, the new round of interactive voice BC continues to discuss weather information, and therefore belongs to the same domain as the previous two rounds of interactive voice.
[0105] Domain intent can also be determined by checking whether the entities (such as locations, people, and items) mentioned in the new round of interactive speech are consistent with those mentioned in the previous two rounds of interactive speech. For example, if the previous two rounds of interactive speech discussed "Kwun Tong Tech Island," and the current round of interactive speech continues to mention "Kwun Tong Tech Island," it indicates that they belong to the same domain.
[0106] For example, the voice prompts for the first two rounds of interaction are as follows:
[0107] A: Where is Kwun Tong Science and Technology Island?
[0108] A': Kwun Tong Science and Technology Island is located in xxx.
[0109] The new interactive voice for this round is:
[0110] BC: What are some fun places to visit on Kwun Tong Tech Island?
[0111] In this example, the new round of interactive voice BC continues to discuss information related to "Kwun Tong Tech Island", and therefore belongs to the same domain as the previous two rounds of interactive voice.
[0112] Interaction sequence is an examination of the coherence between the new, current round of interaction and the previous two rounds of interaction. Interaction sequence can be determined by checking whether the new, current round of interaction forms a coherent dialogue with the previous two rounds. For example, if the previous two rounds of interaction raised a question, and the current round provides a relevant answer, then coherence exists.
[0113] For example, the voice prompts for the first two rounds of interaction are as follows:
[0114] A: How do I get to Kwun Tong Science and Technology Island?
[0115] A': You can take MTR Line X to Kwun Tong Science Island.
[0116] The new interactive voice for this round is:
[0117] BC: When is the first train on subway line x?
[0118] In this example, the new round of interactive voice BC continues to ask for information related to traveling to Kwun Tong Tech Island, thus maintaining continuity with the previous two rounds of interactive voice.
[0119] The sequential relationship of interactions can also be determined by checking whether the new round of interactive speech completes the information from the previous two rounds. For example, if the previous two rounds of interactive speech provided some basic information, while the current round of interactive speech continues to provide more detailed information, it indicates that there is coherence.
[0120] For example, the voice prompts for the first two rounds of interaction are as follows:
[0121] A: What are some fun places to visit on Kwun Tong Tech Island?
[0122] A': Kwun Tong Science Island has many attractions, such as xxx.
[0123] The new interactive voice for this round is:
[0124] BC: What are the opening hours for these attractions?
[0125] In this example, the new round of interactive voice BC continues to ask for information about attractions on Kwun Tong Science Island, thus maintaining continuity with the previous two rounds of interactive voice.
[0126] In other words, step a2 includes: if there is no contextual relationship, no interactive meaning, inconsistent interactive domain intent, and no interactive sequence relationship between the current round of interactive speech and the previous round of interactive speech, then the current round of interactive speech is identified as non-independent interactive speech, and the current round of interactive speech and the previous round of interactive speech are merged into a new current round of interactive speech.
[0127] Specifically, steps a1 and a2 determine whether the current round of interactive voice is a non-independent interactive voice based on the interaction integrity judgment result between the current round of interactive voice and the previous round of interactive voice. If the current round of interactive voice is determined to be non-independent, the current round of interactive voice and the previous round of interactive voice are merged into a new current round of interactive voice. In this way, when the interaction integrity judgment result between the current round of interactive voice and the previous round of interactive voice does not meet the second preset condition, the response content can be directly generated based on the current round of interactive voice, thereby reducing the user's waiting time and improving the response efficiency.
[0128] In some embodiments, the current round of interactive speech and the previous round of interactive speech can also be directly merged into a new round of interactive speech. For example, the timing of the merged new multi-round interactive speech is: AA'->BC, specifically:
[0129] A: What's the weather like in Jingxian these past few days?
[0130] A': The weather in Jingxian County these past two days has been [unclear].
[0131] BC: Please give me directions to Kwun Tong Tech Island.
[0132] Continue reading Figure 4 After step S401, steps S402 and S403 may also be included.
[0133] S402. Obtain the interaction integrity judgment result between multi-round interactive voice and new current round interactive voice.
[0134] Specifically, a large model can be used to obtain the interaction integrity judgment results between multi-turn interactive speech and new current-turn interactive speech. For example, the interaction integrity judgment results of AA'->BB', BB'->C, and AA'->BC can be obtained.
[0135] In some embodiments, when obtaining the interaction integrity judgment result through a large model, a specific example can be as follows: Suppose you are an intelligent voice assistant, please combine the user's historical interaction voices to determine the contextual relevance, whether the interaction is meaningful, the interaction domain intent, and the interaction sequence relationship of two rounds of interaction respectively. Among them, the contextual relevance output is yes or no, the interaction is meaningful output is yes or no, the interaction domain intent output is one of weather, stocks, car control, food, travel, etc., if none of them match, the output is none, and the interaction sequence relationship output is sequential or not sequential.
[0136] S403. If the interaction integrity judgment result between the multi-round interactive voice and the new current round interactive voice meets the first preset condition, then the previous round interactive voice is determined as the associated voice of the current round interactive voice.
[0137] The first preset condition indicates that the new round of interactive speech is independent of the previous two rounds of interactive speech. In other words, if the interaction integrity judgment result between the multi-round interactive speech and the new round of interactive speech meets the first preset condition, it indicates that the new round of interactive speech is an independent interaction relative to the previous two rounds of interactive speech.
[0138] The interaction integrity assessment results between multi-turn interactive voice and the new current-turn interactive voice include assessment results under multiple interaction integrity dimensions. These multiple interaction integrity dimensions include at least two of the following: contextual relevance, interaction meaning, interaction domain intent, and interaction sequence relationship. The assessment results under each interaction integrity dimension include whether they meet the evaluation requirements or not. The first preset condition includes: for at least one of the multiple interaction integrity dimensions, the assessment results between the previous two rounds of interactive voice and the previous round of interactive voice, as well as between the previous round of interactive voice and the current round of interactive voice, are both non-compliant with the evaluation requirements under the at least one interaction integrity dimension, while the assessment results between the previous two rounds of interactive voice and the new current-turn interactive voice are compliant with the evaluation requirements under the at least one interaction integrity dimension; or, the assessment results between the previous two rounds of interactive voice and the previous round of interactive voice are both compliant with the evaluation requirements under the multiple interaction integrity dimensions, the assessment results between the previous round of interactive voice and the current round of interactive voice are both non-compliant with the evaluation requirements under the multiple interaction integrity dimensions, and the assessment results between the previous two rounds of interactive voice and the new current-turn interactive voice are both compliant with the evaluation requirements under the multiple interaction integrity dimensions.
[0139] The evaluation criteria are as follows: at least one dimension of interaction integrity includes contextual relevance, interaction meaning, and interaction sequence; if the evaluation result for at least one dimension is not met, it includes: no contextual relevance, no interaction meaning, and no interaction sequence; if the evaluation result for at least one dimension is met, it includes: contextual relevance, interaction meaning, or interaction sequence; if the evaluation result for multiple dimensions is met, it includes: contextual relevance, interaction meaning, and interaction sequence; and if the interaction domain intent between the first two rounds of interactive voice and the new current round of interactive voice is equal to the interaction domain intent between the first two rounds of interactive voice and the previous round of interactive voice, then the evaluation criteria are met. Figure 1 The results of the judgment under the multiple interaction integrity dimensions are all inconsistent with the evaluation requirements, including: no contextual relevance, no interactive meaning, no interactive sequence relationship, and the interaction domain intention between the previous round of interactive voice and the current round of interactive voice is inconsistent with the interaction domain intention between the previous two rounds of interactive voice and the previous round of interactive voice, as well as the interaction domain intention between the previous two rounds of interactive voice and the new current round of interactive voice.
[0140] In other words, the first presupposition condition includes at least one of the following four conditions: condition (1), condition (2), and condition (3), or condition (4):
[0141] (1) There is no contextual relationship between the first two rounds of interactive speech and the previous round of interactive speech, there is no contextual relationship between the previous round of interactive speech and the current round of interactive speech, and there is a contextual relationship between the first two rounds of interactive speech and the new current round of interactive speech.
[0142] (2) There is no interactive meaning between the first two rounds of interactive speech and the previous round of interactive speech, there is no interactive meaning between the previous round of interactive speech and the current round of interactive speech, and there is interactive meaning between the first two rounds of interactive speech and the new current round of interactive speech.
[0143] (3) There is no sequential relationship between the first two rounds of interactive speech and the previous round of interactive speech, there is no sequential relationship between the previous round of interactive speech and the current round of interactive speech, and there is a sequential relationship between the first two rounds of interactive speech and the new current round of interactive speech.
[0144] (4) There is contextual relevance, interactive meaning, and interactive sequence between the first two rounds of interactive speech and the previous round of interactive speech; there is no contextual relevance, interactive meaning, or interactive sequence between the previous round of interactive speech and the current round of interactive speech. Furthermore, there is contextual relevance, interactive meaning, and interactive sequence between the first two rounds of interactive speech and the new current round of interactive speech. The interactive domain intention between the first two rounds of interactive speech and the previous round of interactive speech is inconsistent with the interactive domain intention between the previous round of interactive speech and the current round of interactive speech, but not consistent with the interactive domain intention between the first two rounds of interactive speech and the new current round of interactive speech. Figure 1 To.
[0145] As mentioned earlier, the interaction between the first two rounds of voice and the previous round of voice can be represented as AA'->BB', the interaction between the previous round of voice and the current round of voice can be represented as BB'->C, and the interaction between the first two rounds of voice and the new current round of voice can be represented as AA'->BC. Therefore, the judgment results of multiple rounds of voice and the new current round of voice in multiple dimensions of interaction integrity include: the judgment results of AA'->BB', BB'->C, and AA'->C in multiple dimensions of interaction integrity. The above cases (1), (2), and (3) can be represented as shown in Table 1 below:
[0146] Table 1 First Preset Conditions
[0147]
[0148] For example, regarding the above situation (1), the following is an example:
[0149] A: Can you tell me what fun things there are to do in Huangshan?
[0150] A': Huangshan has........
[0151] B: The stinky mandarin fish is quite delicious.
[0152] B': Stinky mandarin fish........
[0153] C: Is it from over there?
[0154] For example, regarding the above situation (2), the following is an example:
[0155] A: What songs has celebrity A sung?
[0156] A': Celebrity A has sung ** songs, *** songs, and *** songs.
[0157] B: The second one.
[0158] B': The second *****
[0159] C: Which album is it from?
[0160] For example, regarding the above situation (3), the following is an example:
[0161] A: Do you know how many parks are nearby?
[0162] A': There are 3 parks nearby.
[0163] B, the first one, is pretty good.
[0164] B': The first *****.
[0165] C: Forget it, let's go to the second one.
[0166] The above situation (4) can be represented as shown in Table 2 below:
[0167] Table 2 First Preset Conditions
[0168]
[0169] In Tables 1 and 2, “F” indicates that the assessment requirements are not met, while “T” indicates that the assessment requirements are met.
[0170] Specifically, if the judgment results of the multi-turn interactive voice and the new current-turn interactive voice in multiple dimensions of interactive integrity satisfy one or more of the following conditions (1), (2) and (3), or satisfy condition (4), then it is considered that the judgment results of the multi-turn interactive voice and the new current-turn interactive voice in multiple dimensions of interactive integrity satisfy the first preset condition.
[0171] Continue reading Figure 1 After step S102, the voice interaction method of this application further includes step S103.
[0172] S103. Generate response content based on the current round of interactive voice and related voice.
[0173] After obtaining the associated voice from the current round of interaction, the new voice from the current round of interaction, obtained by splicing the current round of interaction voice and the associated voice, can be sent to the subsequent interaction system to generate accurate response content.
[0174] In some embodiments, new voice messages in the current round of interaction and their corresponding responses can also be stored in the interaction history stack.
[0175] Subsequently, when a new round of interactive voice is detected, it can be determined whether the new round of interactive voice and the aforementioned new round of interactive voice satisfy short-term continuous interaction; if they satisfy, then steps S101 to S103 are executed; if they do not satisfy, the detection continues until a new round of interactive voice and the aforementioned new round of interactive voice satisfy short-term continuous interaction, and then steps S101 to S103 are executed.
[0176] Information stored in the history stack can serve as part of long-term memory, helping the system quickly recall previous contexts when encountering similar problems in the future, thus providing more personalized services. Storing new current-round interaction voice messages and their corresponding responses in the interaction history stack helps the system better understand the user's context in future conversations, thereby improving its ability to understand user intent and the quality of its responses.
[0177] In summary, the technical solution proposed in this application, through evaluation across different dimensions, can help the system more accurately determine the context, meaning, and coherence of the current dialogue, thereby better understanding user intent and making appropriate responses. Therefore, it can improve the accuracy of recognizing the correlation between the current round of interactive speech and historical interactive speech, improve the accuracy of recognizing user intent, and thus improve the accuracy of speech recognition and the accuracy of voice interaction response.
[0178] Corresponding to the above-described voice interaction method, this application also proposes a voice interaction device, such as... Figure 5 As shown, the device includes: an acquisition unit 501, a determination unit 502, and a response unit 503; wherein, the acquisition unit 501 is used to acquire multi-turn interactive voice, which includes the current round of interactive voice and historical interactive voice; the determination unit 502 is used to acquire the interaction integrity judgment result between the current round of interactive voice and the historical interactive voice, and determine the associated voice of the current round of interactive voice from the historical interactive voice based on the interaction integrity judgment result; the response unit 503 is used to generate response content based on the current round of interactive voice and the associated voice.
[0179] In some embodiments, the historical interactive voice includes the previous round of interactive voice and the previous two rounds of interactive voice; the determining unit 502 obtains the interaction integrity judgment result between the current round of interactive voice and the historical interactive voice, and determines the associated voice of the current round of interactive voice from the historical interactive voice based on the interaction integrity judgment result, including: merging the current round of interactive voice and the previous round of interactive voice into a new current round of interactive voice, and obtaining the interaction integrity judgment result between the multiple rounds of interactive voice and the new current round of interactive voice; if the interaction integrity judgment result between the multiple rounds of interactive voice and the new current round of interactive voice satisfies a first preset condition, then the previous round of interactive voice is determined as the associated voice of the current round of interactive voice; the first preset condition indicates that the new current round of interactive voice is independent of the previous two rounds of interactive voice.
[0180] In some embodiments, when the determining unit 502 merges the current round of interactive voice and the previous round of interactive voice into a new round of interactive voice, it is specifically used to: obtain the interaction integrity judgment result between the current round of interactive voice and the previous round of interactive voice; if the interaction integrity judgment result between the current round of interactive voice and the previous round of interactive voice meets a second preset condition, then merge the current round of interactive voice and the previous round of interactive voice into a new round of interactive voice.
[0181] In some embodiments, the interaction integrity judgment result includes judgment results under multiple interaction integrity dimensions, and the multiple interaction integrity dimensions include at least two of contextual relevance, interaction meaning, interaction domain intent, and interaction sequence relationship; the interaction integrity judgment result between the current round of interaction voice and the previous round of interaction voice satisfies the second preset condition, including: there is no contextual relevance, no interaction meaning, inconsistent interaction domain intent, and no interaction sequence relationship between the previous round of interaction voice and the current round of interaction voice.
[0182] In some embodiments, the interaction integrity judgment result includes judgment results under multiple interaction integrity dimensions, wherein the multiple interaction integrity dimensions include at least two of contextual relevance, interaction meaning, interaction domain intent, and interaction sequence relationship; the judgment result under each interaction integrity dimension includes whether it meets the evaluation requirements or does not meet the evaluation requirements; the first preset condition includes: for at least one of the multiple interaction integrity dimensions, the judgment results between the previous two rounds of interactive voice and the previous round of interactive voice, and between the previous round of interactive voice and the current round of interactive voice, are both non-compliant with the evaluation requirements under the at least one interaction integrity dimension, and the judgment results between the previous two rounds of interactive voice and the new current round of interactive voice are compliant with the evaluation requirements under the at least one interaction integrity dimension; or, the judgment results between the previous two rounds of interactive voice and the previous round of interactive voice are both compliant with the evaluation requirements under the multiple interaction integrity dimensions, the judgment results between the previous round of interactive voice and the current round of interactive voice are both non-compliant with the evaluation requirements under the multiple interaction integrity dimensions, and the judgment results between the previous two rounds of interactive voice and the new current round of interactive voice are both compliant with the evaluation requirements under the multiple interaction integrity dimensions.
[0183] In some embodiments, the at least one interaction integrity dimension includes contextual relevance, interaction meaning, and interaction sequence; the judgment result under the at least one interaction integrity dimension is not in compliance with the evaluation requirements, including: no contextual relevance, no interaction meaning, and no interaction sequence; the judgment result under the at least one interaction integrity dimension is in compliance with the evaluation requirements, including: contextual relevance, interaction meaning, or interaction sequence; the judgment result under multiple interaction integrity dimensions is in compliance with the evaluation requirements, including: contextual relevance, interaction meaning, and interaction sequence; and the interaction domain intent between the first two rounds of interactive voice and the new current round of interactive voice is the same as the interaction domain intent between the first two rounds of interactive voice and the previous round of interactive voice. Figure 1 The results of the judgment under the multiple interaction integrity dimensions are all inconsistent with the evaluation requirements, including: no contextual relevance, no interactive meaning, no interactive sequence relationship, and the interaction domain intention between the previous round of interactive voice and the current round of interactive voice is inconsistent with the interaction domain intention between the previous two rounds of interactive voice and the previous round of interactive voice, as well as the interaction domain intention between the previous two rounds of interactive voice and the new current round of interactive voice.
[0184] In some embodiments, the historical interactive voice includes the previous round of interactive voice, and the current round of interactive voice and the previous round of interactive voice are short-term continuous interactions; the short-term continuous interactions between the current round of interactive voice and the previous round of interactive voice are determined by the following steps: determining the interval length between the current round of interactive voice and the previous round of interactive voice; if the interval length is less than a preset length, then the short-term continuous interactions between the current round of interactive voice and the previous round of interactive voice are determined.
[0185] In some embodiments, the preset duration includes personalized preset durations for different users; wherein, the personalized preset duration for any user is determined based on the duration of silence when the user speaks continuously.
[0186] The voice interaction device provided in this embodiment belongs to the same concept as the voice interaction method provided in the above embodiments of this application. It can execute the voice interaction method provided in any of the above embodiments of this application and has the corresponding functional modules and beneficial effects for executing the voice interaction method. Technical details not described in detail in this embodiment can be found in the specific processing content of the voice interaction method provided in the above embodiments of this application, and will not be repeated here.
[0187] The functions implemented by each unit in the above-mentioned voice interaction device can be implemented by the same or different processors, and this application embodiment does not limit this.
[0188] It should be understood that the units in the above-described voice interaction device can be implemented by a processor calling software. For example, the device includes a processor connected to a memory containing instructions. The processor calls the instructions stored in the memory to implement any of the above methods or to implement the functions of each unit in the device. The processor can be a general-purpose processor, such as a CPU or microprocessor, and the memory can be internal or external to the device. Alternatively, the units in the device can be implemented as hardware circuits. By designing the hardware circuits, some or all of the unit functions can be implemented. The hardware circuits can be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units are implemented by designing the logical relationships between the components within the circuit. In another implementation, the hardware circuit can be implemented using a PLD, such as an FPGA, which can include a large number of logic gates. The connection relationships between the logic gates are configured through a configuration file to implement the functions of some or all of the above units. All units in the above device can be implemented entirely by a processor calling software, entirely by hardware circuits, or partially by a processor calling software with the remaining parts implemented by hardware circuits.
[0189] In this application embodiment, a processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and execution capabilities, such as a CPU, microprocessor, GPU, or DSP. In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. These logical relationships are fixed or reconfigurable. For example, the processor may be a hardware circuit implemented as an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the processor loading instructions to implement the functions of some or all of the above units. Furthermore, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as an NPU, TPU, or DPU.
[0190] As can be seen, each unit in the above device can be one or more processors (or processing circuits) configured to implement the above methods, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.
[0191] Furthermore, the units in the above devices can be integrated in whole or in part, or they can be implemented independently. In one implementation, these units are integrated together and implemented in the form of a System-on-Chip (SoC). The SoC may include at least one processor for implementing any of the above methods or implementing the functions of the units in the device. The at least one processor may be of different types, such as CPU and FPGA, CPU and artificial intelligence processor, CPU and GPU, etc.
[0192] This application also proposes a control device, which includes a processor and an interface circuit. The processor in the control device is connected to a data input component through the interface circuit of the control device.
[0193] The data input component specifically refers to a functional component that can input or collect voice data, such as a microphone, etc.
[0194] The aforementioned interface circuit can be any interface circuit capable of implementing data communication functions, such as a USB interface circuit, a Type-C interface circuit, a serial port circuit, a PCIe circuit, etc.
[0195] The processor in this control device is also a circuit with signal processing capabilities, which executes the voice interaction method described in the above embodiments. For specific implementation details of the processor, please refer to the processor implementation methods described above; this application does not impose strict limitations on these implementations.
[0196] Another embodiment of this application also provides an electronic device, see [link to relevant documentation] Figure 6 As shown, the device includes:
[0197] Memory 200 and processor 210;
[0198] The memory 200 is connected to the processor 210 and is used to store programs;
[0199] The processor 210 is configured to implement the voice interaction method disclosed in any of the above embodiments by running the program stored in the memory 200.
[0200] Specifically, the aforementioned electronic device may also include: a bus, a communication interface 220, an input device 230, and an output device 240.
[0201] The processor 210, memory 200, communication interface 220, input device 230, and output device 240 are interconnected via a bus. Among them:
[0202] A bus can include a pathway for transmitting information between various components of a computer system.
[0203] The processor 210 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present invention. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0204] Processor 210 may include a main processor, as well as a baseband chip, modem, etc.
[0205] The memory 200 stores a program that executes the technical solution of this invention, and may also store an operating system and other key business functions. Specifically, the program may include program code, which includes computer operation instructions. More specifically, the memory 200 may include read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), other types of dynamic storage devices capable of storing information and instructions, disk storage, flash memory, etc.
[0206] Input device 230 may include a device for receiving user input data and information, such as a keyboard, mouse, camera, scanner, light pen, voice input device, touch screen, pedometer, or gravity sensor.
[0207] Output device 240 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.
[0208] The communication interface 220 may include a device that uses any transceiver to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.
[0209] The processor 210 executes the program stored in the memory 200 and calls other devices, which can be used to implement the various steps of any of the voice interaction methods provided in the above embodiments of this application.
[0210] This application also proposes a chip, which includes a processor and a data interface. The processor reads and runs a program stored in a memory through the data interface to execute the voice interaction method described in any of the above embodiments. For details of the processing and its beneficial effects, please refer to the above-described embodiments of the voice interaction method.
[0211] In addition to the methods and devices described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps of the voice interaction methods described in any of the above embodiments of this specification.
[0212] The computer program product can be written in any combination of one or more programming languages to perform the operations of the embodiments of this application. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0213] Furthermore, embodiments of this application may also be storage media storing a computer program, which is executed by a processor using the steps of the voice interaction method described in any of the above embodiments of this specification.
[0214] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0215] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0216] The steps in the methods of the various embodiments of this application can be adjusted, merged, or deleted in order according to actual needs, and the technical features described in each embodiment can be replaced or combined.
[0217] The modules and sub-modules in the various embodiments of the present application's devices and terminals can be merged, divided, and deleted according to actual needs.
[0218] It should be understood that the disclosed terminals, devices, and methods can be implemented in other ways, given the several embodiments provided in this application. For example, the terminal embodiments described above are merely illustrative. For instance, the division of modules or sub-modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple sub-modules or modules may be combined or integrated into another module, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.
[0219] The modules or submodules described as separate components may or may not be physically separate. The components that constitute a module or submodule may or may not be physical modules or submodules; that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules can be selected to achieve the purpose of this embodiment's solution, depending on actual needs.
[0220] Furthermore, the functional modules or sub-modules in the various embodiments of this application can be integrated into one processing module, or each module or sub-module can exist physically separately, or two or more modules or sub-modules can be integrated into one module. The integrated modules or sub-modules described above can be implemented in hardware or in the form of software functional modules or sub-modules.
[0221] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0222] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software unit executed by a processor, or a combination of both. The software unit can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0223] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0224] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A voice interaction method, characterized in that, include: Acquire multi-round interactive voice, which includes the current round of interactive voice and historical interactive voice, and the historical interactive voice includes the previous round of interactive voice and the previous two rounds of interactive voice. Obtain the interaction integrity judgment result between the current round of interactive voice and the historical interactive voice, and determine the associated voice of the current round of interactive voice from the historical interactive voice based on the interaction integrity judgment result; The process of obtaining the interaction integrity judgment result between the current round of interactive voice and the historical interactive voice, and determining the associated voice of the current round of interactive voice from the historical interactive voice based on the interaction integrity judgment result, includes: The current round of interactive voice and the previous round of interactive voice are merged into a new round of interactive voice, and the interaction integrity judgment result between the multi-round interactive voice and the new round of interactive voice is obtained. If the interaction integrity judgment result between the multi-round interactive voice and the new current round interactive voice meets the first preset condition, then the previous round interactive voice is determined as the associated voice of the current round interactive voice; the first preset condition indicates that the new current round interactive voice is independent of the previous two rounds of interactive voice.
2. The method of claim 1, wherein, Merging the current round of interactive voice and the previous round of interactive voice into a new round of interactive voice includes: Obtain the interaction integrity judgment result between the current round of interactive voice and the previous round of interactive voice; If the interaction integrity judgment result between the current round of interactive voice and the previous round of interactive voice meets the second preset condition, then the current round of interactive voice and the previous round of interactive voice will be merged into a new current round of interactive voice.
3. The method of claim 2, wherein, The interaction integrity judgment result includes judgment results under multiple interaction integrity dimensions, which include at least two of the following: contextual relevance, interaction meaning, interaction domain intent, and interaction sequence relationship; The interaction integrity judgment result between the current round of interactive voice and the previous round of interactive voice meets the second preset condition, including: there is no contextual association, no interactive meaning, inconsistent interaction domain intention and no interaction sequence relationship between the previous round of interactive voice and the current round of interactive voice.
4. The method of claim 1, wherein, The interaction integrity judgment result includes judgment results under multiple interaction integrity dimensions, which include at least two of the following: contextual relevance, interaction meaning, interaction domain intent, and interaction sequence relationship; The judgment results for each dimension of interaction integrity include whether the evaluation requirements are met or not. The first preset conditions include: For at least one of the multiple interaction integrity dimensions, the judgment results between the first two rounds of interactive voice and the previous round of interactive voice, as well as between the previous round of interactive voice and the current round of interactive voice, are all non-compliant with the evaluation requirements under the at least one interaction integrity dimension, while the judgment result between the first two rounds of interactive voice and the new current round of interactive voice is compliant with the evaluation requirements under the at least one interaction integrity dimension. or, The judgment results of the first two rounds of interactive voice and the previous round of interactive voice under the multiple interaction integrity dimensions all meet the evaluation requirements. The judgment results of the previous round of interactive voice and the current round of interactive voice under the multiple interaction integrity dimensions all fail to meet the evaluation requirements. Furthermore, the judgment results of the first two rounds of interactive voice and the new current round of interactive voice under the multiple interaction integrity dimensions all meet the evaluation requirements.
5. The method of claim 4, wherein, The at least one dimension of interaction integrity includes contextual relevance, interaction meaning, and interaction sequence; The judgment results under at least one dimension of interaction integrity are all unqualified for evaluation, including: no contextual relevance, no interactive meaning, and no sequential interaction relationship; The judgment result under at least one dimension of interaction integrity is considered to meet the evaluation requirements, including: the existence of contextual relevance, the existence of interactive meaning, or the existence of interactive sequence; The judgment results under the multiple interaction integrity dimensions all meet the evaluation requirements, including: there is contextual relevance, there is interactive meaning, there is interactive sequence, and the interaction domain intent between the first two rounds of interactive voice and the new current round of interactive voice is consistent with the interaction domain intent between the first two rounds of interactive voice and the previous round of interactive voice. The judgment results under the multiple interaction integrity dimensions all fail to meet the evaluation requirements, including: no contextual relevance, no interactive meaning, no interactive sequence relationship, and the interaction domain intent between the previous round of interactive voice and the current round of interactive voice is inconsistent with the interaction domain intent between the previous two rounds of interactive voice and the previous round of interactive voice, as well as the interaction domain intent between the previous two rounds of interactive voice and the new current round of interactive voice.
6. The method according to any one of claims 1 to 4, characterized in that, The historical interactive voice includes the previous round of interactive voice, and the current round of interactive voice and the previous round of interactive voice are short-term continuous interactions. The current round of interactive voice and the previous round of interactive voice are short-term continuous interactions, which are determined by the following steps: Determine the time interval between the current round of interactive voice and the previous round of interactive voice; If the interval is less than the preset duration, then the current round of interactive voice is determined to be a short-term continuous interaction with the previous round of interactive voice.
7. The method of claim 6, wherein, The preset duration includes personalized preset durations for each user. The personalized preset duration for any user is determined based on the duration of silence when the user speaks continuously.
8. A voice interaction device, characterized by include: The acquisition unit is used to acquire multi-round interactive voice, which includes the current round of interactive voice and historical interactive voice, and the historical interactive voice includes the previous round of interactive voice and the previous two rounds of interactive voice. The determining unit is used to obtain the judgment results of the current round of interactive voice and the historical interactive voice under multiple interaction integrity dimensions, and to determine the associated voice of the current round of interactive voice from the historical interactive voice based on the judgment results; the multiple interaction integrity dimensions include at least two of the following: contextual relevance, interaction meaning, interaction domain intent, and interaction sequence relationship; Specifically, when the determining unit obtains the interaction integrity judgment result between the current round of interactive voice and the historical interactive voice, and determines the associated voice of the current round of interactive voice from the historical interactive voice based on the interaction integrity judgment result, it includes: The current round of interactive speech and the previous round of interactive speech are merged into a new round of interactive speech, and the interaction integrity judgment result between the multi-round interactive speech and the new round of interactive speech is obtained; if the interaction integrity judgment result between the multi-round interactive speech and the new round of interactive speech meets a first preset condition, then the previous round of interactive speech is determined as the associated speech of the current round of interactive speech; the first preset condition indicates that the new round of interactive speech is independent of the previous two rounds of interactive speech; The response unit is used to generate response content based on the current round of interactive voice and the associated voice.
9. An electronic device, comprising: Including memory and processor; The memory is connected to the processor and is used to store programs; The processor is configured to implement the method as described in any one of claims 1 to 7 by running a program in the memory.
10. A storage medium, characterized by The storage medium stores a computer program, which, when executed by a processor, implements the method as described in any one of claims 1 to 7.
11. A computer program product, characterised in that, It includes computer program instructions that, when executed by a processor, cause the processor to perform the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-round interaction semantic understanding method and device, and computer storage medium
CN111429895A