Voice interaction method, device and apparatus for vehicle cabin
By configuring the current session cache and voice zone cache in the vehicle cabin and combining the historical semantic content of multiple voice zones, the contextual relevance problem during multi-user interaction in the in-vehicle voice interaction system is solved, improving the accuracy of voice intent recognition and user experience.
Patent Information
- Application Number
- CN202210963975.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-11
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2042-08-11
AI Technical Summary
Existing in-vehicle voice interaction systems cannot effectively handle the contextual relationships between voice signals from multiple voice zones, resulting in low accuracy of voice intent recognition during multi-user interaction and affecting user experience.
By configuring the current session cache and audio zone cache in the vehicle cabin and combining the historical semantic content of multiple audio zones, the current multi-turn session is updated to achieve contextual association links between different user sessions and avoid the separate processing of the voice intent of different users.
It improves the accuracy of multi-zone, multi-user interaction in the vehicle cabin, enhances the user experience, and achieves more accurate voice response and intent understanding.
Smart Images

Figure CN115240677B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of vehicles, and in particular to a voice interaction method for a vehicle cockpit, a voice interaction device for a vehicle cockpit, a computer device for a vehicle cockpit, a vehicle including the aforementioned voice interaction device or computer device, a storage medium, and a computer program product. Background Technology
[0002] In recent years, with the rapid development of artificial intelligence technology, voice interaction technology has also reached a practical level. With the continuous increase in the number of private cars, in-vehicle voice interaction systems have become standard equipment. At the same time, people increasingly desire to communicate naturally and conveniently with in-vehicle systems, thereby improving the interactive experience and driving safety. Therefore, further improving in-vehicle voice interaction systems is one of the important tasks in realizing intelligent voice interaction in vehicle cockpits.
[0003] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention
[0004] This disclosure provides a voice interaction method for a vehicle cockpit, a voice interaction device for a vehicle cockpit, a computer device, a vehicle including the above-described voice interaction device or voice interaction equipment, a storage medium, and a computer program product.
[0005] According to one aspect of this disclosure, a voice interaction method for a vehicle cockpit is provided. The vehicle cockpit includes multiple voice zones corresponding to multiple vehicle seats. The multiple voice zones are jointly configured with a current session cache for storing current multi-turn conversations, and each of the multiple voice zones is configured with a corresponding voice zone cache for storing historical semantic content from that voice zone within a preset time window. The method includes: obtaining current semantic content from a first voice zone among the multiple voice zones; updating the current multi-turn conversation based on the current session cache, the voice zone cache of the first voice zone, and the voice zone caches of the other voice zones among the multiple voice zones excluding the first voice zone, wherein the updated current multi-turn conversation is associated with the current semantic content; and processing the current semantic content associated with the updated current multi-turn conversation.
[0006] According to another aspect of this disclosure, a voice interaction device for a vehicle cockpit is provided. The vehicle cockpit includes multiple voice zones corresponding to multiple vehicle seats. The multiple voice zones are jointly configured with a current session cache for storing current multi-turn conversations, and each of the multiple voice zones is configured with a corresponding voice zone cache for storing historical semantic content from that voice zone within a preset time window. The device includes: an acquisition module configured to acquire current semantic content from a first voice zone among the multiple voice zones; an update module configured to update the current multi-turn conversation based on the current session cache, the voice zone cache of the first voice zone, and the voice zone caches of the other voice zones among the multiple voice zones excluding the first voice zone, wherein the updated current multi-turn conversation is associated with the current semantic content; and a processing module configured to process the current semantic content associated with the updated current multi-turn conversation.
[0007] According to another aspect of this disclosure, a computer device for a vehicle cockpit is provided, the vehicle cockpit including a plurality of audio zones corresponding to a plurality of vehicle seats, the plurality of audio zones being jointly configured with a current session cache for storing current multi-turn sessions, and each of the plurality of audio zones being configured with a corresponding audio zone cache for storing historical semantic content from that audio zone within a preset time window, the computer device including: at least one processor; and at least one memory storing a computer program thereon, the computer program causing the at least one processor to implement the above-described method when executed by the at least one processor.
[0008] According to another aspect of this disclosure, a vehicle is provided that includes the aforementioned voice interaction device or computer equipment.
[0009] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided for storing a computer program including instructions that, when executed by a processor, cause the processor to perform the methods described above.
[0010] According to another aspect of this disclosure, a computer program product is provided, the computer program product including instructions that, when executed by a processor, cause the processor to perform the methods described above.
[0011] According to embodiments of this disclosure, simultaneous interaction between multiple users in multiple voice zones and the vehicle's infotainment system can be achieved. Furthermore, it is possible to link different user sessions that have contextual relevance to avoid mapping the voices of different users as separate intentions, thereby improving the accuracy of voice interaction in the vehicle cabin and helping to improve the user's driving experience.
[0012] These and other aspects of this disclosure will be apparent from the embodiments described below, and will be elucidated with reference to the embodiments described below. Attached Figure Description
[0013] Further details, features, and advantages of this disclosure are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings. The accompanying drawings exemplarily illustrate embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements. In the drawings:
[0014] Figure 1 This is a schematic diagram illustrating an example system in which various methods described herein may be implemented according to exemplary embodiments;
[0015] Figure 2 This is a flowchart illustrating a voice interaction method for a vehicle cockpit according to an exemplary embodiment;
[0016] Figure 3 This is a flowchart illustrating a method for updating the current multi-turn session according to an exemplary embodiment;
[0017] Figure 4 This is a logic block diagram illustrating the voice interaction process in a vehicle cockpit according to an exemplary embodiment;
[0018] Figure 5 This is a block diagram illustrating a voice interaction device for a vehicle cockpit according to an exemplary embodiment; and
[0019] Figure 6 This is a block diagram illustrating an exemplary computer device that can be applied to an exemplary embodiment. Detailed Implementation
[0020] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.
[0021] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. As used herein, the term "multiple" means two or more, and the term "based on" should be interpreted as "at least partially based on". Furthermore, the terms "and / or" and "at least one of..." cover any one of the listed items and all possible combinations thereof.
[0022] Before introducing exemplary embodiments of this disclosure, several terms used herein will first be explained.
[0023] As used herein, the term "sound zone" refers to a spatial division equipped with a microphone or array of microphones for capturing user voice and / or ambient sound, and associated with seating positions within the vehicle cabin. The vehicle cabin space can be divided into sound zones according to actual needs. For example, the vehicle cabin space can be divided into two sound zones based on the front and rear seats (actual situations include, but are not limited to, vehicle cabin spaces with two rows of seats; for example, when a vehicle cabin has more than two rows of seats, the vehicle cabin space can be divided into more than two sound zones accordingly), or the vehicle cabin space can be divided into different sound zones based on the driver's seat, front passenger seat, and passenger seats, and so on.
[0024] As used herein, the term "sound zone buffer" refers to a storage device corresponding to a sound zone that stores user speech and / or ambient sound or its processed signals acquired by the microphone or microphone array provided with that sound zone. Unless otherwise specified, when referring to a sound zone buffer, it may refer to the objects stored in the sound zone buffer, such as, but not limited to, semantic content from that sound zone that has been processed by Natural Language Understanding (NLU).
[0025] As used herein, the term "current session cache" refers to a storage device associated with a current session of (multiple) users within the vehicle cabin space, storing user speech and / or ambient sounds, or their processed signals, captured by microphones or microphone arrays in (multiple) audio zones involved in that session. As an example, in actual signal processing, the objects stored in the current session cache may include a chronological cascade of user speech and / or ambient sounds, or their processed signals, captured by microphones or microphone arrays in individual audio zones that constitute the current session. Unless otherwise specified, reference to the current session cache may refer to objects stored in the current session cache, such as, but not limited to, current multi-turn sessions processed by Natural Language Understanding (NLU).
[0026] As used herein, the term "semantic content" refers to the processed signal obtained after the original speech signal has been processed by Natural Language Understanding (NLU) and contains rich details such as contextual information (e.g., whether there is a connection with other speech), conversation state (e.g., whether it is the beginning or end of a dialogue), and intent (e.g., whether it includes a query for specific information).
[0027] As used herein, the term "multi-turn conversation" refers to a conversation that includes more than one voice signal that makes up the conversation. It is usually completed by multiple users, but the possibility of a single user conducting multiple turns of conversation is not excluded. For example, when a single user speaks a voice statement with a query purpose, he or she may then speak another voice statement that supplements the aforementioned query.
[0028] From traditional question-and-answer voice interaction systems to the more popular multi-turn question-and-answer voice interaction systems, human-computer interaction is trending towards human-to-human interaction. Voice interaction typically involves speech recognition, natural language understanding, and speech synthesis. In smart cockpits, the speaker is often not unique. To address this, related technologies include in-vehicle voice interaction systems that enable simultaneous interaction among multiple passengers. These systems achieve this by creating multiple voice interaction links for multiple sound zones within the cockpit (i.e., dividing the cockpit into multiple sound zones (seats) and collecting voice signals from different areas via microphone arrays). When one or more voice interaction links detect a sound signal from a sound zone, the vehicle terminal switches that link to voice processing mode. This voice processing mode is used to process the voice signal input from the passenger in the corresponding sound zone. Although this method supports switching between conversation objects, unfortunately, only one voice interaction link is in voice processing mode at a time, so this method is still a serial, single-conversation interaction. In addition, among the related technologies, there is a method that sends signals from multiple audio zones to a preset cloud server for semantic recognition in order to generate corresponding voice commands to achieve in-vehicle voice interaction. However, this method cannot determine whether there is a contextual relationship between the voice signals from multiple audio zones.
[0029] To address the aforementioned technical problems, a novel voice interaction method for vehicle cockpits is proposed according to one or more embodiments of this disclosure. This method updates the current multi-turn conversation based on the current session cache, the voice region cache of a first voice region, and the voice region caches of the remaining voice regions (excluding the first voice region) among multiple voice regions. This allows for better handling of the current semantic content associated with the updated current multi-turn conversation while taking into account the contextual relevance of the multi-turn conversation. Through this method, simultaneous interaction between multiple users in multiple voice regions and the vehicle's infotainment system can be achieved. Furthermore, it enables linking different user conversations with contextual relevance to avoid mapping different users' voices as separate intentions, thereby improving the accuracy of voice interaction in the vehicle cockpit and contributing to a better user experience. Exemplary embodiments of this disclosure are described in detail below with reference to the accompanying drawings.
[0030] Figure 1 This is a schematic diagram illustrating an example system 100 in which various methods described herein may be implemented according to exemplary embodiments.
[0031] refer to Figure 1 The system 100 includes an in-vehicle system 110, a server 120, and a network 130 that communicatively couples the in-vehicle system 110 and the server 120.
[0032] The in-vehicle system 110 includes a display 114 and an application (APP) 112 that can be displayed via the display 114. The application 112 can be an application that is pre-installed on the in-vehicle system 110 or downloaded and installed by the user 102, or it can be a lightweight applet. When the application 112 is an applet, the user 102 can run the application 112 directly on the in-vehicle system 110 without installing it, by searching for the application 112 in the host application (e.g., by the name of the application 112) or by scanning the graphic code of the application 112 (e.g., a barcode, QR code, etc.). In some embodiments, the in-vehicle system 110 may include one or more processors and one or more memories (not shown), and the in-vehicle system 110 is implemented as an in-vehicle computer. In some embodiments, the in-vehicle system 110 may include more or fewer displays 114 (e.g., no display 114), and / or one or more speakers or other human-machine interaction devices. In some embodiments, the in-vehicle system 110 may not communicate with the server 120.
[0033] Server 120 can represent a single server, a cluster of multiple servers, a distributed system, or a cloud server providing basic cloud services (such as cloud databases, cloud computing, cloud storage, and cloud communication). It will be understood that, although... Figure 1The diagram shows that server 120 communicates with only one vehicle system 110, but server 120 can provide background services for multiple vehicle systems simultaneously.
[0034] Network 130 allows wireless communication and information exchange between vehicles ("X" meaning vehicles, roads, pedestrians, or the Internet, etc.) according to agreed communication protocols and data exchange standards. Examples of network 130 include combinations of local area networks (LANs), wide area networks (WANs), personal area networks (PANs), and / or communication networks such as the Internet. Network 130 can be wired or wireless. In one example, network 130 can be an in-vehicle network, a vehicle-to-vehicle network, and / or an in-vehicle mobile Internet.
[0035] For the purposes of this disclosure's embodiments, Figure 1 In the example, application 112 can be an electronic map application that provides various functions based on electronic maps, such as navigation, route query, location search, etc. Correspondingly, server 120 can be a server used in conjunction with the electronic map application. Server 120 can provide online map services, such as online navigation, online route query, and online location search, to application 112 running in vehicle system 110 based on road network data. Alternatively, server 120 can also provide road network data to vehicle system 110, allowing application 112 running in vehicle system 110 to provide local map services based on this road network data.
[0036] Figure 2 This is a flowchart illustrating a voice interaction method 200 for a vehicle cockpit according to an exemplary embodiment. Method 200 can be implemented in in-vehicle systems (e.g., Figure 1 The execution is performed at the vehicle system 110 shown, that is, the execution entity of each step of method 200 can be... Figure 1 The vehicle-mounted system 110 shown. In some embodiments, method 200 can be performed on a server (e.g., Figure 1 The method 200 is executed at server 120 (as shown in the diagram). In some embodiments, method 200 may be executed in combination by an in-vehicle system (e.g., in-vehicle system 110) and a server (e.g., server 120). Hereinafter, the steps of method 200 will be described using in-vehicle system 110 as an example. Here, the vehicle cabin where in-vehicle system 110 is located includes multiple audio zones corresponding to multiple vehicle seats. These multiple audio zones are jointly configured with a current session cache to store current multi-turn sessions, and each of these multiple audio zones is configured with a corresponding audio zone cache to store historical semantic content from that audio zone within a preset time window.
[0037] like Figure 2 As shown, method 200 includes:
[0038] Step S210: Obtain the current semantic content of the first vocal region from multiple vocal regions;
[0039] Step S220: Based on the current session cache, the register cache of the first register, and the register caches of the other registers besides the first register, update the current multi-turn session, wherein the updated current multi-turn session is associated with the current semantic content; and
[0040] Step S230: Process the current semantic content associated with the updated current multi-turn session.
[0041] The steps of method 200 are described in detail below.
[0042] In step S210, the current semantic content of a first audio region from multiple audio regions is obtained. The first audio region can be any one of multiple audio regions within the vehicle cabin that correspond to multiple vehicle seats. For example, the first audio region can be the audio region corresponding to the driver's seat or the audio region corresponding to the front passenger seat and / or rear passenger seats; this disclosure does not impose any limitations in this regard. The current semantic content can be a processed signal obtained by processing the current raw speech signal from the first audio region using Natural Language Understanding (NLU). In some examples, the current raw speech signal from the first audio region can be acquired by a microphone or microphone array equipped with the first audio region.
[0043] In step S220, the current multi-turn session is updated based on the current session cache, the first voice zone's voice zone cache, and the voice zone caches of the other voice zones besides the first voice zone. Further, the updated current multi-turn session is associated with the current semantic content. The content stored in the current session cache is associated with the current sessions of (multiple) users within the entire vehicle cabin space; the content stored in the first voice zone's voice zone cache is associated with the user's currently spoken speech or historical speech within the first voice zone; and the content stored in the voice zone caches of the other voice zones besides the first voice zone is associated with the user's currently spoken speech or historical speech within the corresponding voice zone. This avoids the NLU crosstalk problem (i.e., the intents included in multiple sessions within multiple voice zones may be merged and understood as the same intent by the NLU) and links different user sessions with contextual relationships to avoid simply mapping the voices of different users into separate intents.
[0044] In step S230, the current semantic content associated with the updated current multi-turn conversation is processed. Since the current multi-turn conversation has been updated in step S220, the current semantic content associated with the updated current multi-turn conversation can reflect the context information of the current multi-turn conversation in which it is located. As a result, the vehicle system can provide more targeted responses to the current semantic content, thereby improving the accuracy of human-machine voice interaction in the vehicle cabin.
[0045] According to embodiments of this disclosure, the method 200 overcomes the shortcomings of in-vehicle voice interaction systems in the related art, which respond to multiple voice inputs in a serial single-session interaction manner and cannot truly respond to multi-voice zone multi-session intents in parallel, as well as the inability to determine whether there is a contextual relationship between the voice signals of multiple voice zones. Method 200 updates the current multi-turn conversation based on the current conversation cache, the voice zone cache of the first voice zone, and the voice zone caches of the remaining voice zones other than the first voice zone. By taking into account the contextual association of the multi-turn conversation, it can better process the current semantic content associated with the updated current multi-turn conversation, thereby better understanding and responding to the user's intent and improving the user's driving experience.
[0046] Figure 3 This is a flowchart illustrating a method 300 for updating the current multi-turn session according to an exemplary embodiment. Method 300 can be implemented in an in-vehicle system (e.g., Figure 1 The execution is performed at the vehicle system 110 shown, that is, the execution entity of each step of method 300 can be... Figure 1 The vehicle-mounted system 110 shown. In some embodiments, method 300 can be performed on a server (e.g., Figure 1 The method 300 is executed at server 120 (as shown in the diagram). In some embodiments, method 300 may be executed in combination by an in-vehicle system (e.g., in-vehicle system 110) and a server (e.g., server 120). Hereinafter, the steps of method 300 will be described using in-vehicle system 110 as an example. Here, the vehicle cabin where in-vehicle system 110 is located includes multiple audio zones corresponding to multiple vehicle seats. These multiple audio zones are jointly configured with a current session cache to store current multi-turn sessions, and each of these multiple audio zones is configured with a corresponding audio zone cache to store historical semantic content from that audio zone within a preset time window.
[0047] like Figure 3 As shown, method 300 includes:
[0048] Step S310: Determine whether historical semantic content is stored in the register buffer of the first register.
[0049] Step S320: In response to determining that the voice region cache of the first voice region stores historical semantic content, determine whether the current semantic content and the historical semantic content belong to the same multi-turn session; and
[0050] In step S330, in response to determining that the current semantic content and the historical semantic content belong to the same multi-turn session, the current semantic content and the historical semantic content are stored together in the current session cache to update the current multi-turn session.
[0051] The following is combined with Figure 4 Let me describe in detail each step of method 300.
[0052] Figure 4 This is a logical block diagram illustrating a voice interaction process 400 in a vehicle cockpit according to an exemplary embodiment. The workflow of process 400 may include acquiring a new sound source signal from one of the multiple sound zones using a microphone or microphone array, processing the sound source signal, and performing semantic content parsing of the sound source signal using Automatic Speech Recognition (ASR) and / or NLU. In a multi-turn session, when multiple sound source signals (which may be accompanied by intent understanding requests) have been received, ASR and / or NLU combine the cache state of the sound region, the cache state of the other sound regions, and the current session cache corresponding to the vehicle cockpit to determine whether the newly received sound source signal belongs to the same multi-turn session as one or more of the multiple sound source signals that have been received. Based on this, the intent understanding at this moment is given, and the current request corresponding to the newly received sound source signal (e.g., new semantic content) is processed. The current session cache and sound region cache are updated. If the newly received sound source signal does not belong to any of the multiple sound source signals that have been received, the newly received sound source signal is newly created as the current multi-turn session in the current session cache (e.g., to replace the content stored in the current session cache) and the sound region cache is updated.
[0053] The following describes each box of process 400 in detail.
[0054] In box 401, wait to receive new semantic content. The semantic content can be the processed signal obtained after the current raw speech signal from the audio region has been processed by Natural Language Understanding (NLU). In some examples, the current raw speech signal from the audio region can be captured by the microphone or microphone array provided with the audio region.
[0055] In box 402, new semantic content for sound zone i is received. Sound zone i can be any one of multiple sound zones that correspond to multiple vehicle seats within the vehicle cabin. For example, sound zone i can be the sound zone corresponding to the driver's seat or the sound zone corresponding to the front passenger seat and / or the rear passenger seat; this disclosure does not impose any limitations on this.
[0056] In box 403, it is determined whether historical semantic content is stored in the cache of voice region i. Box 403 can correspond to Figure 3 Step S310: Determine whether the sound buffer of the first sound region stores historical semantic content. At this time, sound region i corresponds to the first sound region. Historical semantic content can refer to the processed signal obtained by natural language understanding (NLU) after the historical speech signal is processed, which contains rich details such as contextual information (e.g., whether there is a continuation with other speech), conversation state (e.g., whether it is the beginning or end of a dialogue), and intent (e.g., whether it includes a query for specific information).
[0057] According to embodiments of this disclosure, the query time window length of the audio zone cache can be customized.
[0058] When the judgment result of box 403 is yes, process 400 proceeds to box 404.
[0059] In box 404, it is determined whether the new semantic content belongs to the same multi-turn session as the historical semantic content in the cache of voice region i. Box 404 can correspond to Figure 3 Step S320: In response to determining that the audio buffer of the first audio region stores historical semantic content, determine whether the current semantic content and the historical semantic content belong to the same multi-turn session. It should be noted that the methods used to determine whether two semantic contents belong to the same multi-turn session may include speech processing technologies known in the art, such as ASR and NLU. To avoid obscuring the inventive concept of this disclosure, they will not be described in detail here.
[0060] When the judgment result of box 404 is yes (that is, the semantic content newly received in sound zone i has a contextual relationship with the historical semantic content previously received in sound zone i, so that the new semantic content in sound zone i and the historical semantic content constitute the same multi-turn conversation), process 400 proceeds to box 405.
[0061] In box 405, update the current session cache. Specifically, store the newly received semantic content along with its associated historical semantic content in the current session cache to overwrite the content stored in the current session cache. Box 405 may correspond to Figure 3Step S330: In response to determining that the current semantic content and the historical semantic content belong to the same multi-turn session, the current semantic content and the historical semantic content are stored together in the current session cache to update the current multi-turn session. Similar to the query time window of the audio region cache, a query time window length that can be customized can be set for the current session cache. It should be noted here that if the query time window length of the current session cache is insufficient to cover both the current semantic content and the historical semantic content to be overwritten in the current session cache, the earlier part of the historical semantic content may not be stored in the current session cache. That is, the content stored in the current session cache is the closest to the current time. At this point, method 300 is complete.
[0062] Continue to refer to Figure 4 Other exemplary embodiments of this disclosure are described.
[0063] Continuing with the above exemplary embodiment, when process 400 proceeds from box 403 to box 404 and the judgment result at box 404 is negative (that is, the semantic content newly received in sound zone i has no contextual relationship with the historical semantic content previously received in sound zone i, so that the new semantic content in sound zone i and the historical semantic content in sound zone i cannot constitute the same multi-turn conversation), process 400 can proceed from box 404 to box 407.
[0064] In box 407, it is determined whether the new semantic content belongs to the same multi-turn session as the current multi-turn session in the current session cache. The content stored in the current session cache is associated with the current sessions of (multiple) users within the entire vehicle cabin space. For example, if the new semantic content is "need parking space", and the current multi-turn sessions stored in the current session cache include "buy bags" and "nearest shopping mall", then it can be determined that the new semantic content and the current multi-turn sessions stored in the current session cache belong to the same multi-turn session. The intent of this multi-turn session can be understood as wanting to go to the nearest shopping mall with parking spaces to buy bags, and the content to be written to the current session cache is "need parking space", "buy bags", and "nearest shopping mall". Therefore, when the determination result of box 407 is yes, process 400 can proceed from box 407 to box 405.
[0065] According to embodiments of this disclosure, the above method may optionally include an additional step: in response to determining that the current semantic content and the historical semantic content do not belong to the same multi-turn session, determining whether the current semantic content and the current multi-turn session belong to the same multi-turn session; and in response to determining that the current semantic content and the current multi-turn session belong to the same multi-turn session, storing the current semantic content and the current multi-turn session together in the current session cache to update the current multi-turn session.
[0066] Continuing with the exemplary embodiment described above, when process 400 proceeds through blocks 403 and 404 to block 407 and the determination result at block 407 is negative, process 400 can proceed from block 407 to block 408. In block 408, it is determined whether the new semantic content and the historical semantic content in the caches of the remaining sound regions j (excluding sound region i) belong to the same multi-turn session. In some cases, the new semantic content currently received in sound region i may not be associated with either the historical semantic content of sound region i or the current multi-turn session stored in the current session cache. For example, if the new semantic content of sound region i is "traffic restriction tail number," the historical semantic content of sound region i is "today's temperature," and the current multi-turn sessions stored in the current session cache include "buy bags" and "nearest shopping mall," then it is impossible to determine the association between the new semantic content currently received in sound region i and the historical semantic content of sound region i, nor is it possible to determine the association between the new semantic content currently received in sound region i and the current multi-turn session stored in the current session cache. Therefore, it is necessary to determine whether the new semantic content in voice region i and the historical semantic content in the cache of the other voice regions j (excluding voice region i) belong to the same multi-turn session.
[0067] When the new semantic content in voice region i and the historical semantic content in the cache of voice region j (excluding voice region i) belong to the same multi-turn session (for example, the historical semantic content in the cache of voice region j is "what day of the week tomorrow", so the intent of the multi-turn session formed by the two can be understood as "what is the number of roads with unlimited traffic and what is the last digit of the traffic restriction number tomorrow"), process 400 can proceed from box 408 to box 405. In box 405, the new semantic content in voice region i and the historical semantic content in the cache of voice region j are stored together in the current session cache to update the content stored in the current session cache.
[0068] According to embodiments of this disclosure, the above method may optionally include an additional step: in response to determining that the current semantic content and the current multi-turn session do not belong to the same multi-turn session, determining whether the current semantic content and the historical semantic content stored in the region caches of the other regions (excluding the first region) belong to the same multi-turn session; and in response to determining that the current semantic content and the historical semantic content stored in the region cache of the second region belong to the same multi-turn session, storing the current semantic content and the historical semantic content stored in the region cache of the second region together in the current session cache to update the current multi-turn session. Here, the second region is a region different from the first region.
[0069] Continuing with the above exemplary embodiment, when process 400 proceeds through blocks 403, 404, and 407 to block 408 and the determination result at block 408 is negative, process 400 may proceed from block 408 to block 409. When the new semantic content in audio region i does not belong to the same multi-turn session as the historical semantic content in the cache of the other audio regions j (excluding audio region i) (in addition, as mentioned above, the new semantic content in audio region i is not associated with the historical semantic content stored in the cache of audio region i or the current multi-turn session stored in the current session cache of the vehicle cockpit), process 400 may proceed from block 408 to block 409, where the new semantic content in audio region i replaces the current multi-turn session in the current session cache.
[0070] According to embodiments of this disclosure, the above method may optionally include an additional step: in response to determining that the current semantic content and the historical semantic content stored in the audio region caches of the other audio regions (excluding the first audio region) do not belong to the same multi-turn session, storing the current semantic content in the current session cache to update the current multi-turn session.
[0071] Returning to box 403, if the result of the judgment in box 403 is negative, process 400 can proceed from box 403 to box 407. Further, if the result of the judgment in box 407 is positive, process 400 can then proceed from box 407 to box 405. This indicates that the new semantic content received in voice region i may be the initial semantic content of that voice region (i.e., the voice region cache of voice region i was previously empty) and is associated with the current multi-turn session stored in the current session cache (e.g., it may consist of semantic content from voice regions other than voice region i).
[0072] According to embodiments of this disclosure, the above method may optionally include an additional step: in response to determining that no historical semantic content is stored in the audio region cache of the first audio region, determining whether the current semantic content and the current multi-turn session belong to the same multi-turn session; and in response to determining that the current semantic content and the current multi-turn session belong to the same multi-turn session, storing the current semantic content and the current multi-turn session together in the current session cache to update the current multi-turn session.
[0073] Continuing with the exemplary embodiment described above, when process 400 proceeds through blocks 403 and 407 to block 408 and the judgment result at block 408 is yes, process 400 can proceed from block 408 to block 405. For example, this situation could be: the new semantic content of voice region i is "traffic restriction tail number", the cache of voice region i has no historical semantic content (i.e., the cache of voice region i is empty), the current multi-turn conversation stored in the current session cache includes "buy bags" and "nearest shopping mall", while the historical semantic content stored in the cache of the other voice regions j is "what day of the week tomorrow". Therefore, the new semantic content of voice region i can constitute the same multi-turn conversation with the historical semantic content stored in the cache of the other voice regions j.
[0074] According to embodiments of this disclosure, the above method may optionally include an additional step: in response to determining that the current semantic content and the current multi-turn session do not belong to the same multi-turn session, determining whether the current semantic content and the historical semantic content stored in the region caches of the other regions (excluding the first region) belong to the same multi-turn session; and in response to determining that the current semantic content and the historical semantic content stored in the region cache of the second region belong to the same multi-turn session, storing the current semantic content and the historical semantic content stored in the region cache of the second region together in the current session cache to update the current multi-turn session. Here, the second region is a region different from the first region.
[0075] Continuing with the above exemplary embodiment, when process 400 proceeds through blocks 403 and 407 to block 408 and the judgment result at block 408 is negative, process 400 can proceed from block 408 to block 405. For example, this situation could be: the new semantic content of voice region i is "traffic restriction tail number", there is no historical semantic content in the cache of voice region i (i.e., the cache of voice region i is empty), the current multi-turn conversation stored in the current session cache includes "buy bags" and "nearest shopping mall", while the historical semantic content stored in the cache of the other voice regions j is either empty or has no contextual relevance to "traffic restriction tail number".
[0076] According to embodiments of this disclosure, the above method may optionally include an additional step: in response to determining that the current semantic content and the historical semantic content stored in the audio region caches of the other audio regions (excluding the first audio region) do not belong to the same multi-turn session, storing the current semantic content in the current session cache to update the current multi-turn session.
[0077] Process 400 also includes box 410, in which the cache of tone zone i is updated.
[0078] According to embodiments of this disclosure, the above method may optionally include an additional step: storing the current semantic content in the register cache of the first register to update the historical semantic content stored in the register cache of the first register.
[0079] It should be noted that process 400 can be executed in a loop. For example, after process 400 reaches box 410, it can return to box 401 to receive new semantic content.
[0080] Although the operations are depicted in the accompanying drawings in a specific order, this should not be construed as requiring that the operations be performed in the specific order shown or in chronological order, nor should it be construed as requiring that all the operations shown be performed to obtain the desired result.
[0081] Figure 5 This is a schematic block diagram illustrating a voice interaction device 500 for a vehicle cockpit according to an exemplary embodiment. The vehicle cockpit includes multiple voice zones corresponding to multiple vehicle seats, the multiple voice zones are jointly configured with a current session cache for storing current multi-turn conversations, and each of the multiple voice zones is configured with a corresponding voice zone cache for storing historical semantic content from that voice zone within a preset time window.
[0082] The apparatus 500 includes: an acquisition module 510 configured to acquire current semantic content from a first audio region among multiple audio regions; an update module 520 configured to update a current multi-turn session based on a current session cache, an audio region cache of the first audio region, and audio region caches of the other audio regions among the multiple audio regions excluding the first audio region, wherein the updated current multi-turn session is associated with the current semantic content; and a processing module 530 configured to process the current semantic content associated with the updated current multi-turn session.
[0083] It should be understood that Figure 5 The various modules of the device 500 shown can be connected to the reference. Figure 2 The steps in method 200 described correspond to each other. Therefore, the operations, features, and advantages described above for method 200 also apply to device 500 and its included modules.
[0084] According to embodiments of this disclosure, the device 500 overcomes the shortcomings of in-vehicle voice interaction systems in the related art, which respond to multiple voice inputs in a serial single-session interaction manner and cannot truly respond to multi-zone multi-session intentions in parallel, as well as the inability to determine whether there is a contextual relationship between the voice signals of multiple zones. The device 500 updates the current multi-turn conversation based on the current conversation cache, the zone cache of the first zone, and the zone caches of the other zones among the multiple zones excluding the first zone. By taking into account the contextual relationship of the multi-turn conversation, it can better process the current semantic content associated with the updated current multi-turn conversation, thereby better understanding and responding to the user's intentions and improving the user's driving experience.
[0085] While specific functions have been discussed above with reference to specific modules, it should be noted that the functions of the modules discussed herein can be divided into multiple modules, and / or at least some functions of multiple modules can be combined into a single module. The specific module performing an action discussed herein includes the specific module itself performing the action, or alternatively, the specific module calling or otherwise accessing another component or module that performs the action (or performs the action in conjunction with the specific module). Therefore, the specific module performing an action can include the specific module performing the action itself and / or another module that the specific module calls or otherwise accesses to perform the action. As used herein, the phrase "perform action Z based on A, B, and C" can mean performing action Z based only on A, only on B, only on C, based on A and B, based on A and C, based on B and C, or based on A, B, and C.
[0086] It should also be understood that this article can describe various technologies in the general context of software and hardware components or program modules. The above regarding... Figure 5 The various modules described can be implemented in hardware or in hardware in combination with software and / or firmware. For example, these modules can be implemented as computer program code / instructions configured to execute in one or more processors and stored in a computer-readable storage medium. Alternatively, these modules can be implemented as hardware logic / circuit. For example, in some embodiments, one or more of the acquisition module 510, update module 520, and processing module 530 can be implemented together in a System on Chip (SoC). The SoC may include an integrated circuit chip (which includes a processor (e.g., a Central Processing Unit (CPU), microcontroller, microprocessor, digital signal processor (DSP), etc.), memory, one or more communication interfaces, and / or one or more components of other circuitry) and may optionally execute received program code and / or include embedded firmware to perform functions.
[0087] According to one aspect of this disclosure, a computer device for a vehicle cockpit is provided. The vehicle cockpit includes multiple audio zones corresponding to multiple vehicle seats. These multiple audio zones are collectively configured with a current session cache for storing current multi-turn conversations, and each audio zone is configured with a corresponding audio zone cache for storing historical semantic content from that audio zone within a preset time window. The computer device includes at least one memory, at least one processor, and a computer program stored on the at least one memory. The at least one processor is configured to execute the computer program to implement the steps of any of the method embodiments described above.
[0088] According to one aspect of this disclosure, a vehicle is provided that includes the voice interaction device or computer equipment described above.
[0089] According to one aspect of this disclosure, a non-transitory computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of any of the method embodiments described above.
[0090] According to one aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the steps of any of the method embodiments described above.
[0091] In the following text, combined with Figure 6 Illustrative examples describing such computer devices, non-transitory computer-readable storage media, and computer program products.
[0092] Figure 6 An example configuration of a computer device 600 that can be used to implement the methods described herein is shown. For example, Figure 1 The server 120 and / or vehicle system 110 shown may include an architecture similar to computer device 600. The aforementioned device 500 or computer device may also be implemented wholly or at least partially by computer device 600 or similar devices or systems.
[0093] Computer device 600 may include at least one processor 602, memory 604, multiple communication interfaces 606, display device 608, other input / output (I / O) devices 610, and one or more mass storage devices 612 capable of communicating with each other, such as via system bus 614 or other suitable connections.
[0094] Processor 602 may be a single processing unit or multiple processing units, and all processing units may include single or multiple computing units or multiple cores. Processor 602 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuits, and / or any device that manipulates signals based on operating instructions. Among other capabilities, processor 602 may be configured to acquire and execute computer-readable instructions stored in memory 604, mass storage device 612, or other computer-readable media, such as program code of operating system 616, program code of application program 618, program code of other program 620, etc.
[0095] Memory 604 and mass storage device 612 are examples of computer-readable storage media for storing instructions executed by processor 602 to perform the various functions described above. For example, memory 604 may generally include both volatile and non-volatile memory (e.g., RAM, ROM, etc.). Furthermore, mass storage device 612 may generally include hard disk drives, solid-state drives, removable media, including external and removable drives, memory cards, flash memory, floppy disks, optical disks (e.g., CDs, DVDs), storage arrays, network-attached storage, storage area networks, etc. Both memory 604 and mass storage device 612 may be collectively referred to herein as memory or computer-readable storage media, and may be non-transitory media capable of storing computer-readable, processor-executable program instructions as computer program code, which may be executed by processor 602 as a specific machine configured to perform the operations and functions described in the examples herein.
[0096] Multiple programs may be stored on mass storage device 612. These programs include operating system 616, one or more application programs 618, other programs 620, and program data 622, and they may be loaded into memory 604 for execution. Examples of such application programs or program modules may include, for example, computer program logic (e.g., computer program code or instructions) for implementing the following components / functions: method 200, method 300 and their optional additional steps, apparatus 500, and / or other embodiments described herein.
[0097] Although Figure 6 The modules 616, 618, 620, and 622, or portions thereof, are illustrated as being stored in memory 604 of computer device 600; however, modules 616, 618, 620, and 622 may be implemented using any form of computer-readable medium accessible by computer device 600. As used herein, “computer-readable medium” includes at least two types of computer-readable media: computer-readable storage media and communication media.
[0098] Computer-readable storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, DVD, or other optical storage devices, magnetic cassettes, magnetic tapes, disk storage devices or other magnetic storage devices, or any other non-transmission medium that can be used to store information for access by a computer device. In contrast, communication media can embody computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms. Computer-readable storage media as defined herein do not include communication media.
[0099] One or more communication interfaces 606 are used for exchanging data with other devices, such as via a network, direct connection, etc. Such communication interfaces can be one or more of the following: any type of network interface (e.g., a network interface card (NIC)), wired or wireless (such as IEEE 802.11 Wireless LAN (WLAN)) wireless interface, Wi-MAX interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth. TM Interfaces, near field communication (NFC) interfaces, etc. Communication interface 606 can facilitate communication across various network and protocol types, including wired networks (e.g., LAN, cable, etc.) and wireless networks (e.g., WLAN, cellular, satellite, etc.), the Internet, etc. Communication interface 606 can also provide communication with external storage devices (not shown) such as storage arrays, network-attached storage, storage area networks, etc.
[0100] In some examples, a display device 608, such as a monitor, may be included for displaying information and images to the user. Other I / O devices 610 may be devices that receive various inputs from the user and provide various outputs to the user, and may include touch input devices, gesture input devices, cameras, keyboards, remote controls, mice, printers, audio input / output devices, and so on.
[0101] The technologies described herein can be supported by these various configurations of computer device 600, and are not limited to specific examples of the technologies described herein. For example, the functionality can also be implemented wholly or partially on a “cloud” using a distributed system. A cloud includes and / or represents a platform for resources. The platform abstracts the underlying functionality of the cloud’s hardware (e.g., servers) and software resources. Resources may include applications and / or data that can be used when performing computational processing on a server remote from computer device 600. Resources may also include services provided via the Internet and / or via subscriber networks such as cellular or Wi-Fi networks. The platform can abstract resources and functionality to connect computer device 600 to other computer devices. Therefore, the implementation of the functionality described herein can be distributed throughout the cloud. For example, the functionality may be implemented partly on computer device 600 and partly through a platform that abstracts the functionality of the cloud.
[0102] Although this disclosure has been described and illustrated in detail in the accompanying drawings and the foregoing description, such description and illustration should be considered illustrative and suggestive, not restrictive; this disclosure is not limited to the disclosed embodiments. By studying the drawings, the disclosure, and the appended claims, those skilled in the art will be able to understand and implement variations of the disclosed embodiments in practice with respect to the claimed subject matter. In the claims, the word "comprising" does not exclude other elements or steps not listed, the indefinite article "a" or "an" does not exclude a plurality, the term "a plurality" means two or more, and the term "based on" should be interpreted as "at least partially based on". The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be beneficial.
[0103] Some exemplary aspects of this disclosure will be described below.
[0104] Aspect 1, a voice interaction method for a vehicle cockpit, the vehicle cockpit including multiple voice zones corresponding to multiple vehicle seats, the multiple voice zones being jointly configured with a current session cache for storing current multi-turn conversations, and each of the multiple voice zones being configured with a corresponding voice zone cache for storing historical semantic content from that voice zone within a preset time window, the method comprising:
[0105] Retrieve the current semantic content of the first vocal region from multiple vocal regions;
[0106] Based on the current session cache, the first voice region's voice region cache, and the voice region caches of the remaining voice regions (excluding the first voice region) across multiple voice regions, update the current multi-turn session, where the updated current multi-turn session is associated with the current semantic content; and
[0107] Process the current semantic content associated with the updated current multi-turn session.
[0108] Aspect 2, the method of aspect 1, wherein updating the current multi-turn session includes:
[0109] Determine whether historical semantic content is stored in the register cache of the first register.
[0110] In response to determining that the voice region cache of the first voice region stores historical semantic content, it is determined whether the current semantic content and the historical semantic content belong to the same multi-turn session; and
[0111] In response to determining that the current semantic content and the historical semantic content belong to the same multi-turn session, the current semantic content and the historical semantic content are stored together in the current session cache to update the current multi-turn session.
[0112] The methods in aspect 3 and aspect 2 also include:
[0113] In response to determining that the current semantic content and the historical semantic content do not belong to the same multi-turn session, determine whether the current semantic content and the current multi-turn session belong to the same multi-turn session; and
[0114] In response to determining that the current semantic content and the current multi-turn session belong to the same multi-turn session, the current semantic content and the current multi-turn session are stored together in the current session cache to update the current multi-turn session.
[0115] The methods in aspect 4 and aspect 3 also include:
[0116] In response to determining that the current semantic content does not belong to the same multi-turn session, it determines whether the current semantic content and the historical semantic content stored in the region caches of the other regions (excluding the first region) belong to the same multi-turn session; and
[0117] In response to determining that the current semantic content and the historical semantic content stored in the second voice region's voice region cache belong to the same multi-turn session, the current semantic content and the historical semantic content stored in the second voice region's voice region cache are stored together in the current session cache to update the current multi-turn session, where the second voice region is a voice region different from the first voice region.
[0118] The methods in aspect 5 and aspect 4 also include:
[0119] In response to determining that the current semantic content and the historical semantic content stored in the audio caches of the other audio regions (excluding the first audio region) do not belong to the same multi-turn session, the current semantic content is stored in the current session cache to update the current multi-turn session.
[0120] The methods in aspect 6 and aspect 2 also include:
[0121] In response to determining that no historical semantic content is stored in the voice region cache of the first voice region, determine whether the current semantic content and the current multi-turn session belong to the same multi-turn session; and
[0122] In response to determining that the current semantic content and the current multi-turn session belong to the same multi-turn session, the current semantic content and the current multi-turn session are stored together in the current session cache to update the current multi-turn session.
[0123] The methods in aspects 7 and 6 also include:
[0124] In response to determining that the current semantic content does not belong to the same multi-turn session, it determines whether the current semantic content and the historical semantic content stored in the region caches of the other regions (excluding the first region) belong to the same multi-turn session; and
[0125] In response to determining that the current semantic content and the historical semantic content stored in the second voice region's voice region cache belong to the same multi-turn session, the current semantic content and the historical semantic content stored in the second voice region's voice region cache are stored together in the current session cache to update the current multi-turn session, where the second voice region is a voice region different from the first voice region.
[0126] The methods in aspects 8 and 7 also include:
[0127] In response to determining that the current semantic content and the historical semantic content stored in the audio caches of the other audio regions (excluding the first audio region) do not belong to the same multi-turn session, the current semantic content is stored in the current session cache to update the current multi-turn session.
[0128] The methods in aspect 9 and aspect 1 also include:
[0129] Store the current semantic content in the first register's register cache to update the historical semantic content stored in the first register's register cache.
[0130] Aspect 10, a voice interaction device for a vehicle cockpit, the vehicle cockpit including multiple voice zones corresponding to multiple vehicle seats, the multiple voice zones being jointly configured with a current session cache for storing current multi-turn conversations, and each of the multiple voice zones being configured with a corresponding voice zone cache for storing historical semantic content from that voice zone within a preset time window, the device comprising:
[0131] The acquisition module is configured to acquire the current semantic content of the first vocal region from multiple vocal regions;
[0132] The update module is configured to update the current multi-turn session based on the current session cache, the first voice region's voice region cache, and the voice region caches of the remaining voice regions (excluding the first voice region) across multiple voice regions. The updated current multi-turn session is associated with the current semantic content.
[0133] The processing module is configured to process the current semantic content associated with the updated current multi-turn session.
[0134] Aspect 11, a computer device for a vehicle cockpit, the vehicle cockpit including multiple audio zones corresponding to multiple vehicle seats, the multiple audio zones being jointly configured with a current session cache for storing current multi-turn sessions, and each audio zone being configured with a corresponding audio zone cache for storing historical semantic content from that audio zone within a preset time window, the computer device comprising:
[0135] At least one processor; and
[0136] At least one memory on which a computer program is stored,
[0137] Wherein, when a computer program is executed by at least one processor, it causes at least one processor to perform any one of aspects 1-9.
[0138] Aspect 12, a vehicle including the voice interaction device of aspect 10 or the computer equipment of aspect 11.
[0139] Aspect 13, a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the processor performs any one of the methods of aspects 1-9.
[0140] Aspect 14, a computer program product comprising a computer program that, when executed by a processor, causes the processor to perform any one of aspects 1-9.
Claims
1. A voice interaction method for a vehicle cockpit, the vehicle cockpit including multiple voice zones corresponding to multiple vehicle seats, the multiple voice zones being jointly configured with a current session cache for storing current multi-turn conversations, and each of the multiple voice zones being configured with a corresponding voice zone cache for storing historical semantic content from that voice zone within a preset time window, the method comprising: Obtain the current semantic content of the first vocal region from the plurality of vocal regions; Based on the current session cache, the region cache of the first voice region, and the region caches of the remaining voice regions (excluding the first voice region) among the plurality of voice regions, the current multi-turn session is updated, wherein the updated current multi-turn session is associated with the current semantic content; and Process the current semantic content associated with the updated current multi-turn session. The content stored in the current session cache is associated with the current sessions of multiple users within the entire vehicle cabin space. Updating the current multi-turn session includes: Determine whether historical semantic content is stored in the register cache of the first register; In response to determining that historical semantic content is stored in the voice region cache of the first voice region, it is determined whether the current semantic content and the historical semantic content belong to the same multi-turn session; and In response to determining that the current semantic content and the historical semantic content belong to the same multi-turn session, the current semantic content and the historical semantic content are stored together in the current session cache to update the current multi-turn session. The method further includes: In response to determining that the current semantic content and the historical semantic content do not belong to the same multi-turn session, determine whether the current semantic content and the current multi-turn session belong to the same multi-turn session; In response to determining that the current semantic content and the current multi-turn session do not belong to the same multi-turn session, it is determined whether the current semantic content and the historical semantic content stored in the region caches of the other regions (excluding the first region) belong to the same multi-turn session; and In response to determining that the current semantic content and the historical semantic content stored in the audio region cache of the second audio region among the plurality of audio regions belong to the same multi-turn session, the current semantic content and the historical semantic content stored in the audio region cache of the second audio region are stored together in the current session cache to update the current multi-turn session, wherein the second audio region is an audio region different from the first audio region.
2. The method according to claim 1, further comprising: In response to determining that the current semantic content and the current multi-turn session belong to the same multi-turn session, the current semantic content and the current multi-turn session are stored together in the current session cache to update the current multi-turn session.
3. The method according to claim 1, further comprising: In response to determining that the current semantic content and the historical semantic content stored in the audio region caches of the other audio regions (excluding the first audio region) do not belong to the same multi-turn session, the current semantic content is stored in the current session cache to update the current multi-turn session.
4. The method according to claim 1, further comprising: In response to determining that no historical semantic content is stored in the audio region cache of the first audio region, it is determined whether the current semantic content and the current multi-turn session belong to the same multi-turn session; as well as In response to determining that the current semantic content and the current multi-turn session belong to the same multi-turn session, the current semantic content and the current multi-turn session are stored together in the current session cache to update the current multi-turn session.
5. The method according to claim 4, further comprising: In response to determining that the current semantic content and the current multi-turn session do not belong to the same multi-turn session, determine whether the current semantic content and the historical semantic content stored in the audio region cache of the other audio regions (excluding the first audio region) belong to the same multi-turn session; as well as In response to determining that the current semantic content and the historical semantic content stored in the audio region cache of the second audio region among the plurality of audio regions belong to the same multi-turn session, the current semantic content and the historical semantic content stored in the audio region cache of the second audio region are stored together in the current session cache to update the current multi-turn session, wherein the second audio region is an audio region different from the first audio region.
6. The method according to claim 5, further comprising: In response to determining that the current semantic content and the historical semantic content stored in the audio region caches of the other audio regions (excluding the first audio region) do not belong to the same multi-turn session, the current semantic content is stored in the current session cache to update the current multi-turn session.
7. The method according to claim 1, further comprising: The current semantic content is stored in the register cache of the first register to update the historical semantic content stored in the register cache of the first register.
8. A voice interaction device for a vehicle cockpit, the vehicle cockpit including multiple voice zones corresponding to multiple vehicle seats, the multiple voice zones being jointly configured with a current session cache for storing current multi-turn conversations, and each of the multiple voice zones being configured with a corresponding voice zone cache for storing historical semantic content from that voice zone within a preset time window, the device comprising: The acquisition module is configured to acquire the current semantic content from the first voice region among the plurality of voice regions; An update module, configured to update the current multi-turn session based on the current session cache, the register cache of the first register, and the register caches of the other registers among the plurality of registers excluding the first register, wherein the updated current multi-turn session is associated with the current semantic content; and A processing module, configured to process the current semantic content associated with the updated current multi-turn session, The content stored in the current session cache is associated with the current sessions of multiple users within the entire vehicle cabin space. Updating the current multi-turn session includes: Determine whether historical semantic content is stored in the register cache of the first register; In response to determining that historical semantic content is stored in the voice region cache of the first voice region, it is determined whether the current semantic content and the historical semantic content belong to the same multi-turn session; and In response to determining that the current semantic content and the historical semantic content belong to the same multi-turn session, the current semantic content and the historical semantic content are stored together in the current session cache to update the current multi-turn session. The device is also used for: In response to determining that the current semantic content and the historical semantic content do not belong to the same multi-turn session, determine whether the current semantic content and the current multi-turn session belong to the same multi-turn session; In response to determining that the current semantic content and the current multi-turn session do not belong to the same multi-turn session, it is determined whether the current semantic content and the historical semantic content stored in the region caches of the other regions (excluding the first region) belong to the same multi-turn session; and In response to determining that the current semantic content and the historical semantic content stored in the audio region cache of the second audio region among the plurality of audio regions belong to the same multi-turn session, the current semantic content and the historical semantic content stored in the audio region cache of the second audio region are stored together in the current session cache to update the current multi-turn session, wherein the second audio region is an audio region different from the first audio region.
9. A computer device for a vehicle cockpit, the vehicle cockpit including multiple audio zones corresponding to multiple vehicle seats, the multiple audio zones being jointly configured with a current session cache for storing current multi-turn sessions, and each of the multiple audio zones being configured with a corresponding audio zone cache for storing historical semantic content from that audio zone within a preset time window, the computer device comprising: At least one processor; as well as At least one memory on which a computer program is stored, When the computer program is executed by the at least one processor, it causes the at least one processor to perform the method according to any one of claims 1-7.
10. A vehicle comprising the voice interaction device as claimed in claim 8 or the computer device as claimed in claim 9.
11. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the processor to perform the method of any one of claims 1-7.
12. A computer program product comprising a computer program that, when executed by a processor, causes the processor to perform the method of any one of claims 1-7.
Citation Information
Patent Citations
Voice interaction method and system
CN111785266A
Multi-sound-zone voice interaction method for vehicle and electronic equipment
CN111816189A