Voice interaction method, device, equipment, storage medium and vehicle
By sorting and strategically selecting multi-region speech information, the problem of missing single-region speech information in multi-person voice interaction is solved, and the simultaneous input and complete interaction of multi-region speech information are realized.
Patent Information
- Application Number
- CN202310269335.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-15
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2043-03-15
AI Technical Summary
In existing technologies, a single recognition engine can only process one channel of data, which leads to the omission of speech from people in other voice regions in scenarios where multiple people are speaking at the same time, making it impossible to respond to the needs of multi-person voice interaction.
By acquiring multi-region speech information, sorting the speech information according to a preset speech caching strategy, forming a speech queue of wake-up and non-wake-up regions, and selecting target speech information from it according to a preset speech sending strategy to send to the recognition engine for recognition, the simultaneous input of multi-region speech information is achieved.
It fulfills the needs of multi-person voice interaction scenarios, avoids the omission problem caused by suppressing other voice regions when inputting in a single voice region, and ensures the integrity of voice interaction to the greatest extent.
Smart Images

Figure CN118675514B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of voice technology, and in particular to a voice interaction method, apparatus, device, storage medium, and vehicle. Background Technology
[0002] With the development of voice technology, vehicles can support voice control services, such as voice control for opening windows. In real-world driving scenarios, there may be situations where multiple people speak simultaneously, meaning users can issue voice commands from multiple audio zones within the vehicle.
[0003] In existing technologies, a single recognition engine can only process one channel of data. Only a single input channel can correctly respond to the speech of a person in a single vocal range. This is because the recognition engine will suppress input from other vocal ranges when inputting from a single vocal range, resulting in the omission of speech from people in other vocal ranges and making it unable to respond to the needs of multi-person voice interaction scenarios. Summary of the Invention
[0004] To address, or at least partially address, the aforementioned technical problems, this disclosure provides a voice interaction method, apparatus, device, storage medium, and vehicle to meet the needs of multi-person voice interaction scenarios.
[0005] In a first aspect, embodiments of this disclosure provide a voice interaction method, including:
[0006] Acquire multiple voice information from multiple voice regions, the multiple voice regions including a wake-up voice region and at least one non-wake-up voice region, the wake-up voice region being the voice region where the vehicle's voice function is activated;
[0007] The multiple voice information are sorted according to a preset voice caching strategy to obtain a wake-up voice queue and at least one non-wake-up voice queue.
[0008] According to a preset sound delivery strategy, target speech information is selected from the wake-up speech queue and the at least one non-wake-up speech queue;
[0009] The target speech information is sent to the recognition engine to obtain the recognition result of the target speech information, and the speech interaction of the target speech information is completed.
[0010] Secondly, embodiments of this disclosure provide a voice interaction device, including:
[0011] The acquisition module is used to acquire multiple voice information from multiple voice regions, the multiple voice regions including a wake-up voice region and at least one non-wake-up voice region, the wake-up voice region being the voice region where the vehicle's voice function is activated;
[0012] The sorting module is used to sort the multiple voice information according to a preset voice caching strategy to obtain a wake-up voice queue and at least one non-wake-up voice queue.
[0013] The selection module is used to select target speech information from the wake-up speech queue and the at least one non-wake-up speech queue according to a preset speech delivery strategy.
[0014] The recognition module is used to send the target speech information to the recognition engine, obtain the recognition result of the target speech information, and complete the voice interaction of the target speech information.
[0015] Thirdly, embodiments of this disclosure provide an electronic device, including:
[0016] Memory;
[0017] Processor; and
[0018] Computer programs;
[0019] The computer program is stored in the memory and configured to be executed by the processor to implement the method as described in the first aspect.
[0020] Fourthly, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the method described in the first aspect.
[0021] Fifthly, embodiments of this disclosure also provide a vehicle, including: a voice interaction device as described in the second aspect; or an electronic device as described in the third aspect; or a computer-readable storage medium as described in the fourth aspect.
[0022] The voice interaction method, apparatus, device, storage medium, and vehicle provided in this disclosure acquire multiple voice information from multiple voice zones; sort the multiple voice information according to a preset voice caching strategy to obtain a wake-up voice zone queue and at least one non-wake-up voice zone queue, thus clarifying the order of voice information in each voice zone; select target voice information from the wake-up voice zone queue and the at least one non-wake-up voice zone queue according to a preset voice transmission strategy, thus clarifying the target voice information to be recognized and interacted with; send the target voice information to the recognition engine to obtain the recognition result of the target voice information, and complete the voice interaction of the target voice information. This achieves simultaneous input of voice information from multiple voice zones, avoiding the problem in the prior art where suppressing input from other voice zones during single-voice zone input leads to the omission of speech from people in other voice zones, and fulfills the needs of multi-person voice interaction scenarios, maximizing the integrity of voice interaction. Attached Figure Description
[0023] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0024] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 A flowchart of a voice interaction method provided in an embodiment of this disclosure;
[0026] Figure 2 A schematic diagram illustrating the receipt of voice information provided in an embodiment of this disclosure;
[0027] Figure 3 A schematic diagram of the voice interaction method provided in the embodiments of this disclosure;
[0028] Figure 4 This is a framework diagram of a voice interaction method provided in an embodiment of the present disclosure;
[0029] Figure 5 A flowchart of a voice interaction method provided in another embodiment of this disclosure;
[0030] Figure 6 This is a schematic diagram of the structure of the voice interaction device provided in the embodiments of this disclosure;
[0031] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0032] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0033] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.
[0034] This disclosure provides a voice interaction method, which will be described below with reference to specific embodiments.
[0035] Figure 1This is a flowchart illustrating a voice interaction method provided in an embodiment of this disclosure. The method can be executed by a voice interaction device, which can be implemented in software and / or hardware. The voice interaction device can be configured in an electronic device, such as a server or terminal, where the terminal specifically includes electric vehicles and non-electric vehicles. Furthermore, this method can be applied to various voice interaction scenarios; it is understood that the voice interaction method provided in this disclosure can also be applied to other scenarios.
[0036] The following is about Figure 1 The voice interaction method shown is described below, and the specific steps of this method are as follows:
[0037] S101. Obtain multiple voice information from multiple voice regions, wherein the multiple voice regions include a wake-up voice region and at least one non-wake-up voice region, and the wake-up voice region is the voice region where the vehicle's voice function is activated.
[0038] A microphone (MIC), also known as a transducer or microphone, is an energy conversion device that converts sound signals into electrical signals. Microphones are classified into dynamic, condenser, electret, and the more recently developed silicon micromicrophones, as well as liquid microphones and laser microphones. Most microphones are electret condenser microphones, which work by using a diaphragm made of polymer material with permanent charge isolation.
[0039] Each seat in the vehicle is equipped with a corresponding microphone to detect voice information. The relative distance between each seat and its corresponding microphone is fixed. When a microphone receives voice information, it sends the information to the vehicle's infotainment system. The system compares the voice information received from each microphone to determine the microphone with the strongest voice amplitude. Simultaneously, the system locates the sound's origin using the voice information received from each microphone. Based on the strength and location of the voice information, the system ultimately determines the sound's region, or frequency zone. The frequency zone where the vehicle's voice function is activated is called the wake-up frequency zone, while other frequency zones are non-wake-up frequency zones. In other words, the vehicle has only one wake-up frequency zone and at least one non-wake-up frequency zone.
[0040] Specifically, taking a four-seater car as an example, such as Figure 2As shown, the microphone corresponding to the driver's seat is MIC21, the microphone corresponding to the passenger seat is MIC22, the microphone corresponding to the left second-row seat is MIC23, and the microphone corresponding to the right second-row seat is MIC24. When a user speaks from the passenger seat, MIC21, MIC22, MIC23, and MIC24 can all receive the voice information. Due to the different distances between the speaking position and MIC21, MIC22, MIC23, and MIC24, the intensity of the voice information received by MIC21, MIC22, MIC23, and MIC24 varies. For the same sound source, the closer the distance, the louder the sound; correspondingly, the closer the microphone, the stronger the received voice information. Therefore, when MIC21 sends its received voice information 1, MIC22 sends its received voice information 2, MIC23 sends its received voice information 3, and MIC24 sends its received voice information 4 to the vehicle's infotainment system 20, the system receives these voice information 1, 2, 3, and 4, analyzes and compares them, and determines the strongest voice information (voice information 2 is the strongest, voice information 1 and 3 are relatively strong, and voice information 4 is the weakest). This determines that the sound source is in the passenger seat, meaning the user's voice information is in the passenger seat's vocal range. It is understood that this applies not only to four-seater vehicles but also to six-seater vehicles and other multi-seat vehicles. Furthermore, the principle is the same when sound is emitted from other locations within the vehicle, and this will not be elaborated upon further. It is also understood that the microphones can be placed in front of or behind the passenger seat; this embodiment does not limit this, and the principle remains the same, so it will not be elaborated upon further.
[0041] The vehicle infotainment system 20 acquires multiple voice information from multiple voice zones within the vehicle. These multiple voice zones refer to the driver's voice zone, passenger's voice zone, second-row left voice zone, and second-row right voice zone. When a user activates the vehicle's voice function, the voice zone corresponding to that user is identified as the activation voice zone. After that user activates the voice function, other users subsequently activate their voices, and the two voice zones in which those other users activate their voices are identified as non-activation voice zones. For example, if a user activates the vehicle's voice function from the passenger's seat, the passenger's voice zone is the activation voice zone. If other users activate their voice information from the driver's voice zone, the driver's voice zone is identified as a non-activation voice zone, resulting in one non-activation voice zone. If other users activate their voice information from the driver's voice zone and the second-row left voice zone, then the driver's voice zone and the second-row left voice zone are identified as non-activation voice zones, resulting in two non-activation voice zones. In other embodiments, there may also be three non-activation voice zones. For example, if a user activates the vehicle's voice function from the passenger's seat, the passenger's voice zone is the activation voice zone, and the driver's voice zone, second-row left voice zone, and second-row right voice zone are identified as non-activation voice zones.
[0042] S102. Sort the multiple voice information according to the preset voice caching strategy to obtain a wake-up voice queue and at least one non-wake-up voice queue.
[0043] The vehicle system 20 sorts multiple voice information from each voice zone according to a preset voice caching strategy, resulting in a wake-up voice zone queue and at least one non-wake-up voice zone queue. The wake-up voice zone queue contains the voice information of the wake-up user, which can be directly sorted according to time order. The at least one non-wake-up voice zone queue contains the voice information of the non-wake-up user. When there is one non-wake-up voice zone, the voice information of the non-wake-up user is sorted according to the preset voice caching strategy. When there are two non-wake-up voice zones, the weights of the two non-wake-up voice zones are determined to clarify the priority of the voice information of the two non-wake-up users. The voice information of the non-wake-up users is then sorted according to the preset voice caching strategy. In this embodiment, to ensure the quality of multi-person voice interaction, there can be at most two non-wake-up voice zones. The priority of the non-wake-up voice zones can be determined according to their weights, which depend on the time order in which the user sends the voice information.
[0044] Optionally, the at least one non-wake-up sound area includes a first non-wake-up sound area and a second non-wake-up sound area; the weights of the first non-wake-up sound area and the second non-wake-up sound area are obtained according to the order in which the vehicle voice function is responded to; based on the fact that the weight of the first non-wake-up sound area is greater than the weight of the second non-wake-up sound area, the priority of the first non-wake-up sound area is higher than that of the second non-wake-up sound area.
[0045] Specifically, when there are two non-wake-up sound zones, namely the first non-wake-up sound zone and the second non-wake-up sound zone, the weights of the first and second non-wake-up sound zones are determined according to the order in which the vehicle's voice function is responded to. The first non-wake-up sound zone is the one whose voice function is responded to first, and the second non-wake-up sound zone is the one whose voice function is responded to later. Therefore, the weight of the first non-wake-up sound zone is greater than the weight of the second non-wake-up sound zone. Based on the fact that the weight of the first non-wake-up sound zone is greater than the weight of the second non-wake-up sound zone, the priority of the first non-wake-up sound zone is determined to be higher than that of the second non-wake-up sound zone.
[0046] It is understandable that the channel of the wake-up voice queue is unique and immutable; while the channel of the non-wake-up voice queue is mutable. Specifically, if the driver's voice queue is the wake-up voice queue, and the passenger's voice queue speaks first, then the passenger's voice queue becomes the first non-wake-up voice queue. Similarly, if the second-row left voice queue speaks first, then the second-row left voice queue becomes the first non-wake-up voice queue. The first and second non-wake-up voice queues are not fixed. Given a determined wake-up voice queue, the voice queue that speaks first becomes the first non-wake-up voice queue, and the subsequent voice queue becomes the non-wake-up voice queue.
[0047] S103. Select target speech information from the wake-up speech queue and the at least one non-wake-up speech queue according to the preset speech delivery strategy.
[0048] According to a preset audio delivery strategy, the vehicle system 20 selects target audio information from multiple audio information in the wake-up audio queue and multiple audio information in at least one non-wake-up audio queue.
[0049] S104. The target speech information is sent to the recognition engine to obtain the recognition result of the target speech information and complete the speech interaction of the target speech information.
[0050] The vehicle system 20 sends the target voice information to a single recognition engine for recognition, obtains the recognition result of the target voice information, and completes the voice interaction of the target voice information.
[0051] This embodiment of the disclosure acquires multiple voice information from multiple voice regions; sorts the multiple voice information according to a preset voice caching strategy to obtain a wake-up voice region voice queue and at least one non-wake-up voice region voice queue, thus clarifying the order of voice information in each voice region; selects target voice information from the wake-up voice region voice queue and the at least one non-wake-up voice region voice queue according to a preset voice transmission strategy, thus clarifying the target voice information to be used for recognition interaction; sends the target voice information to the recognition engine to obtain the recognition result of the target voice information, and completes the voice interaction of the target voice information. This realizes the simultaneous input of voice information from multiple voice regions, avoiding the problem in the prior art where suppressing other voice region inputs during single voice region input leads to the omission of speech from other voice region users, thus meeting the needs of multi-person voice interaction scenarios and maximizing the integrity of voice interaction.
[0052] Based on the above embodiments, the step of sorting the plurality of voice information according to a preset voice caching strategy to obtain a wake-up voice queue and at least one non-wake-up voice queue includes: for each voice information, determining whether the voice information is in the wake-up voice queue; if the voice information is in the wake-up voice queue, placing the voice information in the wake-up voice queue; if the voice information is not in the wake-up voice queue, determining whether the voice information is in at least one non-wake-up voice queue; if the voice information is in at least one non-wake-up voice queue, placing the voice information in the corresponding non-wake-up voice queue; and if the voice information is not in at least one non-wake-up voice queue, discarding the voice information.
[0053] For each voice message, the vehicle system 20 determines whether the voice message is in the wake-up voice zone. If it is, the voice message is placed in the wake-up voice zone queue in chronological order. If it is not, the vehicle system determines whether the voice message is in the non-wake-up voice zone. If it is, the voice message is placed in the non-wake-up voice zone queue corresponding to the voice message. If it is not, the voice message is discarded.
[0054] This disclosure, through a detailed description of how to obtain the wake-up voice queue and at least one non-wake-up voice queue, clarifies the preset voice caching strategy, provides basic data for subsequent acquisition of target voice information, and improves the flexibility of the voice interaction method.
[0055] In some embodiments, the voice queue includes a voice detection start point, voice information, and a voice detection end point;
[0056] Voice Activity Detection (VAD), also known as voice endpoint detection or voice boundary detection, aims to identify and eliminate long periods of silence in the audio signal stream. This saves voice channel resources without degrading service quality and is a crucial component of IP telephony applications. Silence suppression can conserve valuable bandwidth resources and reduce perceived end-to-end latency.
[0057] The speech queue includes a speech detection start point, speech information, and a speech detection end point. The speech detection start point (VADbegin) refers to the beginning of the speech information, and the speech detection end point (VAD end) refers to the end of the speech information.
[0058] The preset sound delivery strategy includes: a wake-up sound zone sound delivery strategy;
[0059] The wake-up sound zone delivery strategy includes: responding to a first speech detection starting point of the wake-up sound zone speech queue, detecting whether the wake-up sound zone speech queue is an empty set; if the wake-up sound zone speech queue is a non-empty set, then the speech information corresponding to the first speech detection starting point is placed into the wake-up sound zone delivery queue; if the wake-up sound zone speech queue is an empty set, and the at least one non-wake-up sound zone is an empty set, then the speech information corresponding to the first speech detection starting point is placed into the wake-up sound zone delivery queue; if the wake-up sound zone speech queue is an empty set, and the at least one non-wake-up sound zone is a non-empty set, then the speech information of the at least one non-wake-up sound zone is cleared, and the speech information corresponding to the first speech detection starting point is placed into the wake-up sound zone delivery queue.
[0060] Specifically, the vehicle system 20 detects the first voice detection starting point of the wake-up voice queue. In response to the first voice detection starting point, it checks whether the wake-up voice queue is an empty set, i.e., whether there is any voice information in the wake-up voice queue. If the wake-up voice queue is a non-empty set, i.e., there is voice information in the wake-up voice queue, the voice information corresponding to the first voice detection starting point is placed into the wake-up voice delivery queue. If the wake-up voice queue is an empty set, and at least one non-wake-up voice queue is an empty set, i.e., neither the wake-up voice queue nor at least one non-wake-up voice queue has any voice information, the voice information corresponding to the first voice detection starting point is placed into the wake-up voice delivery queue. If the wake-up voice queue is an empty set, and the at least one non-wake-up voice queue is a non-empty set, i.e., there is no message in the wake-up voice queue and at least one non-wake-up voice queue has voice information, the voice information of at least one non-wake-up voice queue is cleared, and the voice information corresponding to the first voice detection starting point is placed into the wake-up voice delivery queue.
[0061] The preset sound delivery strategy also includes: a non-wake-up sound zone sound delivery strategy;
[0062] The non-wake-up voice region delivery strategy includes: responding to the second speech detection starting point of the non-wake-up voice region speech queue, detecting whether the non-wake-up voice region speech queue is an empty set; if the non-wake-up voice region speech queue is an empty set, then arranging the speech information corresponding to the second speech detection starting point into the non-wake-up voice region delivery queue.
[0063] Specifically, the vehicle system 20 detects the second voice detection starting point of the non-wake-up voice queue. In response to the second voice detection starting point, it detects whether the non-wake-up voice queue is an empty set, that is, whether there is voice information in the non-wake-up voice queue. If the non-wake-up voice queue is an empty set, that is, whether there is voice information in the non-wake-up voice queue, then the voice information corresponding to the second voice detection starting point is placed into the non-wake-up voice transmission queue.
[0064] This disclosure, through a detailed description of a preset speech delivery strategy, lays the foundation for subsequent determination of target speech information, further improving the flexibility of the voice interaction method.
[0065] In some embodiments, sending the target voice information to a recognition engine to obtain a recognition result of the target voice information and completing the voice interaction of the target voice information includes: sending the target voice information to a recognition engine to obtain a semantic recognition result of the target voice information; controlling the vehicle to execute the semantic recognition result and completing the voice interaction of the target voice information.
[0066] Voice recognition, also known as Automatic Speech Recognition (ASR), is a technology that enables machines to convert speech signals into corresponding text or commands through recognition and understanding. It belongs to the interdisciplinary subfield of computational linguistics and is essentially a pattern recognition process. The pattern of unknown speech is compared with a reference pattern of known speech, and the best-matching reference pattern is used as the recognition result. It mainly includes three aspects: feature extraction techniques, pattern matching criteria, and model training techniques.
[0067] Semantic recognition, also known as semantic understanding, is the process by which machines understand textual content and identify the intended message. The fields involved in semantic recognition technology include signal processing, pattern recognition, probability theory and information theory, vocal and auditory mechanisms, artificial intelligence, and more.
[0068] The difference between speech recognition and semantic recognition can be vividly illustrated using human organs: speech recognition is like the mouth and ears, responsible for expression and acquisition; semantic recognition is like the brain, responsible for thinking and information processing.
[0069] The vehicle infotainment system 20 sends the target voice information to the recognition engine, which can be a single recognition engine. The recognition engine performs speech recognition on the target voice information to obtain the speech recognition result. The recognition engine then performs semantic recognition on the speech recognition result to obtain the semantic recognition result. The recognition engine sends the semantic recognition result to the vehicle infotainment system 20, so that the vehicle infotainment system 20 obtains the semantic recognition result of the target voice information. The vehicle infotainment system controls the vehicle to execute the semantic recognition result and complete the voice interaction of the target voice information.
[0070] This embodiment of the disclosure sends the target voice information to the recognition engine to obtain the semantic recognition result of the target voice information; controls the vehicle to execute the semantic recognition result to complete the voice interaction of the target voice information, realizing the scenario requirements of multi-person voice interaction and ensuring the integrity of voice interaction to the maximum extent.
[0071] Figure 3 A schematic diagram of the voice interaction method provided in the embodiments of this disclosure, such as... Figure 3 As shown, register 1, register 2, register 3, and register 4 correspond to respectively Figure 2 After acquiring multiple voice information from four voice zones (voice zone 1, voice zone 2, voice zone 3, and voice zone 4) using MIC21, MIC22, MIC23, and MIC24, the system sends these multiple voice information to the recognition engine for recognition according to a preset voice transmission strategy. The recognition result is then obtained, and the vehicle system executes the recognition result to complete the voice interaction of the target voice information.
[0072] Figure 4This is a framework diagram of the voice interaction method provided in the embodiments of this disclosure, such as... Figure 4 As shown, register 1, register 2, register 3, and register 4 correspond to respectively Figure 2 The vehicle system uses MIC21, MIC22, MIC23, and MIC24. After acquiring multiple voice information from four voice zones (zone 1, zone 2, zone 3, and zone 4), it determines the wake-up and non-wake-up voice zones within these four zones. The wake-up voice zone is the zone where the vehicle's voice function is activated. The vehicle system sorts the multiple voice information from multiple zones according to a preset voice caching strategy, resulting in a wake-up voice zone queue, a non-wake-up voice zone queue 1, and a non-wake-up voice zone queue 2. The non-wake-up voice zone queue 1 is the aforementioned first non-wake-up voice zone queue, and the non-wake-up voice zone queue 2 is... The second non-wake-up voice zone is described above. The vehicle system selects target voice information from the wake-up voice zone queue, non-wake-up voice zone queue 1, and non-wake-up voice zone queue 2 according to a preset voice delivery strategy. The target voice information is sent to the recognition engine, which performs speech recognition on the target voice information to obtain a speech recognition result. The recognition engine then performs semantic recognition on the speech recognition result to obtain a semantic recognition result, which is sent to the vehicle system, allowing the vehicle system to obtain the semantic recognition result of the target voice information. The vehicle system controls the vehicle to execute the semantic recognition result, completing the voice interaction of the target voice information. This achieves simultaneous input of multi-voice zone voice information, avoiding the problem in existing technologies where single-voice zone input suppresses other voice zone inputs, leading to the omission of speech from other voice zone users. It fulfills the needs of multi-person voice interaction scenarios.
[0073] Figure 5 A flowchart of a voice interaction method provided in another embodiment of this disclosure is shown below. Figure 5 As shown, the specific steps included in this method are as follows:
[0074] S501. Obtain multiple voice information from multiple voice zones, wherein the multiple voice zones include a wake-up voice zone and at least one non-wake-up voice zone, and the wake-up voice zone is the voice zone where the vehicle's voice function is activated.
[0075] Specifically, the implementation process and principle of S501 and S101 are the same, and will not be repeated here.
[0076] S502. For each voice message, determine whether the voice message is in the wake-up voice zone. If yes, proceed to step S503; otherwise, proceed to step S504.
[0077] S503. The voice information is placed in the voice queue of the wake-up voice area.
[0078] The speech queue includes a speech detection start point, speech information, and a speech detection end point. The speech detection start point (VADbegin) refers to the beginning of the speech information, and the speech detection end point (VAD end) refers to the end of the speech information.
[0079] S504. Determine whether the voice information is in at least one non-wake-up voice zone. If yes, proceed to step S505; otherwise, proceed to step S506.
[0080] S505. The voice information is placed in the non-wake-up voice queue corresponding to the voice information.
[0081] S506. Discard the voice information.
[0082] S507. Select target speech information from the wake-up speech queue and the at least one non-wake-up speech queue according to the preset speech delivery strategy.
[0083] The preset sound delivery strategies include: wake-up sound zone sound delivery strategy and non-wake-up sound zone sound delivery strategy.
[0084] The wake-up sound region delivery strategy includes: responding to a first speech detection starting point in the wake-up sound region speech queue, detecting whether the wake-up sound region speech queue is an empty set; if the wake-up sound region speech queue is a non-empty set, then the speech information corresponding to the first speech detection starting point is placed into the wake-up sound region delivery queue; if the wake-up sound region speech queue is an empty set, and the at least one non-wake-up sound region is an empty set, then the speech information corresponding to the first speech detection starting point is placed into the wake-up sound region delivery queue; if the wake-up sound region speech queue is an empty set, and the at least one non-wake-up sound region is a non-empty set, then the speech information of the at least one non-wake-up sound region is cleared, and the speech information corresponding to the first speech detection starting point is placed into the wake-up sound region delivery queue.
[0085] Specifically, the vehicle system 20 detects the first voice detection starting point of the wake-up voice queue. In response to the first voice detection starting point, it checks whether the wake-up voice queue is an empty set, i.e., whether there is any voice information in the wake-up voice queue. If the wake-up voice queue is a non-empty set, i.e., there is voice information in the wake-up voice queue, the voice information corresponding to the first voice detection starting point is placed into the wake-up voice delivery queue. If the wake-up voice queue is an empty set, and at least one non-wake-up voice queue is an empty set, i.e., neither the wake-up voice queue nor at least one non-wake-up voice queue has any voice information, the voice information corresponding to the first voice detection starting point is placed into the wake-up voice delivery queue. If the wake-up voice queue is an empty set, and the at least one non-wake-up voice queue is a non-empty set, i.e., there is no message in the wake-up voice queue and at least one non-wake-up voice queue has voice information, the voice information of at least one non-wake-up voice queue is cleared, and the voice information corresponding to the first voice detection starting point is placed into the wake-up voice delivery queue.
[0086] The non-wake-up voice region delivery strategy includes: in response to the second speech detection starting point of the non-wake-up voice region speech queue, detecting whether the non-wake-up voice region speech queue is an empty set; if the non-wake-up voice region speech queue is an empty set, then the speech information corresponding to the second speech detection starting point is placed into the non-wake-up voice region delivery queue.
[0087] Specifically, the vehicle system 20 detects the second voice detection starting point of the non-wake-up voice queue. In response to the second voice detection starting point, it detects whether the non-wake-up voice queue is an empty set, that is, whether there is voice information in the non-wake-up voice queue. If the non-wake-up voice queue is an empty set, that is, whether there is voice information in the non-wake-up voice queue, then the voice information corresponding to the second voice detection starting point is placed into the non-wake-up voice transmission queue.
[0088] According to a preset audio delivery strategy, the vehicle system 20 selects target audio information from multiple audio information in the wake-up audio queue and multiple audio information in at least one non-wake-up audio queue.
[0089] S508. The target speech information is sent to the recognition engine to obtain the semantic recognition result of the target speech information.
[0090] Voice recognition, also known as Automatic Speech Recognition (ASR), is a technology that enables machines to convert speech signals into corresponding text or commands through recognition and understanding. It belongs to the interdisciplinary subfield of computational linguistics and is essentially a pattern recognition process. The pattern of unknown speech is compared with a reference pattern of known speech, and the best-matching reference pattern is used as the recognition result. It mainly includes three aspects: feature extraction techniques, pattern matching criteria, and model training techniques.
[0091] Semantic recognition, also known as semantic understanding, is the process by which machines understand textual content and identify the intended message. The fields involved in semantic recognition technology include signal processing, pattern recognition, probability theory and information theory, vocal and auditory mechanisms, artificial intelligence, and more.
[0092] The difference between speech recognition and semantic recognition can be vividly illustrated using human organs: speech recognition is like the mouth and ears, responsible for expression and acquisition; semantic recognition is like the brain, responsible for thinking and information processing.
[0093] The vehicle system 20 sends the target voice information to the recognition engine, which can be a single recognition engine. The recognition engine performs speech recognition on the target voice information to obtain the speech recognition result. The recognition engine then performs semantic recognition on the speech recognition result to obtain the semantic recognition result. The recognition engine sends the semantic recognition result to the vehicle system 20, so that the vehicle system 20 obtains the semantic recognition result of the target voice information.
[0094] S509. Control the vehicle to execute the semantic recognition result and complete the voice interaction of the target voice information.
[0095] The vehicle's infotainment system executes the semantic recognition results to complete the voice interaction of the target voice information.
[0096] This embodiment of the disclosure acquires multiple voice information from multiple voice regions; sorts the multiple voice information according to a preset voice caching strategy to obtain a wake-up voice region voice queue and at least one non-wake-up voice region voice queue, thus clarifying the order of voice information in each voice region; selects target voice information from the wake-up voice region voice queue and the at least one non-wake-up voice region voice queue according to a preset voice transmission strategy, thus clarifying the target voice information to be used for recognition interaction; sends the target voice information to the recognition engine to obtain the recognition result of the target voice information, and completes the voice interaction of the target voice information. This realizes the simultaneous input of voice information from multiple voice regions, avoiding the problem in the prior art where suppressing other voice region inputs during single voice region input leads to the omission of speech from other voice region users, thus meeting the needs of multi-person voice interaction scenarios and maximizing the integrity of voice interaction.
[0097] Figure 6 This is a schematic diagram of the structure of a voice interaction device provided in an embodiment of this disclosure. The voice interaction device may be a vehicle infotainment system as described in the above embodiment, or it may be a component or assembly within the vehicle infotainment system. The voice interaction device provided in this embodiment can execute the processing flow provided in the voice interaction method embodiment, such as... Figure 6As shown, the voice interaction device 60 includes: an acquisition module 61, a sorting module 62, a selection module 63, and a recognition module 64; wherein, the acquisition module 61 is used to acquire multiple voice information from multiple voice regions, the multiple voice regions including a wake-up voice region and at least one non-wake-up voice region, the wake-up voice region being the voice region where the vehicle's voice function is activated; the sorting module 62 is used to sort the multiple voice information according to a preset voice caching strategy to obtain a wake-up voice region voice queue and at least one non-wake-up voice region voice queue; the selection module 63 is used to select target voice information from the wake-up voice region voice queue and the at least one non-wake-up voice region voice queue according to a preset voice transmission strategy; the recognition module 64 is used to send the target voice information to a recognition engine to obtain the recognition result of the target voice information and complete the voice interaction of the target voice information.
[0098] Optionally, the sorting module 62 is further configured to, for each voice message, determine whether the voice message is in the wake-up sound zone; based on the voice message being in the wake-up sound zone, place the voice message in the wake-up sound zone voice queue; based on the voice message not being in the wake-up sound zone, determine whether the voice message is in at least one non-wake-up sound zone; based on the voice message being in at least one non-wake-up sound zone, place the voice message in the non-wake-up sound zone voice queue corresponding to the voice message; and based on the voice message not being in at least one non-wake-up sound zone, discard the voice message.
[0099] Optionally, the voice queue includes a voice detection start point, voice information, and a voice detection end point;
[0100] The preset sound delivery strategy includes: a wake-up sound zone sound delivery strategy;
[0101] The wake-up sound zone delivery strategy includes: responding to a first speech detection starting point of the wake-up sound zone speech queue, detecting whether the wake-up sound zone speech queue is an empty set; if the wake-up sound zone speech queue is a non-empty set, then the speech information corresponding to the first speech detection starting point is placed into the wake-up sound zone delivery queue; if the wake-up sound zone speech queue is an empty set, and the at least one non-wake-up sound zone is an empty set, then the speech information corresponding to the first speech detection starting point is placed into the wake-up sound zone delivery queue; if the wake-up sound zone speech queue is an empty set, and the at least one non-wake-up sound zone is a non-empty set, then the speech information of the at least one non-wake-up sound zone is cleared, and the speech information corresponding to the first speech detection starting point is placed into the wake-up sound zone delivery queue.
[0102] Optionally, the preset sound delivery strategy further includes: a non-wake-up sound zone sound delivery strategy;
[0103] The non-wake-up voice region delivery strategy includes: responding to the second speech detection starting point of the non-wake-up voice region speech queue, detecting whether the non-wake-up voice region speech queue is an empty set; if the non-wake-up voice region speech queue is an empty set, then arranging the speech information corresponding to the second speech detection starting point into the non-wake-up voice region delivery queue.
[0104] Optionally, the at least one non-wake-up sound zone includes a first non-wake-up sound zone and a second non-wake-up sound zone; the weights of the first non-wake-up sound zone and the second non-wake-up sound zone are obtained according to the order in which the vehicle voice function is responded to; based on the fact that the weight of the first non-wake-up sound zone is greater than the weight of the second non-wake-up sound zone, the priority of the first non-wake-up sound zone is higher than that of the second non-wake-up sound zone.
[0105] Optionally, the recognition module 64 is used to send the target voice information to the recognition engine to obtain the semantic recognition result of the target voice information; control the vehicle to execute the semantic recognition result to complete the voice interaction of the target voice information.
[0106] Figure 6 The voice interaction device shown in the embodiment can be used to execute the technical solution of the above-described voice interaction method embodiment. Its implementation principle and technical effect are similar, and will not be described again here.
[0107] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. The electronic device can be a terminal as described in the above embodiments. The electronic device provided in this embodiment of the present disclosure can execute the processing flow provided in the voice interaction method embodiments, such as… Figure 7 As shown, the electronic device 70 includes: a memory 71, a processor 72, a computer program, and a communication interface 73; wherein the computer program is stored in the memory 71 and configured to be executed by the processor 72 using the voice interaction method described above.
[0108] In addition, this disclosure also provides a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the voice interaction method described in the above embodiments.
[0109] Furthermore, this disclosure also provides a vehicle that includes a voice interaction device as described in the above embodiments; or an electronic device as described in the above embodiments; or a computer-readable storage medium as described in the above embodiments.
[0110] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0111] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A voice interaction method, characterized in that, The method includes: Acquire multiple voice information from multiple voice regions, the multiple voice regions including a wake-up voice region and at least one non-wake-up voice region, the wake-up voice region being the voice region where the vehicle's voice function is activated; The multiple voice information are sorted according to a preset voice caching strategy to obtain a wake-up voice queue and at least one non-wake-up voice queue. According to a preset sound delivery strategy, target speech information is selected from the wake-up speech queue and the at least one non-wake-up speech queue; The target speech information is sent to the recognition engine to obtain the recognition result of the target speech information, and the speech interaction of the target speech information is completed.
2. The method according to claim 1, characterized in that, The step of sorting the multiple voice information according to a preset voice caching strategy to obtain a wake-up voice region voice queue and at least one non-wake-up voice region voice queue includes: For each voice message, determine whether the voice message is in the wake-up sound zone; Based on the voice information in the wake-up sound area, the voice information is arranged in the voice queue of the wake-up sound area; Based on the fact that the voice information is not in the wake-up sound zone, determine whether the voice information is in at least one non-wake-up sound zone; Based on the voice information in at least one non-wake-up voice region, the voice information is placed in the voice queue of the non-wake-up voice region corresponding to the voice information; The voice information is discarded if it is not in at least one non-wake-up voice zone.
3. The method according to claim 1, characterized in that, The voice queue includes a voice detection start point, voice information, and a voice detection end point; The preset sound delivery strategy includes: a wake-up sound zone sound delivery strategy; The wake-up sound zone delivery strategy includes: In response to the first speech detection starting point of the wake-up speech queue, it is detected whether the wake-up speech queue is an empty set; If the wake-up sound zone speech queue is a non-empty set, then the speech information corresponding to the first speech detection starting point is placed into the wake-up sound zone speech queue. If the wake-up sound zone speech queue is an empty set and the at least one non-wake-up sound zone is an empty set, then the speech information corresponding to the first speech detection starting point is placed into the wake-up sound zone speech queue. If the wake-up voice queue is an empty set and the at least one non-wake-up voice region is a non-empty set, then the voice information of the at least one non-wake-up voice region is cleared, and the voice information corresponding to the first voice detection starting point is placed into the wake-up voice region voice transmission queue.
4. The method according to claim 3, characterized in that, The preset sound delivery strategy also includes: a non-wake-up sound zone sound delivery strategy; The non-wake-up voice zone delivery strategy includes: In response to the second speech detection starting point of the non-wake-up speech queue, it is detected whether the non-wake-up speech queue is an empty set; If the non-wake-up voice region speech queue is an empty set, then the speech information corresponding to the second speech detection starting point is placed into the non-wake-up voice region speech delivery queue.
5. The method according to claim 1, characterized in that, The at least one non-wake-up tone region includes a first non-wake-up tone region and a second non-wake-up tone region; The weights of the first and second non-wake-up voice regions are obtained based on the order in which the vehicle's voice function is responded to. Based on the fact that the weight of the first non-wake-up tone is greater than the weight of the second non-wake-up tone region, the priority of the first non-wake-up tone region is higher than that of the second non-wake-up tone region.
6. The method according to claim 1, characterized in that, Sending the target speech information to the recognition engine to obtain the recognition result of the target speech information, and completing the voice interaction of the target speech information, including: The target speech information is sent to the recognition engine to obtain the semantic recognition result of the target speech information; The vehicle is controlled to execute the semantic recognition result and complete the voice interaction of the target voice information.
7. A voice interaction device, characterized in that, The device includes: The acquisition module is used to acquire multiple voice information from multiple voice regions, the multiple voice regions including a wake-up voice region and at least one non-wake-up voice region, the wake-up voice region being the voice region where the vehicle's voice function is activated; The sorting module is used to sort the multiple voice information according to a preset voice caching strategy to obtain a wake-up voice queue and at least one non-wake-up voice queue. The selection module is used to select target speech information from the wake-up speech queue and the at least one non-wake-up speech queue according to a preset speech delivery strategy. The recognition module is used to send the target speech information to the recognition engine, obtain the recognition result of the target speech information, and complete the voice interaction of the target speech information.
8. An electronic device, characterized in that, include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the method as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.
10. A vehicle, characterized in that, include: The voice interaction device as described in claim 7; Or, the electronic device as described in claim 8; Alternatively, the computer-readable storage medium as described in claim 9.
Citation Information
Patent Citations
Method, device and equipment for awakening playing equipment and storage medium
CN112634890A