Voice control methods, devices, systems, equipment, and storage media for in-vehicle devices

By placing microphones in different areas of the vehicle cabin and locating the sound source, the problem of mutual interference between voice control information in multiple areas was solved, enabling independent response and improving the user experience.

CN115985295BActive Publication Date: 2025-12-02HUIZHOU DESAY SV AUTOMOTIVE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211492344.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-25
Publication Date
2025-12-02
Estimated Expiration
2042-11-25

AI Technical Summary

Technical Problem

In multi-zone vehicle cabins, the problem of voice control information interfering with each other leads to a degraded user experience.

Method used

By setting up microphones in different areas of the vehicle cabin, voice control information is collected, and sound source localization technology is used to determine the target area where the sound source of the information is located. The information is then sent to the voice interaction device in the target area for response.

Benefits of technology

This allows for independent voice control in different areas, avoiding interference and improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115985295B_ABST
    Figure CN115985295B_ABST
Patent Text Reader

Abstract

This invention discloses a voice control method, device, system, equipment, and storage medium for in-vehicle devices, applied to in-vehicle control equipment in an in-vehicle voice control system. The method includes: receiving first voice control information collected by at least two microphones; wherein each microphone is respectively located in a different area within the vehicle cabin; performing sound source localization based on the first voice control information collected by each microphone to determine the target area where the sound source of the first voice control information is located; and sending the first voice control information to a voice interaction device corresponding to the area number of the target area, causing the voice interaction device to respond to the first voice control information. This technical solution solves the problem of mutual interference between multi-zone voice control information, enabling independent and non-interfering voice control in different areas, thus improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicles, and more particularly to a voice control method, apparatus, system, device, and storage medium for in-vehicle equipment. Background Technology

[0002] With the development of vehicle technology, voice control in smart cockpits has become widespread, and voice source localization technology has evolved from dual-zone to four-zone, and even multi-zone. Simultaneously, the screen coverage within smart cockpits is increasing, evolving from an instrument panel + central control screen to an instrument panel + central control screen + passenger screen + dual rear headrest screens. This development trend clearly indicates that every passenger in every seat has interaction needs based on both voice and visual input.

[0003] In a multi-zone vehicle cabin, once an interactive device is activated, it generally won't respond to other activation messages. Furthermore, during voice interaction with the device, users may be interrupted by activation messages, thus degrading the user experience. Summary of the Invention

[0004] This invention provides a voice control method, device, system, equipment, and storage medium for in-vehicle devices to solve the problem of mutual interference between voice control information in multiple voice zones, realize independent voice control in different zones without interference, and improve the user experience.

[0005] According to a first aspect of the present invention, a voice control method for an in-vehicle device is provided, applied to an in-vehicle control device in an in-vehicle voice control system, comprising:

[0006] Receive first voice control information collected by at least two microphones; wherein each microphone is located in a different area of ​​the vehicle cabin.

[0007] Based on the first voice control information collected by each of the microphones, the sound source is located to determine the target area where the sound source of the first voice control information is located.

[0008] The first voice control information is sent to the voice interaction device corresponding to the area number of the target area, so that the voice interaction device responds to the first voice control information.

[0009] Furthermore, the step of locating the sound source based on the first voice control information collected by each of the microphones, and determining the target area where the sound source of the first voice control information is located, includes:

[0010] Based on the first voice control information collected by each microphone, the sound source is located to determine the angle of arrival and distance of the sound source point of the first voice control information relative to each microphone.

[0011] The microphone with the smallest angle of arrival or the smallest distance relative to the sound source point is identified as the target microphone;

[0012] Based on the target microphone number, query the mapping relationship between the microphone number and the area number of each area in the vehicle cabin to obtain the area number corresponding to the microphone number of the target microphone.

[0013] The region corresponding to the region number of the target microphone is determined as the target region where the sound source of the first voice control information is located.

[0014] According to a second aspect of the present invention, a voice control device for an in-vehicle device is provided, comprising an in-vehicle control device integrated into an in-vehicle voice control system, including:

[0015] A voice information acquisition module is used to receive first voice control information collected by at least two microphones; wherein each microphone is respectively located in a different area of ​​the vehicle cabin;

[0016] The target area determination module is used to locate the sound source based on the first voice control information collected by each microphone and determine the target area where the sound source point of the first voice control information is located.

[0017] A voice control response module is used to send the first voice control information to the voice interaction device corresponding to the area number of the target area, so that the voice interaction device responds to the first voice control information.

[0018] According to a third aspect of the present invention, an in-vehicle voice control system is provided, comprising: an in-vehicle control device, voice interaction devices respectively disposed in various areas of the vehicle cabin, and a microphone;

[0019] First voice control information is collected by at least two microphones, and the first voice control information collected by each microphone is sent to the vehicle control device.

[0020] The vehicle control device receives the first voice control information collected by each microphone; it performs sound source localization based on the first voice control information collected by each microphone to determine the target area where the sound source of the first voice control information is located; and it sends the first voice control information to the voice interaction device corresponding to the area number of the target area.

[0021] The voice interaction device in the target area responds to the first voice control information.

[0022] According to a fourth aspect of the present invention, an in-vehicle control device is provided, comprising:

[0023] At least one processor; and

[0024] A memory that is communicatively connected to at least one processor; wherein,

[0025] The memory stores a computer program that can be executed by at least one processor, such that the at least one processor is able to execute the voice control method of the vehicle device according to any embodiment of the present invention.

[0026] According to a fifth aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions, the computer instructions being configured to cause a processor to execute and implement the voice control method of the vehicle-mounted device or the voice control method of the vehicle-mounted device according to any embodiment of the present invention.

[0027] The voice control method, apparatus, system, device, and storage medium for in-vehicle devices of this invention receive first voice control information collected by at least two microphones, each microphone being disposed in a different area within the vehicle cabin. Sound source localization is performed based on the first voice control information collected by each microphone to determine the target area where the sound source of the first voice control information is located. The first voice control information is then sent to a voice interaction device corresponding to the area number of the target area, causing the voice interaction device to respond to the first voice control information. By employing the above technical solution, the target area is located based on the voice control information collected by different microphones, and the corresponding voice interaction device is determined. The voice interaction device responds to the voice control information, generating at least one voice avatar within an application operating system. This solves the problem of mutual interference between multi-zone voice control information, enabling independent and non-interfering voice control in different areas, thus improving the user experience.

[0028] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This is a flowchart of a voice control method for an in-vehicle device provided in Embodiment 1 of the present invention;

[0031] Figure 2This is a flowchart of a voice control method for an in-vehicle device provided in Embodiment 2 of the present invention;

[0032] Figure 3 This is a schematic diagram of the structure of a voice control device for an in-vehicle device provided in Embodiment 3 of the present invention;

[0033] Figure 4 This is a system schematic diagram of an in-vehicle voice control system provided in Embodiment 4 of the present invention;

[0034] Figure 5 This is a specific example diagram of an in-vehicle voice control system provided in Embodiment 4 of the present invention;

[0035] Figure 6 This is a schematic diagram of the structure of an in-vehicle control device provided in Embodiment 5 of the present invention. Detailed Implementation

[0036] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0037] It should be noted that the terms "first," "second," "target," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0038] Example 1

[0039] Figure 1 This is a flowchart of a voice control method for an in-vehicle device provided in Embodiment 1 of the present invention. This embodiment is applicable to situations where voice control information from multiple regions is responded to independently. The method can be executed by a voice control device of the in-vehicle device, which can be implemented in hardware and / or software and can be integrated into the in-vehicle control equipment in the in-vehicle voice control system.

[0040] like Figure 1 As shown, the method includes:

[0041] S101, Receive first voice control information collected by at least two microphones respectively; wherein each of the microphones is respectively located in a different area of ​​the vehicle cabin.

[0042] The microphone, or other sound acquisition device, is used to collect voice control information within the cabin and transmit it to the in-vehicle equipment. The vehicle cabin can be divided into multiple areas based on location or function, and a microphone can be installed in each area. At least one microphone can be installed in each of at least two different areas within the vehicle cabin. For example, the vehicle cabin areas include: a driver's seat area, a front passenger seat area, and a rear seat area, with one microphone installed in each area; alternatively, the rear seat area can be further divided into a left rear seat area and a right rear seat area, with one microphone installed in each of these areas. It should be noted that each area also needs to be equipped with corresponding interactive devices. The specific location and number of microphones and interactive devices can be set according to the actual needs of the vehicle; this embodiment does not impose any limitations on this.

[0043] The first voice control information can be understood as the voice control commands collected by the microphone, such as wake-up commands for waking up the vehicle voice control system and commands for interacting with the vehicle voice control system.

[0044] Specifically, microphones are installed in at least two areas within the vehicle cabin. Within the vehicle cabin, the in-vehicle voice control system collects the user's initial voice control information through these microphones located in at least two different areas.

[0045] For example, the first voice control information can be either a wake-up message or a command message. The wake-up message can be determined by the factory settings of the in-vehicle voice control system or can be set according to the user's needs; for example, a preset word such as "Hello" can be used as the wake-up message. The command message includes specific function commands, such as audio playback commands, air conditioning control commands, or seat adjustment commands.

[0046] It is understandable that users in different areas of the vehicle cabin can issue different first voice control information. Therefore, multiple microphones can collect multiple first voice control information and transmit them to the vehicle voice control system.

[0047] For example, the user in the driver's seat area can issue a first voice control message to adjust the seat position, and the user in the passenger seat area can also issue a first voice control message to turn on the air conditioning. The first voice control messages issued by the two different users can be collected and transmitted by the microphone, that is, the vehicle voice control system receives two first voice control messages.

[0048] S102. Based on the first voice control information collected by each of the microphones, perform sound source localization to determine the target area where the sound source point of the first voice control information is located.

[0049] In this embodiment, the sound source point can be understood as the location of the user who issued the first voice control information, and the target area can be understood as the specific area in the vehicle cabin where the sound source point is located, such as the passenger area.

[0050] Specifically, the first voice control information collected by each microphone is located using a sound source localization algorithm to determine the location of the sound source from which the first voice control information is emitted, and the cabin area corresponding to the sound source location is determined as the target area.

[0051] For example, sound source localization algorithms involve the intersection of multiple disciplines, including signal processing, computer technology, biology, pattern recognition, compressed sensing, neural networks, and artificial intelligence. The main research area of ​​sound source localization technology is determining the relative direction and distance of the received signal source to the receiving sensor, i.e., direction estimation (BE) and distance estimation (RE). Direction estimation is sometimes also called direction of arrival (DOA) localization technology.

[0052] Understandably, the microphone can collect multiple different first voice control messages. Accordingly, based on the sound source localization results of the first voice control messages, the target area corresponding to each first voice control message can be determined.

[0053] S103. Send the first voice control information to the voice interaction device corresponding to the area number of the target area, so that the voice interaction device responds to the first voice control information.

[0054] In this embodiment, the voice interaction device can be understood as the application device used to realize voice control information interaction in the vehicle voice control system, and each area is equipped with a corresponding interaction device.

[0055] Specifically, multiple areas within the vehicle cabin are numbered. For example, if there are four cabin areas, they can be numbered using Arabic numerals 1-4, or they can be randomly assigned numbers based on UUIDs (Universally Unique Identifiers). This embodiment does not impose any limitations on this. The target area corresponding to the first voice control information is determined, and the specific content of the first voice control information and the area number corresponding to the target area are sent to the corresponding voice interaction device. The in-vehicle voice control system recognizes the specific command information of the first voice control information and the area number of the target area. The voice interaction device responds accordingly based on the first voice control information, implementing the specific command information of the first voice control information within the target area.

[0056] It is understandable that multiple target areas can be determined based on multiple first voice control information, and the voice interaction devices corresponding to the multiple target areas of the vehicle voice control system can respond to the specific instructions corresponding to the multiple first voice control information respectively.

[0057] In this embodiment, first voice control information collected by at least two microphones is received, with each microphone positioned in a different area within the vehicle cabin. Sound source localization is performed based on the first voice control information collected by each microphone to determine the target area where the sound source of the first voice control information is located. The first voice control information is then sent to the voice interaction device corresponding to the area number of the target area, causing the voice interaction device to respond to the first voice control information. By employing the above technical solution, the target area is located based on the voice control information collected by different microphones, and the corresponding voice interaction device is determined. The voice interaction device responds to the voice control information, generating at least one voice avatar within an application operating system. This solves the problem of mutual interference between multi-zone voice control information, enabling independent and non-interfering voice control in different areas, thus improving the user experience.

[0058] Example 2

[0059] Figure 2 This is a flowchart of a voice control method for an in-vehicle device provided in Embodiment 2 of the present invention. It is a further optimization of any of the above embodiments and can be applied to situations where voice control information from multiple regions is responded to independently. The method can be executed by the voice control device of the in-vehicle device, which can be implemented in hardware and / or software and can be integrated into the in-vehicle control equipment in the in-vehicle voice control system.

[0060] like Figure 2 As shown, the method includes:

[0061] S201, Receive the first voice control information collected by at least two microphones respectively.

[0062] Each microphone is located in a different area within the vehicle's cabin.

[0063] S202. Based on the first voice control information collected by each microphone, perform sound source localization and determine the angle of arrival and distance of the sound source point of the first voice control information relative to each microphone.

[0064] In this embodiment, each microphone can be taken as the central origin. The angle of arrival of the sound source point of the first voice control information relative to each microphone can be understood as the angle between the sound wave ray of the sound source point and the direction of the microphone (horizontal plane or horizontal plane normal); the distance of the sound source point of the first voice control information relative to each microphone can be understood as the distance between the position of the sound source point and the position of the microphone.

[0065] For example, four microphones are installed in the vehicle cabin, namely the first microphone, the second microphone, the third microphone, and the fourth microphone. Based on each microphone as the central origin, the position of the sound source is calculated according to the sound source localization algorithm, and the angle of arrival and distance of the sound source relative to each of the four microphones are determined.

[0066] S203. The microphone with the smallest angle of arrival or the smallest distance relative to the sound source point is identified as the target microphone.

[0067] In this embodiment, the target microphone can be understood as a front-end acquisition device used to perform voice control, and is the microphone with the smallest distance to the sound source point of the first voice control information.

[0068] Specifically, for the same sound source point, multiple microphones collect the first voice control information corresponding to the sound source point. Among the multiple microphones, the microphone closest to the sound source point or with the smallest angle of arrival is determined as the target microphone.

[0069] For example, the distances between the first voice control information sound source point and the first pickup, the second pickup, the third pickup, and the fourth pickup are 0.3m, 0.8m, 0.6m, and 1.2m, respectively. Therefore, it can be determined that the distance between the sound source point and the first pickup is the smallest, which is 0.3m, and the first pickup is determined as the target pickup.

[0070] It is understandable that there can be multiple sound source points corresponding to the first voice control information, each with a different angle of arrival and distance relative to different microphones. If there exists a first sound source point with the smallest distance to the first microphone and a second sound source point with the smallest distance to the second microphone, then the first microphone can be determined as the target microphone relative to the first sound source point, and the second microphone as the target microphone relative to the second sound source point. Therefore, it can be determined that there is at least one target microphone.

[0071] S204. Based on the target microphone number, query the mapping relationship between the microphone number and the area number of each area in the vehicle cabin to obtain the area number corresponding to the microphone number of the target microphone.

[0072] In this embodiment, multiple microphones within the vehicle cabin are assigned corresponding numbers. These microphone numbers can be assigned using Arabic numerals or unique identifiers (UUIDs), and this embodiment does not impose any limitations on this. The target microphone number can be understood as the unique number corresponding to the target microphone.

[0073] Specifically, each area within the vehicle's cabin has a corresponding area number, and there is a one-to-one mapping between each area number and the microphone number. Based on this mapping, the corresponding number of the other area can be determined by querying one area number.

[0074] For example, microphone numbers can be named in the format "s + Arabic numerals", and area numbers can be named in the format "q + Arabic numerals". If the microphone number is set to s1, a corresponding driver's area, q1, exists. Each microphone can have a corresponding cockpit area. When the target microphone is determined to be the one with the number s1, the area number corresponding to s1 can be found through the unique mapping relationship between the two numbers.

[0075] S205. The area corresponding to the area number of the target microphone is determined as the target area where the sound source of the first voice control information is located.

[0076] In this embodiment, the cockpit area corresponding to the target microphone can be determined by querying the area number corresponding to the target microphone number. The target microphone is the microphone closest to the sound source of the first voice control information. Therefore, the cockpit area corresponding to the target microphone can be determined as the target area where the sound source of the first voice control information is located.

[0077] It is understandable that there can be multiple target microphones, and correspondingly, there can also be multiple area numbers corresponding to the microphone numbers of the target microphones, that is, there can be multiple target areas.

[0078] S206. Send the first voice control information to the voice interaction device corresponding to the area number of the target area, so that the voice interaction device responds to the first voice control information.

[0079] In this embodiment, the system receives first voice control information collected by at least two microphones; performs sound source localization based on the first voice control information collected by each microphone to determine the angle of arrival and distance of the sound source point of the first voice control information relative to each microphone; identifies the microphone with the smallest angle of arrival or the smallest distance relative to the sound source point as the target microphone; queries the mapping relationship between the microphone number and the area number of each area in the vehicle cabin based on the target microphone number to obtain the area number corresponding to the target microphone number; identifies the area corresponding to the area number of the target microphone as the target area where the sound source of the first voice control information is located; and sends the first voice control information to the voice interaction device corresponding to the area number of the target area, causing the voice interaction device to respond to the first voice control information. By adopting the above technical solution, the distance between the sound source location of the first voice control information and the microphone is determined using sound source localization, further determining the cabin area where the sound source is located. Different application cabin areas can be determined based on the sound source location of different first voice control information, supporting independent dialogue between multiple users and the voice interaction device of the in-vehicle voice control system. By adopting the above technical solution, the corresponding voice interaction device for each voice control information is determined. The voice interaction device responds to the voice control information separately, which solves the problem of mutual interference between voice control information in multiple voice zones. This enables voice control in different areas to be independent and non-interfering, thereby improving the user experience.

[0080] As a first optional embodiment of this example, based on the above embodiment, this first optional embodiment further optimizes and adds a description of the first voice control information, specifically:

[0081] The first voice control information includes: first voice wake-up information; correspondingly, the voice interaction device responds to the first voice control information by: waking up the voice interaction device set up in the target area by the vehicle voice control system based on the first voice wake-up information.

[0082] In this embodiment, the first voice wake-up information can be understood as the information that first appears within the current time period to wake up the voice interaction device of the vehicle voice control system. The first voice control information may include at least the first voice wake-up information, such as "Hello." The voice interaction device can be understood as a device for displaying and responding to the first voice control information, such as a display screen.

[0083] Specifically, within the vehicle cabin, each area can be equipped with corresponding voice interaction devices. For example, the driver's area has a central control screen, the passenger's area has a passenger-side screen, the left rear seat area has a left rear headrest screen, and the right rear seat area has a right rear headrest screen, etc. Based on the first voice wake-up information, sound source localization is performed. The target area is determined based on the location of the sound source point of the first voice control information. Further, the corresponding voice interaction device within the target area is identified. Based on the first voice wake-up information from the first voice control information, the voice interaction device of the in-vehicle voice control system within the target area is activated.

[0084] The embodiments of the present invention can be based on sound source localization, so that users in different areas can wake up the interactive devices in the corresponding areas respectively, so that the wake-up work of interactive devices in different areas does not interfere with each other.

[0085] Furthermore, after the voice interaction device of the in-vehicle equipment located in the first target area is activated by the first voice wake-up information, it also includes:

[0086] (1) If the second voice wake-up information collected by at least two microphones is received, the sound source is located according to the second voice wake-up information, and the second target area corresponding to the sound source point of the second voice wake-up information is determined.

[0087] In this embodiment, the second voice wake-up information can be understood as the voice wake-up information collected by the microphone when a voice interaction device is activated. The first target area can be understood as the cabin area where the sound source of the first voice wake-up information is located. The second target area can be understood as the cabin area where the sound source of the second voice wake-up information is located.

[0088] Specifically, the first voice control information may further include second voice wake-up information. If a target area is determined based on the first voice wake-up information, the second voice wake-up information can be used to continue enabling voice interaction between the user and the in-vehicle equipment. If at least two microphones collect second voice wake-up information, the second voice wake-up information is located using a sound source localization algorithm. The microphone with the smallest distance and / or angle of arrival to the sound source point of the second voice wake-up information is identified as the target microphone. The second target area corresponding to the sound source point of the second voice wake-up information is determined based on the mapping between the target microphone number and the area number. The first target area and the second target area may overlap to form the same area or may be different areas. Whether the sound source points of the first and second voice wake-up information can be located in the same cabin area is determined according to requirements; this embodiment does not impose any limitations on this.

[0089] (2) If the area number of the second target area is different from the area number of the first target area, the second voice wake-up information is sent to the voice interaction device corresponding to the area number of the second target area to wake up the voice interaction device set in the second target area.

[0090] In this embodiment, the sound source points of the first voice wake-up information and the second voice wake-up information correspond to the first target area and the second target area. If the area numbers of the first target area and the second target area are different, it can be determined that the target cockpit areas corresponding to the first voice wake-up information and the second voice wake-up information are different. Different cockpit areas are equipped with different voice interaction devices. The second voice wake-up information and the corresponding area number of the second target area are sent to the voice interaction device corresponding to the second target area, and the device is woken up by voice according to the wake-up command of the second voice wake-up information.

[0091] It is understandable that the voice interaction devices activated by the first voice wake-up message and the second voice wake-up message are different voice interaction devices in different cockpit areas. Therefore, there may be multiple first voice control messages that wake up different voice interaction devices respectively.

[0092] This invention enables the activation of interactive devices in other areas of a vehicle cabin when an interactive device in one area is activated. Furthermore, when multiple wake-up voice messages from users in different areas are simultaneously collected, the corresponding interactive devices in those areas can be activated concurrently.

[0093] Optionally, the method further includes:

[0094] If there is only one activated voice interaction device in the vehicle cabin, and if at least one microphone collects second voice control information, and the second voice control information is not voice wake-up information, then the second voice control information is sent to the activated voice interaction device so that the activated voice interaction device responds to the second voice control information.

[0095] In this embodiment, the second voice control information can be understood as the voice control information collected by the microphone when there is only one awakened voice interaction device in the vehicle cabin.

[0096] Specifically, if only one voice interaction device is activated within the vehicle cabin, and at least one microphone picks up a second voice control message, and this second voice control message is not a voice wake-up message, then there is no need to locate the sound source of the second voice control message; the message is directly sent to the only activated voice interaction device. This voice interaction device then responds to the second voice control message and executes the corresponding interactive operation.

[0097] For example, if only one interactive device in the vehicle cabin is in a wake-up state, and the microphone collects audio playback information, since the audio playback information is not a wake-up information, there is no need to locate the sound source. The second voice control information collected by the microphone is directly sent to the wake-up interactive device so that the wake-up interactive device responds to the command information, displays music playback information on the screen interface, and plays music.

[0098] In this embodiment of the invention, when only one voice interaction device is activated, if the collected voice control information is not a wake-up command, there is no need to perform sound source localization on the collected voice control information. Therefore, users in any area of ​​the vehicle cabin can interact with the activated voice device. It is understood that if the second voice control information is a wake-up command, which has the potential to wake up other voice interaction devices, it is necessary to determine a further target area through sound source localization before sending the wake-up command to the interaction device in the target area to activate the interaction device in the target area.

[0099] Example 3

[0100] Figure 3 This is a structural schematic diagram of a voice control device for an in-vehicle device provided in Embodiment 3 of the present invention. The voice control device for this in-vehicle device is integrated into the in-vehicle control equipment within the in-vehicle voice control system, such as... Figure 3 As shown, the device includes:

[0101] The voice information acquisition module 31 is used to receive first voice control information collected by at least two microphones respectively; wherein each of the microphones is respectively set in a different area inside the vehicle cabin;

[0102] The target area determination module 32 is used to locate the sound source based on the first voice control information collected by each of the microphones, and determine the target area where the sound source point of the first voice control information is located.

[0103] The voice control response module 33 is used to send the first voice control information to the voice interaction device corresponding to the area number of the target area, so that the voice interaction device responds to the first voice control information.

[0104] By adopting the above technical solution, the target area is located based on the voice control information collected by different microphones, and the corresponding voice interaction device is determined. The voice interaction device responds to the voice control information and generates at least one voice avatar within an application operating system. This solves the problem of mutual interference between voice control information from multiple voice zones, enabling voice control in different areas to be independent and non-interfering, thus improving the user experience.

[0105] Optionally, the target area determination module 32 includes:

[0106] The sound source localization unit is used to perform sound source localization based on the first voice control information collected by each of the microphones, and to determine the angle of arrival and distance of the sound source point of the first voice control information relative to each microphone.

[0107] The pickup determination unit is used to determine the pickup with the smallest angle of arrival or the smallest distance relative to the sound source point as the target pickup;

[0108] The area number determination unit is used to query the mapping relationship between the pickup number and the area number of each area in the vehicle cabin based on the target pickup number of the target pickup, so as to obtain the area number corresponding to the pickup number of the target pickup.

[0109] The target region determination unit is used to determine the region corresponding to the region number of the target microphone as the target region where the sound source of the first voice control information is located.

[0110] Optionally, the first voice control information includes: first voice wake-up information; correspondingly, the voice control response module 33 is specifically applied to: waking up the voice interaction device set up in the target area by the vehicle voice control system based on the first voice wake-up information.

[0111] Optionally, the voice control response module 33 is also specifically applied to:

[0112] If at least two microphones collect second voice wake-up information respectively, then the sound source is located based on the second voice wake-up information to determine the second target area corresponding to the sound source point of the second voice wake-up information;

[0113] If the area number of the second target area is different from the area number of the first target area, the second voice wake-up information is sent to the voice interaction device corresponding to the area number of the second target area to wake up the voice interaction device set in the second target area.

[0114] Optionally, the device is also specifically used in:

[0115] If there is only one activated voice interaction device in the vehicle cabin, and if at least one microphone collects second voice control information, and the second voice control information is not voice wake-up information, then the second voice control information is sent to the activated voice interaction device to respond to the second voice control information.

[0116] The voice control device for vehicle-mounted equipment provided in this embodiment of the invention can execute the voice control method for vehicle-mounted equipment provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0117] Example 4

[0118] Figure 4 This is a system schematic diagram of an in-vehicle voice control system provided in Embodiment 4 of the present invention. Figure 4 As shown, the vehicle voice control system includes: vehicle control equipment 42, voice interaction equipment 43 and microphone 41 respectively installed in various areas of the vehicle cabin;

[0119] First voice control information is collected by at least two microphones 41, and the first voice control information collected by each microphone 41 is sent to the vehicle control device 43.

[0120] The vehicle control device 42 receives the first voice control information collected by each of the microphones 41; it performs sound source localization based on the first voice control information collected by each of the microphones 41 to determine the target area where the sound source of the first voice control information is located; and it sends the first voice control information to the voice interaction device 43 corresponding to the area number of the target area.

[0121] The voice interaction device 43 in the target area responds to the first voice control information.

[0122] The microphone 41 is used to collect various voice control information and transmit it to the vehicle control device 41.

[0123] The vehicle control device 42 is used to control and send the first voice control information to the voice interaction device 43 in the target area;

[0124] The voice interaction device 43 is used to respond to the voice control information collected by the microphone 41, and exists in the form of a display screen.

[0125] In this embodiment, first voice control information is collected by a microphone. The first voice control information is then used to locate the sound source using a sound source localization algorithm to determine the target microphone closest to the sound source. Based on the mapping relationship between the microphone and the cabin area, a target area within the cabin is determined, i.e., the cabin area where the sound source of the first voice control information is located. A corresponding voice interaction device is then determined based on the target area. This voice interaction device responds to the specific commands of the first voice control information and can specifically present the commands of the first voice control information on the corresponding voice interaction device, realizing voice interaction between the user and the vehicle.

[0126] Optional, Figure 5This is a specific example diagram of an in-vehicle voice control system provided in Embodiment 4 of the present invention, as shown below. Figure 5 As shown, the voice interaction device 43 in the vehicle voice control system may specifically include, for example, a central control screen, a passenger-side screen, a left rear headrest screen, and a right rear headrest screen. The vehicle control device 42 in the vehicle voice control system includes: a processor, which may be a system-on-chip (SoC), on which a microphone driver (audio driver), voice services, and voice applications are loaded;

[0127] The pickup driver receives the first voice control information collected by the pickup 41, and then sends the PCM (Pulse Code Modulation) data stream of the first voice control information and the number of the pickup (target pickup) that is closest to the sound source of the first voice control information to the voice service.

[0128] The voice service receives the PCM data stream and target microphone number sent by the audio driver, and sends the speech recognition information of the first voice control information obtained from the PCM data stream, as well as the area number (voice zone number) of the target area determined by the target microphone number, to the voice application.

[0129] The voice application generates a corresponding voice image based on different first voice control information and sends it to the voice interaction device corresponding to each target microphone that collects the first voice control information.

[0130] In this embodiment, DSP (Digital Signal Processing) technology is used to number each microphone 41. After the microphone collects the sound of the first voice control information, it generates a PCM data stream and sends it to the SoC. Each segment of the PCM data stream is distinguished by a corresponding microphone number, and multiple areas in the vehicle cabin also have corresponding area numbers. When multiple users simultaneously wake up the voice using the first voice control information, the voice service listens to the PCM data stream and the microphone number, determines multiple sound-emitting areas as target areas through a sound source localization algorithm, converts the target areas into area numbers, and sends the PCM data stream and area numbers to the vehicle control device 42. The vehicle control device 42 obtains the PCM data stream, performs a voice wake-up algorithm, recognizes the wake-up command, and pops up a voice image on the voice interaction device 43 corresponding to the target area according to the area number to interact with the user. After voice wake-up, multiple users can simultaneously converse with the corresponding voice interaction devices. The commands generated by each conversation execute different action scenarios according to the area number.

[0131] For example, the first microphone acquires the first voice control information "play music," parses the "play music" command information using DSP technology, generates a PCM data stream, and sends it to the vehicle control device 42. Based on sound source localization, the first microphone is identified as the target microphone. According to the first voice control information acquired by the target microphone, the processor processes the information, first receiving the PCM data stream of the first voice control information and the target microphone number, and then sending it to the voice service via audio driver. Simultaneously, the cabin area corresponding to the target microphone is identified as the target area. The audio driver sends the area number of the target area and the voice recognition information of the first voice control information to the voice application. The voice application uses the voice interaction device 43 corresponding to the area number of the target area. For example, if the target area is the driver's area, the corresponding voice interaction device 43 is the central control screen. The central control screen interface displays voice icon 1, i.e., "play music," and responds based on the "play music" voice command, realizing voice interaction between the user and the vehicle.

[0132] In this embodiment, at least two microphones respectively collect first voice control information, and the first voice control information collected by each microphone is sent to the vehicle control device; the vehicle control device receives the first voice control information collected by each microphone; sound source localization is performed based on the first voice control information collected by each microphone to determine the target area where the sound source of the first voice control information is located; the first voice control information and the corresponding area number of the target area are sent to the voice interaction device; the voice interaction device in the target area responds to the first voice control information. By adopting the above technical solution, combined with voice source localization technology, multiple voice avatars are generated in one operating system and projected onto multiple voice interaction devices, enabling simultaneous multi-screen dialogue that is independent and does not interfere with each other. This allows for timely response to voice interactions from multiple locations, effectively improving the user experience.

[0133] Example 5

[0134] Figure 6 A schematic diagram of an in-vehicle control device 42, which can be used to implement embodiments of the present invention, is shown. It is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, and other suitable computers. The in-vehicle control device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0135] like Figure 6 As shown, the vehicle control device 42 includes at least one processor 51 and a memory, such as a read-only memory (ROM) 52 or a random access memory (RAM) 53, communicatively connected to the at least one processor 51. The memory stores computer programs executable by the at least one processor. The processor 51 can perform various appropriate actions and processes based on the computer program stored in the ROM 52 or loaded from storage unit 58 into the RAM 53. The RAM 53 can also store various programs and data required for the operation of the vehicle control device 42. The processor 51, ROM 52, and RAM 53 are interconnected via a bus 54. An input / output (I / O) interface 55 is also connected to the bus 54.

[0136] Multiple components in the vehicle control device 42 are connected to the I / O interface 55, including: an input unit 56, such as a keyboard, mouse, etc.; an output unit 57, such as various types of displays, speakers, etc.; a storage unit 58, such as a disk, optical disk, etc.; and a communication unit 59, such as a network card, modem, wireless transceiver, etc. The communication unit 59 allows the vehicle control device 42 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0137] Processor 51 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 51 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 51 performs the various methods described above, such as voice control methods for in-vehicle devices.

[0138] In some embodiments, the voice control method of the in-vehicle device may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 58. In some embodiments, part or all of the computer program may be loaded and / or installed on the in-vehicle control device 42 via ROM 52 and / or communication unit 59. When the computer program is loaded into RAM 53 and executed by processor 51, one or more steps of the voice control method of the in-vehicle device described above may be performed. Alternatively, in other embodiments, processor 51 may be configured to perform the voice control method of the in-vehicle device by any other suitable means (e.g., by means of firmware).

[0139] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0140] Computer programs for implementing the voice control method of the in-vehicle device of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs can be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0141] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0142] To provide interaction with the user, the systems and techniques described herein can be implemented on an in-vehicle control device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the in-vehicle control device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0143] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0144] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0145] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0146] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A voice control method for an in-vehicle device, characterized in that, An in-vehicle control device applied in an in-vehicle voice control system, the method comprising: The system receives first voice control information collected by at least two microphones; wherein each microphone is located in a different area of ​​the vehicle cabin, and the first voice control information includes: first voice wake-up information and / or command information. Based on the first voice control information collected by each of the microphones, the sound source is located to determine the target area where the sound source of the first voice control information is located. The first voice control information is sent to the voice interaction device corresponding to the area number of the target area, so that the voice interaction device responds to the first voice control information. The voice interaction device responds to the first voice control information by: waking up the voice interaction device set up in the target area by the vehicle voice control system based on the first voice wake-up information; After the voice interaction device of the in-vehicle equipment set in the first target area is awakened by the first voice wake-up information, it also includes: If at least two microphones collect the second voice wake-up information respectively, then the sound source is located based on the second voice wake-up information to determine the second target area corresponding to the sound source point of the second voice wake-up information; If the area number of the second target area is different from the area number of the first target area, the second voice wake-up information is sent to the voice interaction device corresponding to the area number of the second target area to wake up the voice interaction device set in the second target area.

2. The method according to claim 1, characterized in that, The step of locating the sound source based on the first voice control information collected by each of the microphones, and determining the target area where the sound source of the first voice control information is located, includes: Based on the first voice control information collected by each microphone, the sound source is located to determine the angle of arrival and distance of the sound source point of the first voice control information relative to each microphone. The microphone with the smallest angle of arrival or the smallest distance relative to the sound source point is identified as the target microphone. Based on the target microphone number, query the mapping relationship between the microphone number and the area number of each area in the vehicle cabin to obtain the area number corresponding to the microphone number of the target microphone. The region corresponding to the region number of the target microphone is determined as the target region where the sound source of the first voice control information is located.

3. The method according to claim 1, characterized in that, Also includes: If there is only one activated voice interaction device in the vehicle cabin, and if at least one microphone collects second voice control information, and the second voice control information is not voice wake-up information, then the second voice control information is sent to the activated voice interaction device so that the activated voice interaction device responds to the second voice control information.

4. A voice control device for vehicle-mounted equipment, characterized in that, The vehicle control device integrated into the vehicle voice control system includes: A voice information acquisition module is used to receive first voice control information collected by at least two microphones respectively; wherein each microphone is respectively set in a different area in the vehicle cabin, and the first voice control information includes: first voice wake-up information and / or command information; The target area determination module is used to locate the sound source based on the first voice control information collected by each microphone and determine the target area where the sound source point of the first voice control information is located. A voice control response module is used to send the first voice control information to a voice interaction device corresponding to the area number of the target area, so that the voice interaction device responds to the first voice control information. The voice control response module is specifically used for: The vehicle voice control system activates the voice interaction device set up in the target area based on the first voice wake-up information. The voice control response module is also specifically used in: If at least two microphones collect the second voice wake-up information respectively, then the sound source is located based on the second voice wake-up information to determine the second target area corresponding to the sound source point of the second voice wake-up information; If the area number of the second target area is different from the area number of the first target area, the second voice wake-up information is sent to the voice interaction device corresponding to the area number of the second target area to wake up the voice interaction device set in the second target area.

5. The apparatus according to claim 4, characterized in that, The target region determination module includes: The sound source localization unit is used to perform sound source localization based on the first voice control information collected by each of the microphones, and to determine the angle of arrival and distance of the sound source point of the first voice control information relative to each microphone. The pickup determination unit is used to determine the pickup with the smallest angle of arrival or the smallest distance relative to the sound source point as the target pickup; The area number determination unit is used to query the mapping relationship between the pickup number and the area number of each area in the vehicle cabin based on the target pickup number of the target pickup, so as to obtain the area number corresponding to the pickup number of the target pickup. The target region determination unit is used to determine the region corresponding to the region number of the target microphone as the target region where the sound source of the first voice control information is located.

6. A vehicle-mounted voice control system, characterized in that, include: In-vehicle control equipment, voice interaction devices and microphones installed in various areas of the vehicle cabin; First voice control information is collected by at least two microphones, and the first voice control information collected by each microphone is sent to the vehicle control device. The vehicle control device receives first voice control information collected by each microphone, the first voice control information including: first voice wake-up information and / or command information; performs sound source localization based on the first voice control information collected by each microphone to determine the target area where the sound source of the first voice control information is located; sends the first voice control information to the voice interaction device corresponding to the area number of the target area; the voice interaction device responds to the first voice control information by: waking up the voice interaction device set up in the target area by the first voice wake-up information; after the voice interaction device of the vehicle device set up in the first target area is woken up by the first voice wake-up information, the device further includes: if it receives second voice wake-up information collected by at least two microphones respectively, it performs sound source localization based on the second voice wake-up information to determine the second target area corresponding to the sound source of the second voice wake-up information; if the area number of the second target area is different from the area number of the first target area, it sends the second voice wake-up information to the voice interaction device corresponding to the area number of the second target area to wake up the voice interaction device set up in the second target area. The voice interaction device in the target area responds to the first voice control information.

7. A vehicle-mounted control device, characterized in that, The vehicle-mounted control device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the voice control method of the vehicle-mounted device according to any one of claims 1-3.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the voice control method of the vehicle-mounted device according to any one of claims 1-3.

Citation Information

Patent Citations

  • Multi-screen voice interaction method and device of vehicle-mounted system, storage medium and vehicle machine

    CN109493871A

  • Voice wake-up processing method and device, and storage medium

    CN109841214A