In-vehicle voice interaction control method, system and vehicle
By using the front and rear host architecture and multi-microphone array to identify the user's position, the problem of in-car partition identification and control is solved, balanced output and independent operation of multiple sound zones in the car are achieved, the convenience and intelligence of voice interaction are improved, and hardware costs are reduced.
Patent Information
- Application Number
- CN202111036246.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-06
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2041-09-06
AI Technical Summary
Existing technologies are unable to achieve true partition recognition and control within the vehicle, resulting in a single voice control function. In particular, it is difficult to control the power system with high safety requirements, and the timeliness is low. Business control is limited, and the system cannot be effectively controlled when directly connected to the voice business.
It adopts a front and rear host architecture, uses a multi-microphone array to identify the user's position of voice commands, and distributes the voice commands to the corresponding host for processing based on the position recognition results. The front host uniformly processes vehicle control commands, and the rear host processes non-vehicle control commands. Ethernet paths are used for communication transmission to achieve rapid response and independent control.
It achieves balanced output and independent operation of multiple sound zones in the car, reduces hardware costs, improves voice response speed, reduces bus burden, is compatible with the front single screen solution without software changes, and improves the convenience and intelligence of voice interaction.
Smart Images

Figure CN115762501B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automobile voice control, and in particular to a control method and system for in-vehicle voice interaction and an automobile. Background Art
[0002] With the development of voice technology, voice interaction has become a mainstream form of in-car interaction. As voice control becomes increasingly complex and its scope of control expands, the demand for voice-controlled zones is growing stronger among different users. Currently, the industry's voice-controlled zones only offer body control functions. In particular, the current mainstream four-zone solution only supports a single voice assistant and only supports a single screen in the front row. For rear-seat users, zoned zones only support body control functions, while other functions cannot be identified or controlled. This represents a "pseudo" four-zone solution, making multiple zones of voice control meaningless. As a result, voice control functionality in the automotive industry is limited, limited to controlling certain functions such as seats and air conditioning. Controlling powertrains, which require high safety standards, is particularly difficult to implement, resulting in a lack of convenience and intelligence.
[0003] Currently, the industry mainly uses a single front-row voice engine + voice assistant (no voice assistant in the back row) to perform voice wake-up, recognition, and semantic understanding, and then distribute commands to the front and rear screens for operation and display. This is equivalent to a single front-row voice command driving four screens, and many front and rear commands need to be processed in parallel. If all rear-row voice commands are recognized and distributed by the first host, and then sent to the front host system, the commands are transmitted between the front and rear host systems, and then responded to by the rear host system for business control and display, this will lead to: 1) low timeliness; 2) business control limitations. The system directly connects to voice services such as voice control of music, lacks interfaces, and cannot effectively control any user commands. Summary of the Invention
[0004] The purpose of the present invention is to propose a control method, system and vehicle for in-vehicle voice interaction, so as to solve the technical problem that existing methods cannot truly realize in-vehicle partition identification and control.
[0005] In one aspect, a method for controlling in-vehicle voice interaction is provided, which is applied to interactive control of an in-vehicle voice assistant, comprising:
[0006] The first host obtains the user's voice command and recognizes the semantics of the voice command, and simultaneously recognizes the location of the user who issued the voice command to obtain a location recognition result;
[0007] Determining the type of the voice command based on the semantic recognition result of the voice command, and processing the voice command by the first host when the voice command is a vehicle control command; and sending the voice command to the first host or the second host for processing when the voice command is not a vehicle control command based on the position recognition result; and collecting the operation command fed back by the user through the first host or the second host;
[0008] The first host responds to the operation instruction fed back by the user and recognizes the position of the terminal input according to the operation instruction and the position result of the voice instruction.
[0009] Preferably, obtaining the user's voice command includes:
[0010] Voice commands inside the car are collected by a multi-microphone array installed in the car, wherein the multi-microphone array forms multiple sound pickup zones in the car, and the sound pickup zones correspond to the seating positions in the car. Each sound zone separately collects the sound signals within the sound zone and blocks the sound signals of other sound zones.
[0011] Preferably, identifying the location of the user issuing the voice command includes:
[0012] Identify the multi-microphone array corresponding to the voice command collection, and determine the corresponding seating position based on the corresponding multi-microphone array;
[0013] The seating position corresponding to the voice command is set as the position of the user who issued the voice command, and a position recognition result is obtained.
[0014] Preferably, sending the voice command to the first host or the second host for processing according to the position recognition result includes:
[0015] If the position recognition result is the front row, the first host processes the voice command;
[0016] If the position recognition result is the back row, the first host sends the voice command to the second host via Ethernet for processing.
[0017] Preferably, sending the voice command to the first host for processing includes:
[0018] When the first host processes the voice command, the semantics of the voice command is forwarded to the central control terminal; the first host controls the central control terminal to independently input or display;
[0019] The sending of the voice command to the second host for processing includes:
[0020] When the second host receives the voice command, it forwards the semantics of the voice command to the rear-row display terminal corresponding to the position recognition result; the second host controls multiple rear-row display terminals at the same time, and controls the rear-row display terminals to input or display independently.
[0021] Preferably, performing TTS broadcasting to users at corresponding locations includes:
[0022] Provide TTS broadcasts to users in corresponding positions and control the broadcast volume to achieve balanced output in multiple sound zones in the car.
[0023] On the other hand, a control system for in-vehicle voice interaction is also provided, which is used to implement the control method for in-vehicle voice interaction, including:
[0024] The first host is configured to obtain a user's voice command and recognize the semantics of the voice command, and simultaneously recognize the location of the user issuing the voice command to obtain a location recognition result; determine the type of the voice command based on the semantic recognition result, and when the voice command is a vehicle control command, the first host processes the command; when the voice command is not a vehicle control command, send the voice command to the first host or the second host for processing based on the location recognition result, and simultaneously collect the operation command fed back by the user; and, in response to the operation command fed back by the user, perform a TTS broadcast to the user at the corresponding location and control the broadcast volume based on the terminal location of the operation command input and the location recognition result of the voice command, so as to achieve balanced output of multiple sound zones in the vehicle;
[0025] The second host is used to receive and process the voice instructions sent by the first host, and collect operation instructions fed back by the user.
[0026] Preferably, the first host is further configured to collect voice commands within the vehicle using a multi-microphone array disposed within the vehicle, identify the multi-microphone array corresponding to the voice command collection, and determine a corresponding seating position based on the corresponding multi-microphone array; set the seating position corresponding to the voice command as the position of the user issuing the voice command to obtain a position recognition result; wherein the multi-microphone array forms a plurality of sound pickup zones within the vehicle, each of which corresponds to a seating position within the vehicle, and each zone separately collects sound signals within the zone and shields sound signals from other zones;
[0027] And when the position recognition result is the front row, the first host processes the voice command and forwards the semantics of the voice command to the central control terminal; the first host controls the central control terminal to independently input or display; when the position recognition result is the back row, the first host sends the voice command to the second host via Ethernet for processing.
[0028] Preferably, the second host is also used to forward the semantics of the voice command to the rear display terminal corresponding to the position recognition result when receiving the voice command; the second host simultaneously controls multiple rear display terminals and controls the rear display terminals to input or display independently.
[0029] On the other hand, a car is also provided, which controls the screen display, permission management and partition control of multiple sound zones in the car by the in-car voice assistant through the control system of the in-car voice interaction.
[0030] In summary, the implementation of the embodiments of the present invention has the following beneficial effects:
[0031] The control method, system and automobile for in-vehicle voice interaction provided by the present invention, through the pioneering front and rear host architecture, voice recognition and TTS playback are uniformly distributed by the front host, and the rear host receives front instructions to control rear business; it can break through the technical limitation of insufficient driving force of a single voice engine, achieve rapid response, and uniformly perform recognition and distribution and TTS broadcasting by the front host; the rear host does not need to add hardware noise reduction, engine processing, DSP chip and connect to the whole vehicle speakers, which can greatly save costs.
[0032] At the same time, vehicle control commands are uniformly processed by the front host, reducing bus load and enabling compatibility with single-screen solutions without requiring any software changes. The front and rear hosts communicate via a private Ethernet interface, eliminating the need for system interaction and reducing command transmission time. After receiving commands from the front host, the rear host, driven by the same rear host, enables independent voice, image, and user-specific display, as well as application control, through the left and right rear screens, without disrupting each other. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, without paying any creative work, other drawings obtained based on these drawings still fall within the scope of the present invention.
[0034] Figure 1 This is a schematic diagram of the main process of a control method for in-vehicle voice interaction in an embodiment of the present invention.
[0035] Figure 2 This is a logical diagram of a control method for in-vehicle voice interaction in an embodiment of the present invention.
[0036] Figure 3The figure is a schematic diagram of the architecture of a control system for in-vehicle voice interaction according to an embodiment of the present invention. DETAILED DESCRIPTION
[0037] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention will be described in further detail below with reference to the accompanying drawings.
[0038] like Figure 1 and Figure 2 FIG2 is a schematic diagram of an embodiment of a method for controlling in-vehicle voice interaction provided by the present invention. In this embodiment, the method is applied to interactive control of an in-vehicle voice assistant, and includes the following steps:
[0039] The first host obtains the user's voice commands and recognizes the semantics of the voice commands, while identifying the location of the user who issued the voice commands to obtain the location recognition result; the first host, also known as the front host, collects the voice commands through the microphone array installed in the car, and performs semantic recognition and sound source positioning, determines the direction of the beam, and specifically identifies which direction the voice command is input from, front, back, left, or right.
[0040] In a specific embodiment, voice commands are collected using a multi-microphone array installed within the vehicle. The multi-microphone array forms multiple pickup zones within the vehicle, corresponding to seating positions within the vehicle. Each zone independently captures the sound signals within its own zone and blocks the sound signals from other zones. Once the collection is complete, the multi-microphone array corresponding to the voice command is identified, and the corresponding seating position is determined based on the multi-microphone array. The seating position corresponding to the voice command is then set as the position of the user issuing the voice command, resulting in a position recognition result.
[0041] Furthermore, after collecting the voice commands, semantic analysis (semantic recognition) and semantic distribution of all voice commands are carried out, and the commands are classified and processed at the same time. That is, TTS (Text To Speech) is used for processing, which is part of the human-computer dialogue and enables the machine to speak. With the support of the built-in chip, the text is intelligently converted into a natural speech stream through the design of the neural network. TTS technology converts text files in real time, and the conversion time can be calculated in seconds. Under the action of its unique intelligent voice controller, the voice output of the text has a smooth rhythm, making the listener feel natural when listening to the information, without the indifference and awkwardness of machine voice output. TTS speech synthesis technology will soon cover the first and second level Chinese characters of the national standard, have an English interface, automatically recognize Chinese and English, and support mixed reading of Chinese and English.
[0042] The type of the voice command is determined based on the semantic recognition result of the voice command. When the type of the voice command is a vehicle control command, it is processed by the first host; when the type of the voice command is not a vehicle control command, the voice command is sent to the first host or the second host for processing based on the position recognition result; and the operation command of the user feedback is collected by the first host or the second host; the semantic recognition of the above steps can distinguish the type of voice command, and it can be determined whether the voice command is a vehicle control command or a non-vehicle control command. Here, the vehicle control command refers to the command to control multiple body parts such as windows, seats, power gears, air conditioners, etc. in the car; and the non-vehicle control command mainly refers to the command to control the multi-screen display system, music, navigation, weather games and other applications in the car. Specifically, the second host, that is, the rear host, communicates and transmits through the Ethernet path private interface (first host→second host), without going through the system (first host→front host system→rear host system→second host), which can reduce the command transmission time.
[0043] In a specific embodiment, when determining the voice command issuance status based on the voice command type, all vehicle control commands are uniformly processed by the front host. Since the front host itself is developed with the CAN protocol, and if the front and rear hosts process simultaneously, more bus signals need to be added, which increases the bus burden. Forwarding to the rear host requires adding front and rear host system interaction commands. The rear host then converts them into CAN commands and sends them to each vehicle body-related node via the CAN protocol for control, resulting in response delays and affecting the interactive experience. Therefore, for vehicle control commands, regardless of front and rear rows, the front host performs unified parsing and CAN command issuance. This not only saves steps but is also compatible with the front single screen configuration solution. For this configuration, the vehicle control partition command control does not require any software changes.
[0044] Furthermore, for non-vehicle control semantic instructions, further judgment is required. The position of the user who issued the instruction can determine whether it is processed by the front-row host or the rear-row host. If the position recognition result is the front row, the first host processes the voice instruction; if the position recognition result is the rear row, the first host sends the voice instruction to the second host via Ethernet for processing. In other words, for voice instructions other than vehicle control, if they are voice instructions from the front-row user, they are sent to the front-row host for processing. If they are rear-row voice instructions, the semantic results are transmitted to the rear-row host via Ethernet TCP / IP protocol. The Ethernet interface is directly used between the front and rear hosts for voice instruction transmission. The real-time transmission efficiency is high and the delay can be reduced to a level of only 10ms. The rear-seat voice response time is shortened, which can greatly improve the timely response accuracy of the voice typewriter.
[0045] Furthermore, when the first host processes the voice command, it forwards the semantics of the voice command to the central control terminal; the first host controls the central control terminal to independently input or display. When the second host receives the voice command, it forwards the semantics of the voice command to the rear-row display terminal corresponding to the position recognition result; the second host controls multiple rear-row display terminals at the same time, and controls the rear-row display terminals to independently input or display. In other words, after the rear-row host receives the voice command from the front row, it needs to forward the semantics according to the direction of the received voice command, such as the left rear. It should be noted that, if Figure 3 As shown, the rear host is based on the dual-application technology deeply developed by the Android system, and forwards it to the user IDs of the identified left and right screens. Dual-application support is for the Android system to support multi-user operation, that is, one host (rear host) only has one Android operating system that needs to drive two screens (rear left screen, rear right screen), and support independent voice input and UI display for the two screens. That is, after the rear host is powered on, it needs to identify the IDs of the main user and secondary user of the system. The main user is the first user ID of the device (left screen, USER 0), and the secondary user ID (right screen, USER 1), distinguish the instructions to the rear left or rear right, and distribute them to the partition application for control.
[0046] Specifically, after the same application in the rear row, such as (music), is opened in pairs, two independent applications are formed on the left and right screens for independent display and operation, which are isolated from each other and run independently. After receiving the voice command of the partition forwarded from the front row, the control response of the corresponding partition will be performed. For example, if the left rear user wakes up the command, the rear host will start the voice assistant to display on the display screen in the corresponding direction (and needs to judge the screen status. If the screen is off, it needs to light up the screen). After receiving the continuous voice input from the user in the front row in real time, the voice recognition content is presented in real time in the form of a typewriter, and the application, such as (music), is called according to the recognition and understanding results to achieve independent display and application response according to the user screen in that direction.
[0047] The first host responds to user feedback and, based on the location of the terminal where the input command was entered and the location of the voice command, broadcasts TTS to the user in the corresponding position and controls the broadcast volume, achieving balanced output across multiple audio zones within the vehicle. As can be seen from the above, the rear host independently controls the two rear screens to capture user input commands; the front host independently controls the central control screen to capture user input commands from the front row. Combined with voice commands, these commands can determine the desired action for the user in the corresponding position, such as opening a music app or controlling the volume in that zone. This allows the first host, also known as the front host, to uniformly distribute commands based on the identified needs, providing unified control of the vehicle's audio zones or functions. This ensures that TTS is uniformly output by the front host, while all voice interaction audio is output from the vehicle's speakers. Multi-channel broadcasting automatically adjusts the sound field balance based on sound source location and on-screen services, achieving control of the four sound fields and volume within the vehicle to meet the hearing feedback needs of the user in that position. It can reduce the number of rear host DSP chips and the need to connect the entire vehicle speaker hardware, reducing hardware and wiring harness costs; and control instructions for permission processing. When unsupported services are triggered in the back row, TTS is required to remind the user that it is not supported and guide the user, and give a unified response to the content that is not supported by the rear screen.
[0048] like Figure 3 As shown, the present invention also provides a schematic diagram of an embodiment of a control system for in-vehicle voice interaction. In this embodiment, the system is used to implement the control method for in-vehicle voice interaction, including:
[0049] The first host is configured to obtain a user's voice command and recognize the semantics of the voice command, and simultaneously recognize the location of the user issuing the voice command to obtain a location recognition result; determine the type of the voice command based on the semantic recognition result, and when the voice command is a vehicle control command, the first host processes the command; when the voice command is not a vehicle control command, send the voice command to the first host or the second host for processing based on the location recognition result, and simultaneously collect the operation command fed back by the user; and, in response to the operation command fed back by the user, perform a TTS broadcast to the user at the corresponding location and control the broadcast volume based on the terminal location of the operation command input and the location recognition result of the voice command, so as to achieve balanced output of multiple sound zones in the vehicle;
[0050] In a specific embodiment, the first host is further configured to collect voice commands within the vehicle using a multi-microphone array disposed within the vehicle, identify the multi-microphone array corresponding to the voice command collection, and determine a corresponding seating position based on the corresponding multi-microphone array; set the seating position corresponding to the voice command as the position of the user issuing the voice command, thereby obtaining a position recognition result; wherein the multi-microphone array forms a plurality of sound pickup zones within the vehicle, each of which corresponds to a seating position within the vehicle, and each zone separately collects sound signals within the zone and shields sound signals from other zones;
[0051] And when the position recognition result is the front row, the first host processes the voice command and forwards the semantics of the voice command to the central control terminal; the first host controls the central control terminal to independently input or display; when the position recognition result is the back row, the first host sends the voice command to the second host via Ethernet for processing.
[0052] The second host is used to receive and process the voice instructions sent by the first host, and collect operation instructions fed back by the user.
[0053] In a specific embodiment, the second host is also used to forward the semantics of the voice command to the rear-row display terminal corresponding to the position recognition result when receiving the voice command; the second host simultaneously controls multiple rear-row display terminals and controls the rear-row display terminals to input or display independently.
[0054] The present invention also provides a car, which controls the screen display, authority management and partition control of multiple sound zones in the car by an in-car voice assistant through the control system of the in-car voice interaction.
[0055] Regarding the control system of the in-vehicle voice interaction and the specific implementation process of the car, please refer to the implementation process of the above-mentioned in-vehicle voice interaction control method, which will not be repeated here.
[0056] In summary, the implementation of the embodiments of the present invention has the following beneficial effects:
[0057] The control method, system and automobile for in-vehicle voice interaction provided by the present invention, through the pioneering front and rear host architecture, voice recognition and TTS playback are uniformly distributed by the front host, and the rear host receives front instructions to control rear business; it can break through the technical limitation of insufficient driving force of a single voice engine, achieve rapid response, and uniformly perform recognition and distribution and TTS broadcasting by the front host; the rear host does not need to add hardware noise reduction, engine processing, DSP chip and connect to the whole vehicle speakers, which can greatly save costs.
[0058] At the same time, vehicle control commands are uniformly processed by the front host, reducing bus load and enabling compatibility with single-screen solutions without requiring any software changes. The front and rear hosts communicate via a private Ethernet interface, eliminating the need for system interaction and reducing command transmission time. After receiving commands from the front host, the rear host, driven by the same rear host, enables independent voice, image, and user-specific display, as well as application control, through the left and right rear screens, without disrupting each other.
[0059] The above disclosure is merely a preferred embodiment of the present invention and certainly cannot be used to limit the scope of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope of the present invention.
Claims
1. A control method for in-vehicle voice interaction, applied to the interactive control of an in-vehicle voice assistant, characterized in that: include: The first host obtains the user's voice command and recognizes the semantics of the voice command, and simultaneously recognizes the location of the user who issued the voice command to obtain a location recognition result; the first host is the front-row host; Determining the type of the voice command based on the semantic recognition result of the voice command, and processing the voice command by the first host when the voice command is a vehicle control command; and sending the voice command to the first host or the second host for processing when the voice command is not a vehicle control command based on the position recognition result; and collecting the operation command fed back by the user through the first host or the second host; The second host is the rear host; The sending of the voice command to the first host or the second host for processing according to the position recognition result includes: If the position recognition result is the front row, the first host processes the voice command; If the position recognition result is the back row, the first host sends the voice command to the second host via Ethernet for processing; The first host responds to the operation instruction fed back by the user and performs TTS broadcast to the user at the corresponding location according to the terminal location input by the operation instruction and the location recognition result of the voice instruction.
2. The method according to claim 1, wherein The obtaining of the user's voice command comprises: Voice commands inside the car are collected by a multi-microphone array installed in the car, wherein the multi-microphone array forms multiple sound pickup zones in the car, and the sound pickup zones correspond to the seating positions in the car. Each sound zone separately collects the sound signals within the sound zone and blocks the sound signals of other sound zones.
3. The method according to claim 2, wherein The identifying the location of the user issuing the voice command includes: Identify the multi-microphone array corresponding to the voice command collection, and determine the corresponding seating position based on the corresponding multi-microphone array; The seating position corresponding to the voice command is set as the position of the user who issued the voice command, and a position recognition result is obtained.
4. The method according to claim 3, wherein The sending of the voice command to the first host for processing includes: When the first host processes the voice command, the semantics of the voice command is forwarded to the central control terminal; the first host controls the central control terminal to independently input or display; The sending of the voice command to the second host for processing includes: When the second host receives the voice command, it forwards the semantics of the voice command to the rear-row display terminal corresponding to the position recognition result; the second host controls multiple rear-row display terminals at the same time, and controls the rear-row display terminals to input or display independently.
5. The method according to claim 1, wherein The performing of TTS broadcasting to users at corresponding locations includes: Provide TTS broadcasts to users in corresponding positions and control the broadcast volume to achieve balanced output in multiple sound zones in the car.
6. A control system for in-vehicle voice interaction, used to implement the method according to any one of claims 1 to 5, characterized in that: include: The first host is configured to obtain a user's voice command and recognize the semantics of the voice command, and simultaneously recognize the location of the user issuing the voice command to obtain a location recognition result; determine the type of the voice command based on the semantic recognition result of the voice command; if the voice command is a vehicle control command, the first host processes the command; if the voice command is not a vehicle control command, the first host sends the voice command to the first host or the second host for processing based on the location recognition result, and simultaneously collects operation instructions fed back by the user; In response to the user's feedback, the system performs TTS announcements to the user at the corresponding location and controls the volume of the announcements based on the terminal location where the operation instruction was input and the position recognition result of the voice command, thereby achieving balanced output in multiple sound zones within the vehicle. The first host is a front-row host; Wherein, if the position recognition result is the front row, the first host processes the voice command and forwards the semantics of the voice command to the central control terminal; the first host controls the central control terminal to independently input or display; if the position recognition result is the back row, the first host sends the voice command to the second host via Ethernet for processing; The second host is used to receive and process the voice instructions sent by the first host, and collect the operation instructions fed back by the user; the second host is the rear host.
7. The system according to claim 6, wherein: The first host is also used to collect voice commands in the car through a multi-microphone array set in the car, identify the multi-microphone array corresponding to the voice command collection, and determine the corresponding seating position according to the corresponding multi-microphone array; set the seating position corresponding to the voice command to the position of the user who issued the voice command to obtain a position recognition result; wherein, the multi-microphone array forms a plurality of sound pickup zones in the car, and the sound pickup zones correspond to the seating positions in the car, and each sound zone separately collects the sound signals within the sound zone and shields the sound signals of other sound zones.
8. The system according to claim 6, wherein: The second host is also used to forward the semantics of the voice command to the rear-row display terminal corresponding to the position recognition result when receiving the voice command; the second host simultaneously controls multiple rear-row display terminals and controls the rear-row display terminals to input or display independently.
9. An automobile, characterized in that: The in-vehicle voice interaction control system as described in any one of claims 6-8 controls the screen display, permission management and partition control of multiple audio zones in the car by the in-vehicle voice assistant.
Citation Information
Patent Citations
Voice control method, terminal and computer readable storage medium
CN110992946A
Vehicle-based voice processing method, voice processor and vehicle-mounted processor
CN112599133A