Vehicle-mounted voice interaction method, device, system, electronic device and storage medium

CN116206605BActive Publication Date: 2026-09-04IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211734714.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2026-09-04
Estimated Expiration
2042-12-30

AI Technical Summary

Technical Problem

由于切换过程以及非控制模式下的各种操作均需要用户手动执行,执行规则繁琐难记,大大降低用户体验

Benefits of technology

[0034] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the in-vehicle voice interaction method as described above.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116206605B_ABST
    Figure CN116206605B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of Internet of Vehicles, and provides a vehicle-mounted voice interaction method, device, system, electronic equipment and storage medium, the method first acquires an audio signal of user voice; if at least one condition of a handheld sound pickup device being in a connected state, a vehicle-mounted screen displaying a related interface of a target application matched with the handheld sound pickup device and the handheld sound pickup device being in a use state being met, the user voice is collected based on the handheld sound pickup device; otherwise, the user voice is collected based on a fixed sound pickup device; then the audio signal is recognized to obtain a recognition result, and the recognition result is subjected to natural language processing to obtain semantic information corresponding to the recognition result; finally, the user voice is responded based on the semantic information. The method can automatically switch the handheld sound pickup device or the fixed sound pickup device for sound pickup based on a scene, and then can realize recognition and seamless connection of different modes, does not need manual operation of a user, and reduces learning cost of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle networking technology, and in particular to an in-vehicle voice interaction method, device, system, electronic device, and storage medium. Background Technology

[0002] With the increasing number of cars on the road, the number of entertainment systems and intelligent assistants in cars is also increasing. However, how to better meet people's entertainment needs in driving scenarios and provide users with a better service experience when using voice services, especially the application of karaoke software in the in-vehicle field, is one of the urgent tasks that the in-vehicle industry needs to accomplish.

[0003] Currently, in-vehicle voice interaction typically includes voice control mode and non-control mode. Voice control mode refers to the user's voice control of the vehicle, while non-control mode refers to other functions achieved through voice interaction, such as a singing mode (in-vehicle karaoke) where users can relax by singing. In non-control mode, voice control is disabled, requiring users to manually switch between different modes and perform various operations within the non-control mode. Because the switching process and all operations in non-control mode require manual execution, the rules are cumbersome and difficult to remember, significantly reducing the user experience.

[0004] Therefore, there is an urgent need to provide a method for in-vehicle voice interaction. Summary of the Invention

[0005] This invention provides an in-vehicle voice interaction method, device, system, electronic device, and storage medium to address the deficiencies in the prior art.

[0006] This invention provides an in-vehicle voice interaction method, comprising:

[0007] The audio signal of the user's voice is acquired. If at least one of the following conditions is met: the handheld microphone is connected, the in-vehicle screen displays a relevant interface of a target application that matches the handheld microphone, and the handheld microphone is in use, then the user's voice is acquired based on the handheld microphone; otherwise, the user's voice is acquired based on a fixed microphone.

[0008] The audio signal is identified to obtain an identification result, and the identification result is subjected to natural language processing to obtain the semantic information corresponding to the identification result;

[0009] Based on the semantic information, respond to the user's voice.

[0010] According to a vehicle-mounted voice interaction method provided by the present invention, the step of responding to the user's voice based on the semantic information includes:

[0011] Determine the sound source region of the user's voice;

[0012] Based on the semantic information, the response information of the user's voice is determined, and the response information is displayed on the target vehicle screen corresponding to the sound source area.

[0013] According to a vehicle-mounted voice interaction method provided by the present invention, the handheld microphone is in use, and this is determined based on the following steps:

[0014] Acquire in-vehicle images captured by the vehicle's onboard camera;

[0015] Based on the in-vehicle image, determine whether the user is using the handheld microphone.

[0016] If the user is using the handheld microphone, then the handheld microphone is determined to be in use.

[0017] According to the present invention, an in-vehicle voice interaction method is provided, wherein the user's voice is acquired based on the following steps:

[0018] Determine whether the handheld microphone is in a connected state;

[0019] If the handheld microphone is not connected, the user's voice is collected based on the fixed microphone.

[0020] If the handheld microphone is connected, then determine whether the vehicle screen displays the relevant interface of the target application and whether the handheld microphone is in use.

[0021] If the vehicle screen displays the relevant interface of the target application or the handheld microphone is in use, the user's voice is collected based on the handheld microphone.

[0022] If the vehicle screen does not display the relevant interface of the target application and the handheld microphone is not in use, the user's voice is collected based on the fixed microphone.

[0023] The present invention also provides an in-vehicle voice interaction device, comprising:

[0024] The audio acquisition module is used to acquire the audio signal of the user's voice. If at least one of the following conditions is met: the handheld microphone is connected, the vehicle screen displays a relevant interface of a target application that matches the handheld microphone, and the handheld microphone is in use, then the user's voice is acquired based on the handheld microphone; otherwise, the user's voice is acquired based on a fixed microphone.

[0025] The recognition and understanding module is used to recognize the audio signal, obtain the recognition result, and perform natural language processing on the recognition result to obtain the semantic information corresponding to the recognition result;

[0026] An interactive response module is used to respond to the user's voice based on the semantic information.

[0027] The present invention also provides an in-vehicle voice interaction system, including: a handheld microphone, a fixed microphone, and an in-vehicle cockpit domain controller, wherein the fixed microphone is connected to the in-vehicle cockpit domain controller, and the handheld microphone can be configured to be connected to the in-vehicle cockpit domain controller;

[0028] The vehicle-mounted cockpit domain controller is used to implement the above-mentioned vehicle-mounted voice interaction method.

[0029] According to the present invention, an in-vehicle voice interaction system is provided, wherein the fixed pickup device is fixed to the roof of the vehicle.

[0030] According to the present invention, an in-vehicle voice interaction system further includes an in-vehicle camera, wherein the in-vehicle camera is connected to the in-vehicle voice interaction device;

[0031] The vehicle-mounted camera is used to capture images inside the vehicle.

[0032] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the in-vehicle voice interaction method as described above.

[0033] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the in-vehicle voice interaction method as described above.

[0034] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the in-vehicle voice interaction method as described above.

[0035] The in-vehicle voice interaction method, device, system, electronic device, and storage medium provided by this invention first acquire the audio signal of the user's voice. If at least one of the following conditions is met: the handheld microphone is connected, the in-vehicle screen displays a relevant interface of a target application matching the handheld microphone, and the handheld microphone is in use, then the user's voice is acquired using the handheld microphone; otherwise, the user's voice is acquired using a fixed microphone. Then, the audio signal is recognized to obtain a recognition result, and the recognition result is processed by natural language to obtain semantic information corresponding to the recognition result. Finally, based on the semantic information, the user's voice is responded to. This method, by configuring both a handheld and a fixed microphone, can automatically switch between using either the handheld or fixed microphone based on the scenario, thereby achieving recognition and seamless integration of different modes without requiring manual user operation, reducing the user's learning cost. Moreover, this method can achieve voice interaction in different modes without requiring manual user operation, which can greatly improve the user experience. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on the drawings described below without creative effort.

[0037] Figure 1 This is one of the flowcharts of the in-vehicle voice interaction method provided by the present invention;

[0038] Figure 2 This is a schematic diagram of the user voice acquisition process in the in-vehicle voice interaction method provided by the present invention;

[0039] Figure 3 This is the second flowchart of the in-vehicle voice interaction method provided by the present invention;

[0040] Figure 4 This is a schematic diagram of the structure of the in-vehicle voice interaction device provided by the present invention;

[0041] Figure 5 This is a schematic diagram of the structure of the in-vehicle voice interaction system provided by the present invention;

[0042] Figure 6 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0043] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0044] Currently, in-vehicle voice interaction typically includes a voice control mode and a non-control mode. In the non-control mode, voice control is disabled, requiring users to manually switch between different modes and perform various operations. Because the switching process and all operations in the non-control mode require manual execution, the rules are cumbersome and difficult to remember, significantly reducing the user experience.

[0045] Existing voice recognition systems for televisions include a voice input system, a data analysis system, and a coordination system connecting the voice input system and the data analysis system. This system introduces voice recognition technology into televisions, simplifying the operation of the TV remote control and easily enabling functions such as voice-activated song selection, voice-activated movie playback, and voice-activated channel selection, thus enriching the television's functionality. However, this voice recognition system requires users to click on the microphone to switch to the TV control. Firstly, its application is primarily limited to television playback control; secondly, the need for manual operation during voice conversations makes the process less smooth, and the user experience needs improvement.

[0046] Existing solutions also include a microphone-based one-click song selection method, system, and storage medium. This method pre-adds an operation button to the microphone for one-click activation of the voice-activated song selection function. When a song needs to be selected, the operation button is used to activate the voice-activated song selection function, sending a command to the TV. The TV's karaoke background service then activates the recording device, receives the song selection voice command through the microphone, and performs the song selection and karaoke operation based on the command. However, this method also requires manual user operation and is not suitable for in-vehicle voice interaction scenarios.

[0047] Based on this, an in-vehicle voice interaction method is provided in this embodiment of the invention.

[0048] Figure 1 This is a flowchart illustrating an in-vehicle voice interaction method provided in an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes:

[0049] S1, acquire the audio signal of the user's voice; if at least one of the following conditions is met: the handheld microphone is connected, the vehicle screen displays a relevant interface of the target application matching the handheld microphone, and the handheld microphone is in use, then the user's voice is acquired based on the handheld microphone; otherwise, the user's voice is acquired based on the fixed microphone.

[0050] S2, the audio signal is identified to obtain an identification result, and the identification result is processed by natural language to obtain the semantic information corresponding to the identification result;

[0051] S3, respond to the user's voice based on the semantic information.

[0052] Specifically, the in-vehicle voice interaction method provided in this embodiment of the invention is executed by an in-vehicle voice interaction device, which can be configured in the in-vehicle cockpit domain controller.

[0053] First, step S1 is executed to acquire the audio signal of the user's voice. This user voice refers to the voice emitted by the user inside the vehicle. It can be control voice, such as voice commands for opening, closing, searching, and page turning of various in-vehicle applications, or even the user singing, depending on the user's actual application scenario. The vehicle can accommodate one or more users, each located in the driver's seat, front passenger seat, left rear seat, or right rear seat, etc. Therefore, the user voice can be the voice of one user or the voice of multiple users.

[0054] The audio signal of user speech refers to the digital signal obtained after sampling, quantization, and encoding. Sampling refers to the process of digitizing user speech on the time axis, quantization refers to the process of digitizing user speech on the amplitude axis, and encoding refers to recording the sampled and quantized digital signals in a certain format, thus obtaining the audio signal of user speech.

[0055] The vehicle can be equipped with both handheld and fixed microphones. The fixed microphone is connected to the in-vehicle voice interaction device, while the handheld microphone can be configured to connect to the in-vehicle voice interaction device. In other words, the fixed microphone is always connected to the in-vehicle voice interaction device, while the handheld microphone is not always connected and requires manual configuration to establish a connection.

[0056] Handheld pickup devices can be handheld microphones, while fixed pickup devices can be microphones in a fixed position, the location of which can be set as needed and is not specifically limited here. For example, fixed pickup devices can be placed on the top or side of a vehicle. Both handheld and fixed pickup devices can be one or more; for example, there can be one handheld pickup device and two fixed pickup devices, namely a left microphone and a right microphone.

[0057] User voice can be captured using either a handheld microphone or a fixed microphone. When capturing user voice using a handheld microphone, at least one of the following three conditions must be met: the handheld microphone is connected; the vehicle screen displays the interface of the target application that matches the handheld microphone; and the handheld microphone is in use.

[0058] "Handheld microphone in connected state" means that the handheld microphone is connected to the in-vehicle voice interaction device, and the two can transmit voice data. This can be achieved by turning on the handheld microphone or manually connecting it to the in-vehicle voice interaction device.

[0059] "Handheld microphone not connected" means that the handheld microphone is not connected to the in-vehicle voice interaction device, and voice data transmission between the two is impossible. Under normal circumstances, the handheld microphone is not connected.

[0060] The in-vehicle screens can include one or more, with multiple screens located in front of the driver's seat, front passenger seat, left rear seat, and right rear seat, respectively, to display content to the users in those seats. The central control screen can be an in-vehicle screen located in front of the driver's seat. The in-vehicle screens can be capacitive touchscreens, allowing users to interact with them via touch.

[0061] The target application (APP) that is compatible with the handheld microphone refers to an in-vehicle application that requires the use of a handheld microphone, such as a karaoke app or an in-vehicle karaoke app. The display of the target application's interface on the in-vehicle screen means that the target application is currently open and its relevant interface is displayed on the screen. This interface can be the target application's homepage, or it can be an interface that appears after the target application has been opened and interacted with the user several times, such as a song selection interface or a playlist interface; no specific limitation is made here.

[0062] The handheld microphone being in use refers to the state in which the handheld microphone is being used by the user, such as when the user is holding the handheld microphone or when the user is using the handheld microphone to sing or speak.

[0063] The handheld microphone is not in use, which means that the handheld microphone is not currently in use by the user, such as when the handheld microphone is placed on the front control panel.

[0064] When at least one of the following conditions is met: the handheld microphone is connected, the in-vehicle screen displays the interface of the target application matched with the handheld microphone, and the handheld microphone is in use, it indicates that the in-vehicle voice interaction is in non-control mode. The user is highly likely currently using or about to use the target application, therefore, it is necessary to use the handheld microphone to collect the user's voice. Otherwise, if none of the above conditions are met, it indicates that the in-vehicle voice interaction is in control mode, and the user is highly likely to have control needs; therefore, it is sufficient to directly use a fixed microphone to collect the user's voice.

[0065] Each handheld microphone captures user voice data, which, after processing by the target application, yields two audio signals. Similarly, each fixed microphone captures user voice data, which, after processing by the audio module, also yields two audio signals.

[0066] Subsequently, to ensure an accurate response to user voice, the engine's model processing can be used to eliminate echoes and reduce noise in the audio signal of the user's voice.

[0067] Then, step S2 is executed, where the audio signal of the user's speech is recognized by the speech recognition engine to obtain the recognition result. Subsequently, the recognition result is processed by a Natural Language Processing (NLP) engine to obtain the semantic information corresponding to the recognition result.

[0068] Finally, step S3 is executed, responding to the user's voice based on semantic information. Here, the semantic information can be parsed and processed, and the response information corresponding to the user's voice can be determined based on the parsing result. This response information can be voice feedback to the user, an automatic operation displayed on the in-vehicle screen, or voice feedback on the result of the automatic operation. For example, if the user's voice is "open playlist," the voice feedback to the user could be "opened," and the automatic operation would be displaying the opened playlist on the in-vehicle screen. As another example, if the user's voice is "switch songs," the automatic operation could be automatically switching songs and displaying the changed song information on the in-vehicle screen; the voice feedback on the result of the automatic operation could be the currently playing instrumental track.

[0069] The in-vehicle voice interaction method provided in this embodiment of the invention first acquires the audio signal of the user's voice. If at least one of the following conditions is met: the handheld microphone is connected, the in-vehicle screen displays a relevant interface of a target application matching the handheld microphone, and the handheld microphone is in use, then the user's voice is acquired using the handheld microphone; otherwise, the user's voice is acquired using a fixed microphone. Then, the audio signal is recognized to obtain a recognition result, and the recognition result is processed using natural language to obtain semantic information corresponding to the recognition result. Finally, based on the semantic information, the method responds to the user's voice. This method, by configuring both a handheld and a fixed microphone, can automatically switch between using either the handheld or fixed microphone based on the scenario, thereby achieving seamless integration and recognition across different modes without requiring manual user operation, reducing the user's learning cost. Moreover, this method can achieve voice interaction in different modes without requiring manual user operation, greatly improving the user experience.

[0070] Based on the above embodiments, the in-vehicle voice interaction method provided in this embodiment of the invention, wherein responding to the user's voice based on the semantic information includes:

[0071] Determine the sound source region of the user's voice;

[0072] Based on the semantic information, the response information of the user's voice is determined, and the response information is displayed on the target vehicle screen corresponding to the sound source area.

[0073] Specifically, when responding to user voice based on semantic information, the sound source region of the user's voice can be determined first. The sound source region can be determined by localizing the user's voice. This sound source region can include the adjacent areas of the driver's seat, front passenger seat, left rear seat, and right rear seat inside the vehicle.

[0074] Then, based on semantic information, the response information of the user's voice is determined and displayed on the target in-vehicle screen corresponding to the sound source area. For example, if the sound source area is near the driver's seat, it means that the user's voice is spoken by the driver, and the target in-vehicle screen can be the main control screen or the in-vehicle screen in front of the driver's seat. Displaying the response information on the target in-vehicle screen allows the user to better observe the response process and the response information.

[0075] Based on the above embodiments, the in-vehicle voice interaction method provided in this embodiment of the invention, wherein the handheld microphone is in use, is determined based on the following steps:

[0076] Acquire in-vehicle images captured by the vehicle's onboard camera;

[0077] Based on the in-vehicle image, determine whether the user is using the handheld microphone.

[0078] If the user is using the handheld microphone, then the handheld microphone is determined to be in use.

[0079] Specifically, to determine whether a handheld microphone is in use, one can first acquire an in-vehicle image captured by a vehicle-mounted camera. This camera can be installed on the ceiling inside the vehicle and can capture a 360-degree view. In other words, the in-vehicle image can include images from various locations within the vehicle.

[0080] By performing facial pose recognition on images inside the vehicle, it can be determined whether the user is using a handheld microphone. If the user is using a handheld microphone, it is determined that the microphone is in use. Otherwise, if the user is not using a handheld microphone, it is determined that the microphone is not in use.

[0081] In this embodiment of the invention, by obtaining in-vehicle images captured by an in-vehicle camera to determine whether a handheld microphone is in use, the microphone scene recognition of the handheld microphone can be made more accurate.

[0082] Based on the above embodiments, the in-vehicle voice interaction method provided in this embodiment of the invention obtains the user's voice through the following steps:

[0083] Determine whether the handheld microphone is in a connected state;

[0084] If the handheld microphone is not connected, the user's voice is collected based on the fixed microphone.

[0085] If the handheld microphone is connected, then determine whether the vehicle screen displays the relevant interface of the target application and whether the handheld microphone is in use.

[0086] If the vehicle screen displays the relevant interface of the target application or the handheld microphone is in use, the user's voice is collected based on the handheld microphone.

[0087] If the vehicle screen does not display the relevant interface of the target application and the handheld microphone is not in use, the user's voice is collected based on the fixed microphone.

[0088] Specifically, when collecting user voice, such as Figure 2 As shown, you can first determine whether the handheld microphone is connected. If the handheld microphone is not connected, you can directly use the fixed microphone to collect the user's voice.

[0089] If the handheld microphone is connected, the system further determines whether the in-vehicle screen displays the relevant interface of the target application and whether the handheld microphone is in use.

[0090] If the in-vehicle screen displays the relevant interface of the target application or the handheld microphone is in use, the user's voice will be collected using the handheld microphone.

[0091] If the in-vehicle screen does not display the relevant interface of the target application and the handheld microphone is not in use, the user's voice will be collected using a fixed microphone.

[0092] In this embodiment of the invention, by making multiple judgments, various situations of collecting user voice based on handheld microphones can be more comprehensively determined.

[0093] like Figure 3 As shown, based on the above embodiments, the in-vehicle voice interaction method provided in this embodiment of the invention includes:

[0094] First, determine whether the user is in a handheld microphone usage scenario. This involves checking if at least one of the following conditions is met: the handheld microphone is connected, the in-vehicle screen displays a relevant interface of the target application matching the handheld microphone, and the handheld microphone is in use. If the user is in a handheld microphone usage scenario, i.e., at least one of the above conditions is met, then the user's voice is collected based on the handheld microphone; otherwise, the user's voice is collected based on the fixed microphone.

[0095] Next, echo cancellation and noise reduction are performed on the user's audio signal. Then, the results of the echo cancellation and noise reduction are recognized, and natural language processing is applied to these results to obtain the corresponding semantic information.

[0096] Finally, the sound source region of the user's voice is determined, and based on semantic information, the response information of the user's voice is determined and displayed on the target in-vehicle screen corresponding to the sound source region.

[0097] In summary, this invention provides an in-vehicle voice interaction method based on speech recognition, semantic understanding, multi-modal recognition, microphone, and multi-screen interaction design in the field of vehicle networking technology. While singing in the car, users can seamlessly perform voice-activated song selection, deletion, and pinning operations using a handheld microphone, eliminating the need for manual song searching. Users can also utilize a fixed microphone for voice interaction, greatly satisfying their need for quick and simple voice interaction while singing.

[0098] like Figure 4 As shown, based on the above embodiments, this embodiment of the invention provides an in-vehicle voice interaction device, including:

[0099] The audio acquisition module 41 is used to acquire the audio signal of the user's voice. If at least one of the following conditions is met: the handheld microphone is in a connected state, the vehicle screen displays a relevant interface of a target application that matches the handheld microphone, and the handheld microphone is in use, then the user's voice is acquired based on the handheld microphone; otherwise, the user's voice is acquired based on the fixed microphone.

[0100] The recognition and understanding module 42 is used to recognize the audio signal, obtain the recognition result, and perform natural language processing on the recognition result to obtain the semantic information corresponding to the recognition result;

[0101] The interactive response module 43 is used to respond to the user's voice based on the semantic information.

[0102] Based on the above embodiments, the in-vehicle voice interaction device provided in this embodiment of the invention, wherein the interaction response module is specifically used for:

[0103] Determine the sound source region of the user's voice;

[0104] Based on the semantic information, the response information of the user's voice is determined, and the response information is displayed on the target vehicle screen corresponding to the sound source area.

[0105] Based on the above embodiments, the in-vehicle voice interaction device provided in this embodiment of the invention further includes a usage status determination module, used for:

[0106] Acquire in-vehicle images captured by the vehicle's onboard camera;

[0107] Based on the in-vehicle image, determine whether the user is using the handheld microphone.

[0108] If the user is using the handheld microphone, then the handheld microphone is determined to be in use.

[0109] Based on the above embodiments, the in-vehicle voice interaction device provided in this embodiment of the invention further includes a user voice acquisition module, used for:

[0110] Determine whether the handheld microphone is in a connected state;

[0111] If the handheld microphone is not connected, the user's voice is collected based on the fixed microphone.

[0112] If the handheld microphone is connected, then determine whether the vehicle screen displays the relevant interface of the target application and whether the handheld microphone is in use.

[0113] If the vehicle screen displays the relevant interface of the target application or the handheld microphone is in use, the user's voice is collected based on the handheld microphone.

[0114] If the vehicle screen does not display the relevant interface of the target application and the handheld microphone is not in use, the user's voice is collected based on the fixed microphone.

[0115] Specifically, the functions of each module in the in-vehicle voice interaction device provided in the embodiments of the present invention correspond one-to-one with the operation flow of each step in the above method embodiments, and the achieved effects are also the same. For details, please refer to the above embodiments, and this will not be repeated in the embodiments of the present invention.

[0116] like Figure 5 As shown, based on the above embodiments, this embodiment of the invention also provides an in-vehicle voice interaction system, including: a handheld microphone 51, a fixed microphone 52, and an in-vehicle cockpit domain controller 53. The fixed microphone 52 is connected to the in-vehicle cockpit domain controller 53, and the handheld microphone 51 can be configured to be connected to the in-vehicle cockpit domain controller 53.

[0117] The vehicle cockpit domain controller 53 is used to implement the vehicle voice interaction methods provided in the above embodiments.

[0118] Specifically, the functions of each module in the vehicle cockpit domain controller provided in this embodiment of the invention correspond one-to-one with the operation flow of each step in the above method-like embodiments, and the achieved effects are also the same. For details, please refer to the above embodiments, and this will not be repeated in this embodiment of the invention.

[0119] Based on the above embodiments, the in-vehicle voice interaction system provided in this embodiment of the invention has the fixed pickup device fixed to the roof of the vehicle.

[0120] Based on the above embodiments, the in-vehicle voice interaction system provided in this embodiment of the invention further includes an in-vehicle camera, which is connected to the in-vehicle voice interaction device; the in-vehicle camera is used to capture images inside the vehicle.

[0121] Figure 6 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 6As shown, the electronic device may include a processor 610, a communications interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communications interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute the in-vehicle voice interaction method provided in the above embodiments. The method includes: acquiring the audio signal of the user's voice; if at least one of the following conditions is met: the handheld microphone is in a connected state, the in-vehicle screen displays a relevant interface of a target application matching the handheld microphone, and the handheld microphone is in use, then the user's voice is acquired based on the handheld microphone; otherwise, the user's voice is acquired based on a fixed microphone; recognizing the audio signal to obtain a recognition result, and performing natural language processing on the recognition result to obtain semantic information corresponding to the recognition result; and responding to the user's voice based on the semantic information.

[0122] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0123] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the in-vehicle voice interaction method provided in the above embodiments. The method includes: acquiring an audio signal of a user's voice; if at least one of the following conditions is met: the handheld microphone is in a connected state, the in-vehicle screen displays a relevant interface of a target application matching the handheld microphone, and the handheld microphone is in use, then the user's voice is acquired based on the handheld microphone; otherwise, the user's voice is acquired based on a fixed microphone; recognizing the audio signal to obtain a recognition result, and performing natural language processing on the recognition result to obtain semantic information corresponding to the recognition result; and responding to the user's voice based on the semantic information.

[0124] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the in-vehicle voice interaction method provided in the above embodiments. The method includes: acquiring an audio signal of a user's voice; if at least one of the following conditions is met: a handheld microphone is in a connected state, an in-vehicle screen displays a relevant interface of a target application matching the handheld microphone, and the handheld microphone is in use, then the user's voice is acquired based on the handheld microphone; otherwise, the user's voice is acquired based on a fixed microphone; recognizing the audio signal to obtain a recognition result, and performing natural language processing on the recognition result to obtain semantic information corresponding to the recognition result; and responding to the user's voice based on the semantic information.

[0125] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0126] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A vehicle-mounted voice interaction method, characterized in that, include: Acquire the audio signal of the user's voice; If at least one of the following conditions is met: the handheld microphone is connected, the in-vehicle screen displays an interface of a target application that matches the handheld microphone, and the handheld microphone is in use, it indicates that the in-vehicle voice interaction is in non-control mode, and the user's voice is collected based on the handheld microphone. Otherwise, it indicates that the in-vehicle voice interaction is in control mode, and the user's voice is collected based on a fixed sound pickup device; the control mode is the user's voice control of the vehicle, and the non-control mode is to achieve other functions through voice interaction; The audio signal is identified to obtain an identification result, and the identification result is subjected to natural language processing to obtain the semantic information corresponding to the identification result; Based on the semantic information, the system responds to the user's voice, and the response information includes voice feedback to the user, automatic operations displayed on the in-vehicle screen, and voice feedback on the results of the automatic operations.

2. The in-vehicle voice interaction method according to claim 1, characterized in that, The step of responding to the user's voice based on the semantic information includes: Determine the sound source region of the user's voice; Based on the semantic information, the response information of the user's voice is determined, and the response information is displayed on the target vehicle screen corresponding to the sound source area.

3. The in-vehicle voice interaction method according to claim 1, characterized in that, The handheld microphone is in use, as determined by the following steps: Acquire in-vehicle images captured by the vehicle's onboard camera; Based on the in-vehicle image, determine whether the user is using the handheld microphone. If the user is using the handheld microphone, then the handheld microphone is determined to be in use.

4. The in-vehicle voice interaction method according to any one of claims 1-3, characterized in that, The user's voice was collected based on the following steps: Determine whether the handheld microphone is in a connected state; If the handheld microphone is not connected, the user's voice is collected based on the fixed microphone. If the handheld microphone is connected, then determine whether the vehicle screen displays the relevant interface of the target application and whether the handheld microphone is in use. If the vehicle screen displays the relevant interface of the target application or the handheld microphone is in use, the user's voice is collected based on the handheld microphone. If the vehicle screen does not display the relevant interface of the target application and the handheld microphone is not in use, the user's voice is collected based on the fixed microphone.

5. A vehicle-mounted voice interaction device, characterized in that, include: The audio acquisition module is used to acquire the audio signal of the user's voice; If at least one of the following conditions is met: the handheld microphone is connected, the in-vehicle screen displays an interface of a target application that matches the handheld microphone, and the handheld microphone is in use, it indicates that the in-vehicle voice interaction is in non-control mode, and the user's voice is collected based on the handheld microphone. Otherwise, it indicates that the in-vehicle voice interaction is in control mode, and the user's voice is collected based on a fixed sound pickup device; The recognition and understanding module is used to recognize the audio signal, obtain the recognition result, and perform natural language processing on the recognition result to obtain the semantic information corresponding to the recognition result; The interactive response module is used to respond to the user's voice based on the semantic information. The response information includes voice feedback to the user, automatic operations displayed on the in-vehicle screen, and voice feedback on the results of the automatic operations.

6. An in-vehicle voice interaction system, characterized in that, include: The system includes a handheld microphone, a fixed microphone, and a vehicle-mounted cockpit domain controller. The fixed microphone is connected to the vehicle-mounted cockpit domain controller, and the handheld microphone can be configured to connect to the vehicle-mounted cockpit domain controller. The vehicle-mounted cockpit domain controller is used to implement the vehicle-mounted voice interaction method as described in any one of claims 1-4.

7. The in-vehicle voice interaction system according to claim 6, characterized in that, The fixed microphone is fixed to the roof of the vehicle.

8. The in-vehicle voice interaction system according to claim 6, characterized in that, It also includes an in-vehicle camera, which is connected to the in-vehicle cockpit domain controller; The vehicle-mounted camera is used to capture images inside the vehicle.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the in-vehicle voice interaction method as described in any one of claims 1-4.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the in-vehicle voice interaction method as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Multi-screen interaction method, device and equipment and computer readable storage medium

    CN115431762A

  • Vehicle-mounted microphone control method, vehicle-mounted microphone and vehicle-mounted microphone storage device

    CN115442683A