In-vehicle voice interaction method, system, storage medium and terminal

By determining the working mode of the voice system in the on-board terminal and combining or independently processing voice information, the voice interaction problem of multiple people or multiple devices is solved, and the user experience is improved.

CN115497467BActive Publication Date: 2025-08-08SHANGHAI PATEO ELECTRONIC EQUIPMENT MANUFACTURING CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110674544.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-17
Publication Date
2025-08-08
Estimated Expiration
2041-06-17

AI Technical Summary

Technical Problem

The voice interaction function of existing vehicle terminals cannot effectively process voice information from multiple people or multiple devices, resulting in poor user experience.

Method used

Determine the voice system working mode by obtaining voice commands, and combine or process the voice information according to the mode, including the whole-vehicle merging radio mode and the whole-vehicle independent radio mode, and process or merge the voice information separately.

Benefits of technology

It improves the continuous voice usage experience and can switch the voice system working mode according to the user's intentions to meet the voice interaction needs of multiple people or multiple devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115497467B_ABST
    Figure CN115497467B_ABST
Patent Text Reader

Abstract

A method, system, storage medium, and terminal for in-vehicle voice interaction include: obtaining a voice command; determining a voice system operating mode based on the voice command; obtaining multiple voice messages; and, based on the voice system operating mode, merging or independently processing the multiple voice messages to obtain processing results. This allows switching the voice system operating mode based on user intent, thereby improving the continuous voice experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of automotive intelligent technology, and in particular to a method, system, storage medium and terminal for in-vehicle voice interaction. Background Art

[0002] With the rapid development of modern civilization, cars have also been rapidly popularized and have become a necessity for modern people. With the popularization of cars, the related vehicle information terminal equipment has also had room for development, and these devices have brought a lot of convenience to modern people.

[0003] During operation, the on-board terminal in the vehicle can provide users with corresponding services by recognizing the voice input by the user, such as playing music, turning on navigation, opening windows, etc., giving users a more convenient and intelligent driving experience.

[0004] However, as user demands increase, the functions of in-vehicle terminals also need to be improved accordingly. Summary of the Invention

[0005] The technical problem solved by the present invention is to provide a method, system, storage medium and terminal for vehicle-mounted voice interaction to enhance the functions of the vehicle-mounted terminal.

[0006] To solve the above technical problems, an embodiment of the present disclosure provides a method for in-vehicle voice interaction, including: obtaining a voice command; determining a voice system operating mode based on the voice command; obtaining multiple voice messages; and merging or independently processing the multiple voice messages based on the voice system operating mode to obtain a processing result.

[0007] Optionally, several pieces of voice information are collected by different sound receiving devices respectively; or, several pieces of voice information are sent by different persons and collected by the same sound receiving device; or, several pieces of voice information are sent by different persons and collected by different sound receiving devices.

[0008] Optionally, the voice system working mode includes a whole-vehicle independent radio mode or a whole-vehicle combined radio mode; in the whole-vehicle independent radio mode, several voice messages are processed independently; in the whole-vehicle combined radio mode, several voice messages are combined for processing.

[0009] Optionally, in the vehicle-wide combined radio mode, when the time interval between any two voice messages is less than a preset value, the two voice messages are combined into one voice message.

[0010] Optionally, in the vehicle-wide combined radio reception mode, after obtaining a preset number of voice messages, the preset number of voice messages are combined into one voice message.

[0011] Optionally, in the whole-vehicle independent sound reception mode, the method for independently processing a plurality of voice messages includes: processing the plurality of voice messages separately according to different sound reception devices to obtain independent processing results.

[0012] Optionally, the method for the sound receiving device to collect voice information includes: the sound receiving device collects voice information with strong signals and filters out voice information with weak signals according to the signal strength of the collected voice information.

[0013] Optionally, in the whole-vehicle independent sound receiving mode, the method for independently processing a plurality of voice messages includes: processing the plurality of voice messages separately according to different senders to obtain independent processing results.

[0014] Optionally, the method for determining the voice system working mode according to the voice command includes: judging the application that executes the voice command according to the voice command; and determining the voice system working mode according to the business type of the application that executes the voice command.

[0015] Optionally, after obtaining the processing result, the method further includes: outputting the processing result, and the method of outputting the processing result includes: displaying the processing result, outputting the processing result by voice, or executing the next step of the process according to the content of the processing result.

[0016] Optionally, the method for merging and processing a plurality of voice messages includes: superimposing and accumulating the contents of the plurality of voice messages to obtain a processing result.

[0017] Correspondingly, an embodiment of the present disclosure also provides an in-vehicle voice interaction system, including: a command receiving module for obtaining voice commands; an analysis module for determining the voice system working mode based on the voice commands; a voice collection module for obtaining multiple voice messages; and a voice processing module for merging or independently processing multiple voice messages according to the voice system working mode to obtain processing results.

[0018] Optionally, several pieces of voice information are collected by different sound receiving devices respectively; or, several pieces of voice information are sent by different persons and collected by the same sound receiving device; or, several pieces of voice information are sent by different persons and collected by different sound receiving devices.

[0019] Optionally, the voice system working mode includes a whole-vehicle independent radio mode or a whole-vehicle combined radio mode; in the whole-vehicle independent radio mode, the voice processing module independently processes several voice messages; in the whole-vehicle combined radio mode, the voice processing module combines and processes several voice messages.

[0020] Optionally, the analysis module includes: an analysis unit, used to determine the application that executes the voice command based on the voice command; and a selection unit, used to determine the voice system working mode based on the business type of the application that executes the voice command.

[0021] Optionally, it also includes: a result output module, used to output the processing result.

[0022] Optionally, the result output module includes: a display unit, configured to display the processing result.

[0023] Optionally, the result output module includes: a voice output unit, used to output the processing result in voice.

[0024] Optionally, the result output module includes: a processing unit, configured to execute the next step of the process according to the content of the processing result.

[0025] Correspondingly, an embodiment of the present disclosure further provides a storage medium on which computer instructions are stored, characterized in that the steps of any of the above methods are executed when the computer instructions are executed.

[0026] Correspondingly, an embodiment of the present disclosure also provides a terminal, including a memory and a processor, wherein the memory stores computer instructions that can be run on the processor, and is characterized in that the processor executes the steps of any of the above methods when running the computer instructions.

[0027] Compared with the prior art, the embodiments of the present disclosure have the following beneficial effects:

[0028] In the disclosed embodiment, the voice system operating mode is determined based on the voice command, and then multiple voice messages are obtained. Based on the voice system operating mode, the multiple voice messages are merged or individually recorded to obtain processing results. This allows the voice system operating mode to be switched based on user intent, thereby improving the continuous voice experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 and Figure 2 is a flow chart of the vehicle-mounted voice interaction method according to an embodiment of the present disclosure;

[0030] Figure 3 is a flow chart of an in-vehicle voice interaction method in another disclosed embodiment;

[0031] Figure 4 and Figure 5 is a schematic structural diagram of the vehicle-mounted voice interaction system in an embodiment of the present disclosure;

[0032] Figure 6It is a structural diagram of an in-vehicle voice interaction system in another disclosed embodiment. DETAILED DESCRIPTION

[0033] As described in the background art, the functions of the vehicle-mounted terminal also need to be improved accordingly.

[0034] Specifically, in smart car applications, some scenarios require multiple people to listen separately, such as the driver booking a plane ticket on the main driver's screen and the co-pilot booking movie tickets on the co-pilot screen. In other scenarios, differentiating between different people is unnecessary; everyone can simply complete the conversation together. For example, on the main driver's screen, we order takeout, the driver moves to the ordering screen, and the person in the back seat specifies their needs, such as a Coke or a fried chicken drumstick. The ordering screen on the main screen then aggregates these orders.

[0035] It is necessary to provide a voice interaction solution that has both parallel and merged conversations.

[0036] In order to solve the above problems, the embodiments of the present disclosure provide a method, system, storage medium and terminal for in-vehicle voice interaction.

[0037] In order to make the above-mentioned objects, features and beneficial effects of the present invention more obvious and easy to understand, specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0038] Figure 1 and Figure 2 It is a flowchart of the method of in-vehicle voice interaction in an embodiment of the present disclosure.

[0039] Please refer to Figure 1 , the in-vehicle voice interaction method includes:

[0040] Step S10: obtaining a voice command;

[0041] Step S11: determining the voice system operating mode according to the voice command;

[0042] Step S12: Obtaining several voice messages;

[0043] Step S13: Based on the voice system working mode, multiple voice messages are processed together or independently to obtain a processing result.

[0044] In this embodiment, the method determines the voice system operating mode based on the voice command, then obtains multiple voice messages. Depending on the voice system operating mode, the multiple voice messages are merged or individually recorded to obtain processing results. This allows the voice system operating mode to be switched based on user intent, thereby improving the continuous voice experience.

[0045] Next, each step will be described separately.

[0046] Please continue to refer to Figure 1 , execute step S10: obtain voice command.

[0047] The voice command is what the user wants to do on the car computer. It is spoken through a voice message and the car audio device receives the voice message. For example, the user can say "order takeout", "book a flight", "book a movie ticket", or "play an encyclopedia quiz".

[0048] The vehicle-mounted radio equipment is arranged at various locations in the vehicle as required. In this embodiment, there are multiple vehicle-mounted radio equipment, which are respectively located on the four door handles of the vehicle.

[0049] Please continue to refer to Figure 1 , executing step S11: determining the voice system working mode according to the voice command.

[0050] The voice system operating mode includes a whole-vehicle independent radio mode or a whole-vehicle integrated radio mode. The vehicle-mounted system determines which voice system operating mode the voice command corresponds to based on the content of the voice command.

[0051] For example, when a user says "order takeout", the needs of multiple people generate one order, which generally corresponds to the combined radio mode of the entire vehicle; when a user says "book a plane ticket", booking a plane ticket is an independent behavior of a single person, which generally corresponds to the independent radio mode of the entire vehicle; when a user says "play an encyclopedia quiz competition", multiple people are required to participate and score separately, which generally corresponds to the independent radio mode of the entire vehicle.

[0052] The whole-vehicle combined radio mode means that each vehicle-mounted radio device receives voice information from all directions of the vehicle, and summarizes and sends it to the application working in the current whole-vehicle combined radio mode, and then the application summarizes and processes the summarized voice information.

[0053] For example, when the main driver says "order takeout", the takeout app is turned on. In the integrated radio mode of the whole car, users in each seat can express their needs. The main driver says "a pizza", and a "pizza" will be added to the order. The co-driver says "I want fried chicken", and another "fried chicken" will be added to the same order. If more than one is needed, you can also add the quantity "order two fried chickens". If you want to cancel, you can also say "cancel one fried chicken". The main driver then places the order and pays. At this time, the in-vehicle radio equipment receives the order requirements of each user and sends the order requirements to the takeout app, completing the takeout ordering process together. The co-driver can also say "order takeout", triggering the integrated radio mode of the whole car, and repeat the above takeout ordering process.

[0054] The independent radio mode for the entire vehicle means that each vehicle-mounted radio device only receives the voice information of the user in the seat corresponding to the vehicle-mounted radio device, that is, each vehicle-mounted radio device only receives voice information with strong signals, filters out voice information with weak signals, and sends them separately to the application working in the current independent radio mode for the entire vehicle, and then the application processes the aggregated voice information separately.

[0055] For example, when a user says "play the encyclopedia quiz competition", the Encyclopedia Quiz application is turned on. In the independent audio reception mode of the entire vehicle, when multiple users speak at the same time, each vehicle-mounted audio device only receives the voice information of the user in the seat corresponding to the vehicle-mounted audio device, and suppresses the voice information from other seats. It recognizes the voice information of the user in the seat corresponding to the vehicle-mounted audio device and sends the voice information to the Encyclopedia Quiz application. The Encyclopedia Quiz application understands and judges the voice information and records the scores respectively.

[0056] For example, if the driver says "book a flight," the ticket booking app opens. In the vehicle's independent audio mode, the driver's seat's onboard audio device receives the driver's voice message, completing the flight booking process. Meanwhile, if the co-driver says "book a movie ticket," the co-driver's onboard audio device, in the vehicle's independent audio mode, receives the co-driver's voice message, completing the movie ticket booking process. Both the driver's and co-driver's flight booking processes can proceed simultaneously.

[0057] Please refer to Figure 2 In this embodiment, the method for determining the working mode of the voice system according to the voice command includes:

[0058] Step S111: determining an application to execute the voice command according to the voice command;

[0059] Step S112: Determine the voice system operating mode according to the service type of the application that executes the voice command.

[0060] Now combined Figure 2 Each step is described in detail.

[0061] Executing step S111: determining an application to execute the voice command according to the voice command.

[0062] The in-vehicle system determines which app to execute the voice command based on the content of the voice command. For example, if a user says "order takeout," the "order takeout" command is executed by a food delivery app, such as Ele.me or Meituan Waimai. If a user says "book a flight," the "book a flight" command is executed by a travel ticketing app, such as Ctrip, Qunar, or Fliggy.

[0063] Execute step S112: determine the voice system operating mode according to the service type of the application executing the voice command.

[0064] In this embodiment, the voice system operating mode used is determined by the service type of the application executing the voice command. Each application has its own corresponding voice system operating mode. This allows the voice system operating mode to be switched based on user intent, thereby improving the continuous voice usage experience.

[0065] For example, when a user says "order takeout", the "order takeout" command is executed by the takeout application. Ordering takeout means ordering multiple items to generate an order, which can be ordered by one person or multiple people. Therefore, the takeout application uses the full-vehicle integrated radio mode.

[0066] For another example, when a user says "book a plane ticket", the "book a plane ticket" command is executed by a travel ticketing application. Booking a plane ticket requires accurate confirmation of personal identity information and travel time, and does not require multiple people to operate. Therefore, travel ticketing applications use an independent radio mode for the entire vehicle.

[0067] For example, when a user says "play the encyclopedia quiz contest", the command "play the encyclopedia quiz contest" is executed by the encyclopedia quiz application. The encyclopedia quiz contest game is set up in a mode where multiple people participate and each person is scored separately. Therefore, the encyclopedia quiz application uses an independent radio reception mode for the entire vehicle.

[0068] Please continue to refer to Figure 1 , execute step S12: obtain several voice messages.

[0069] Several voice messages are sent by one or more users and collected by various in-vehicle audio equipment.

[0070] In one embodiment, the plurality of voice messages are collected by different audio receivers. In this case, the plurality of voice messages may be sent by one person or multiple people, and each is collected by the vehicle-mounted audio receiver corresponding to the user's seat. This utilizes a fully independent audio receiver mode.

[0071] In another embodiment, several voice messages are sent by different people and collected by the same audio receiving device. In this case, the voice messages are sent by multiple users and collected by a single vehicle-mounted audio receiving device. In this case, the vehicle-wide combined audio receiving mode is used.

[0072] In another embodiment, several voice messages are sent by different people and collected by different audio receivers. In this case, the voice messages are sent by multiple users and collected by the in-vehicle audio receivers corresponding to the seats where the users are sitting. In this case, the whole vehicle independent audio receiver mode is used.

[0073] In the vehicle-wide independent audio reception mode, the audio reception device collects voice information by collecting strong voice information and filtering out weak voice information based on the signal strength of the collected voice information. That is, any audio reception device only collects the voice of the user in the seat corresponding to the audio reception device. For example, the audio reception device on the driver's door collects the voice of the user in the driver's seat and filters out the voice of users in other seats.

[0074] Please continue to refer to Figure 1 , execute step S13: according to the working mode of the voice system, merge or independently process the multiple voice messages to obtain the processing result.

[0075] The processing result is a combination of several voice messages or independent processing, and the final user demand is obtained, which requires the corresponding application to perform the next step of processing.

[0076] In the independent radio reception mode of the entire vehicle, several voice messages are processed independently.

[0077] In one embodiment, in the whole-vehicle independent sound reception mode, the method for independently processing a plurality of voice messages includes: processing the plurality of voice messages separately according to different sound reception devices to obtain independent processing results.

[0078] That is, the voice information collected by any sound receiving device is regarded as a set, and the voice information in the set is processed to obtain a processing result. The processing results corresponding to each sound receiving device are independent.

[0079] For example, when the driver says "book a flight," the ticket booking app opens. In the independent car-wide audio mode, the onboard audio device in the driver's seat receives the voice information of the user in the driver's seat, completing the flight booking process. If the co-driver wants to "book movie tickets," the co-driver says "book movie tickets." In the independent car-wide audio mode, the onboard audio device in the co-driver's seat receives the voice information of the user in the co-driver's seat, completing the movie ticket booking process. At this time, the voice of the driver collected by the onboard audio device in the driver's seat is processed for flight booking, and the voice of the co-driver collected by the onboard audio device in the co-driver's seat is processed for movie ticket booking, obtaining independent processing results. The processing results at this time are the operations of booking flights or movie tickets in the travel ticketing software.

[0080] In another embodiment, in the whole-vehicle independent sound receiving mode, the method for independently processing a plurality of voice messages includes: processing the plurality of voice messages separately according to different senders to obtain independent processing results.

[0081] That is, among several voice messages, the voice messages sent by one user are taken as a set, and the voice messages in the set are processed to obtain a processing result. The sound receiving device has a voiceprint recognition function and can distinguish different users by recognizing voiceprints.

[0082] In the whole-vehicle combined radio mode, several voice messages are combined and processed.

[0083] The method for merging and processing a plurality of voice messages includes: superimposing and accumulating the contents of the plurality of voice messages to obtain a processing result.

[0084] At this time, regardless of whether the voice information is collected by one or more sound receiving devices, or whether the voice information is sent by one user or multiple users, the contents of multiple voice information are superimposed and accumulated to obtain a processing result.

[0085] In one embodiment, in the vehicle-wide combined sound reception mode, when the time interval between any two voice messages is less than a preset value, the two voice messages are combined into one voice message.

[0086] For example, if one person says "today" and another says "weather" after the word "day", they will be merged into one voice message "today's weather".

[0087] In another embodiment, in the vehicle-wide combined recording mode, after a preset number of voice messages are acquired, the preset number of voice messages are combined into one voice message. That is, during the continuous recording process, several consecutive voice messages are set as one voice message, and the contents of the multiple voice messages are then superimposed and accumulated to obtain a processing result.

[0088] Figure 3 It is a flowchart of a method for in-vehicle voice interaction in another disclosed embodiment.

[0089] Please refer to Figure 3 , Figure 3 For Figure 1 Based on the schematic diagram, the in-vehicle voice interaction method includes:

[0090] Step S10: obtaining a voice command;

[0091] Step S11: determining the voice system operating mode according to the voice command;

[0092] Step S12: Obtaining several voice messages;

[0093] Step S13: Based on the voice system working mode, multiple voice messages are combined or processed independently to obtain processing results;

[0094] Step S14: outputting the processing result. The methods of outputting the processing result include: displaying the processing result, outputting the processing result by voice, or executing the next step according to the content of the processing result.

[0095] For detailed description of steps S10 to S13, please refer to Figure 1 and Figure 2 , I will not go into details here.

[0096] Please continue to refer to Figure 3 , execute step S14: output the processing result.

[0097] The way of outputting the processing result includes: displaying the processing result, outputting the processing result by voice, or executing the next step of the process according to the content of the processing result.

[0098] In one embodiment, the method of outputting the processing result includes: displaying the processing result, that is, displaying the results of independent processing or combined processing of multiple voice messages on the vehicle's central control display interface.

[0099] For example, when a user says "play the encyclopedia quiz competition", the independent radio mode of the entire car is turned on. The encyclopedia quiz competition game is set up in a mode where multiple people participate and each person is scored separately. Each car-mounted radio device only receives the voice information of the user in the seat corresponding to the car-mounted radio device, and sends the voice information to the encyclopedia quiz application. The encyclopedia quiz application understands and judges the voice information, and records the scores separately. Finally, the scores of the users corresponding to each car-mounted radio device are displayed on the car's central control display interface.

[0100] In another embodiment, the method of outputting the processing result includes: outputting the processing result by voice, that is, announcing the result of independent processing or combined processing of multiple voice messages by voice broadcasting through an in-vehicle voice device.

[0101] For example, when the user says "book a flight", the independent radio mode of the entire vehicle is turned on. The user states his or her travel time and other requirements under the guidance of the travel ticketing application. The travel ticketing application selects eligible flights based on the travel time and other requirements and announces them by voice. The user then confirms which flight to book based on the announced content and makes a reservation.

[0102] In another embodiment, the method of outputting the processing result includes: executing the next step of the process according to the content of the processing result, that is, after merging or independently processing several voice messages, the user's final demand is obtained, and the corresponding application needs to perform the next step of processing.

[0103] For example, when a user says "order takeout," the food delivery app opens. With the car's integrated audio system, users in each seat can express their needs. If the driver says "a pizza," "pizza" is added to the order. If the passenger says "fried chicken," "fried chicken" is added to the order. If they want more, they can add the quantity, "order two fried chickens." To cancel, they can say "cancel one fried chicken." Finally, multiple items are ordered into one order. After confirming the order, the user fills in the delivery address and instructs the food delivery app to place the order for delivery.

[0104] Figure 4 and Figure 5 It is a structural diagram of the vehicle-mounted voice interaction system in an embodiment of the present disclosure.

[0105] Please refer to Figure 4 The structure of the vehicle voice interaction system includes:

[0106] A command receiving module 100 is used to obtain voice commands;

[0107] An analysis module 200 is configured to determine a voice system operating mode based on the voice command;

[0108] The voice collection module 300 is used to obtain a number of voice messages;

[0109] The voice processing module 400 is used to combine or independently process a plurality of voice messages according to the working mode of the voice system to obtain a processing result.

[0110] The system can switch the voice system working mode according to the user's intention, thereby improving the continuous voice usage experience.

[0111] Next, each module will be described separately.

[0112] Please continue to refer to Figure 4 The command receiving module 100 is used to obtain voice commands.

[0113] The voice command is what the user wants to do on the car computer. It is spoken through a voice message and the car audio device receives the voice message. For example, the user can say "order takeout", "book a flight", "book a movie ticket", or "play an encyclopedia quiz".

[0114] The vehicle-mounted radio equipment is arranged at various locations in the vehicle as required. In this embodiment, there are multiple vehicle-mounted radio equipment, which are respectively located on the four door handles of the vehicle.

[0115] Please continue to refer to Figure 4 The analysis module 200 is used to determine the voice system working mode according to the voice command.

[0116] The voice system operating mode includes a whole-vehicle independent radio mode or a whole-vehicle integrated radio mode. The vehicle-mounted system determines which voice system operating mode the voice command corresponds to based on the content of the voice command.

[0117] For example, when a user says "order takeout", the needs of multiple people generate one order, which generally corresponds to the combined radio mode of the entire vehicle; when a user says "book a plane ticket", booking a plane ticket is an independent behavior of a single person, which generally corresponds to the independent radio mode of the entire vehicle; when a user says "play an encyclopedia quiz competition", multiple people are required to participate and score separately, which generally corresponds to the independent radio mode of the entire vehicle.

[0118] The whole-vehicle combined radio mode means that each vehicle-mounted radio device receives voice information from all directions of the vehicle, and summarizes and sends it to the application working in the current whole-vehicle combined radio mode, and then the application summarizes and processes the summarized voice information.

[0119] For example, when the main driver says "order takeout", the takeout app is turned on. In the integrated radio mode of the whole car, users in each seat can express their needs. The main driver says "a pizza", and a "pizza" will be added to the order. The co-driver says "I want fried chicken", and another "fried chicken" will be added to the same order. If more than one is needed, you can also add the quantity "order two fried chickens". If you want to cancel, you can also say "cancel one fried chicken". The main driver then places the order and pays. At this time, the in-vehicle radio equipment receives the order requirements of each user and sends the order requirements to the takeout app, completing the takeout ordering process together. The co-driver can also say "order takeout", triggering the integrated radio mode of the whole car, and repeat the above takeout ordering process.

[0120] The independent radio mode for the entire vehicle means that each vehicle-mounted radio device only receives the voice information of the user in the seat corresponding to the vehicle-mounted radio device, that is, each vehicle-mounted radio device only receives voice information with strong signals, filters out voice information with weak signals, and sends them separately to the application working in the current independent radio mode for the entire vehicle, and then the application processes the aggregated voice information separately.

[0121] For example, when a user says "play the encyclopedia quiz competition", the Encyclopedia Quiz application is turned on. In the independent audio reception mode of the entire vehicle, when multiple users speak at the same time, each vehicle-mounted audio device only receives the voice information of the user in the seat corresponding to the vehicle-mounted audio device, and suppresses the voice information from other seats. It recognizes the voice information of the user in the seat corresponding to the vehicle-mounted audio device and sends the voice information to the Encyclopedia Quiz application. The Encyclopedia Quiz application understands and judges the voice information and records the scores respectively.

[0122] For example, if the driver says "book a flight," the ticket booking app opens. In the vehicle's independent audio mode, the driver's seat's onboard audio device receives the driver's voice message, completing the flight booking process. Meanwhile, if the co-driver says "book a movie ticket," the co-driver's onboard audio device, in the vehicle's independent audio mode, receives the co-driver's voice message, completing the movie ticket booking process. Both the driver's and co-driver's flight booking processes can proceed simultaneously.

[0123] Please refer to Figure 5 In this embodiment, the analysis module 200 includes: an analysis unit 2001, used to determine the application that executes the voice command based on the voice command; and a selection unit 2002, used to determine the voice system working mode according to the business type of the application that executes the voice command.

[0124] The analysis unit 2001 determines which application should execute the voice command based on the content of the voice command. For example, if a user says "order takeout," the "order takeout" command is executed by a takeout application, such as Ele.me or Meituan Waimai. If a user says "book a flight," the "book a flight" command is executed by a travel ticketing application, such as Ctrip, Qunar, or Fliggy.

[0125] In this embodiment, the voice system operating mode used is determined by the service type of the application executing the voice command. Each application has its own corresponding voice system operating mode, and the selection unit 2002 selects the voice system operating mode based on the application type. This allows the voice system operating mode to be switched based on user intent, thereby improving the continuous voice usage experience.

[0126] For example, when a user says "order takeout", the "order takeout" command is executed by the takeout application. Ordering takeout means ordering multiple items to generate an order, which can be ordered by one person or multiple people. Therefore, the takeout application uses the full-vehicle integrated radio mode.

[0127] For another example, when a user says "book a plane ticket", the "book a plane ticket" command is executed by a travel ticketing application. Booking a plane ticket requires accurate confirmation of personal identity information and travel time, and does not require multiple people to operate. Therefore, travel ticketing applications use an independent radio mode for the entire vehicle.

[0128] For example, when a user says "play the encyclopedia quiz contest", the command "play the encyclopedia quiz contest" is executed by the encyclopedia quiz application. The encyclopedia quiz contest game is set up in a mode where multiple people participate and each person is scored separately. Therefore, the encyclopedia quiz application uses an independent radio reception mode for the entire vehicle.

[0129] Please continue to refer to Figure 4 The voice collection module 300 is used to obtain a number of voice messages.

[0130] Several voice messages are sent by one or more users and collected by various in-vehicle audio equipment.

[0131] In one embodiment, the plurality of voice messages are collected by different audio receivers. In this case, the plurality of voice messages may be sent by one person or multiple people, and each is collected by the vehicle-mounted audio receiver corresponding to the user's seat. This utilizes a fully independent audio receiver mode.

[0132] In another embodiment, several voice messages are sent by different people and collected by the same audio receiving device. In this case, the voice messages are sent by multiple users and collected by a single vehicle-mounted audio receiving device. In this case, the vehicle-wide combined audio receiving mode is used.

[0133] In another embodiment, several voice messages are sent by different people and collected by different audio receivers. In this case, the voice messages are sent by multiple users and collected by the in-vehicle audio receivers corresponding to the seats where the users are sitting. In this case, the whole vehicle independent audio receiver mode is used.

[0134] In the vehicle-wide independent audio reception mode, the audio reception device collects voice information by collecting strong voice information and filtering out weak voice information based on the signal strength of the collected voice information. That is, any audio reception device only collects the voice of the user in the seat corresponding to the audio reception device. For example, the audio reception device on the driver's door collects the voice of the user in the driver's seat and filters out the voice of users in other seats.

[0135] Please continue to refer to Figure 4 The voice processing module 400 is used to merge or independently process several voice messages according to the working mode of the voice system to obtain processing results.

[0136] The processing result is a combination of several voice messages or independent processing, and the final user demand is obtained, which requires the corresponding application to perform the next step of processing.

[0137] In the whole-vehicle independent sound receiving mode, the voice processing module 400 processes a plurality of voice messages independently.

[0138] In one embodiment, in the vehicle-wide independent sound reception mode, the voice processing module 400 treats the voice information collected by any sound reception device as a set and processes the voice information in the set to obtain a processing result. The processing results corresponding to each sound reception device are independent.

[0139] For example, when the driver says "book a flight," the ticket booking app opens. In the independent car-wide audio mode, the onboard audio device in the driver's seat receives the voice information of the user in the driver's seat, completing the flight booking process. If the co-driver wants to "book movie tickets," the co-driver says "book movie tickets." In the independent car-wide audio mode, the onboard audio device in the co-driver's seat receives the voice information of the user in the co-driver's seat, completing the movie ticket booking process. At this time, the voice of the driver collected by the onboard audio device in the driver's seat is processed for flight booking, and the voice of the co-driver collected by the onboard audio device in the co-driver's seat is processed for movie ticket booking, obtaining independent processing results. The processing results at this time are the operations of booking flights or movie tickets in the travel ticketing software.

[0140] In another embodiment, in the vehicle-wide independent sound reception mode, the voice processing module 400 treats the voice information sent by a user as a set and processes the voice information in the set to obtain a processing result. The sound reception device has a voiceprint recognition function and can distinguish different users by recognizing voiceprints.

[0141] In the whole-vehicle combined sound reception mode, the voice processing module 400 combines and processes a plurality of voice messages.

[0142] At this time, regardless of whether the voice is collected by one or multiple sound receiving devices, or is emitted by one user or multiple users, the voice processing module 400 superimposes and accumulates the contents of multiple voice messages to obtain a processing result.

[0143] In one embodiment, in the vehicle-wide combined sound recording mode, when the time interval between any two voice messages is less than a preset value, the voice processing module 400 combines the two voice messages into one voice message.

[0144] For example, if one person says "today" and another says "weather" after the word "day", they will be merged into one voice message "today's weather".

[0145] In another embodiment, in the vehicle-wide combined recording mode, after acquiring a preset number of voice messages, the voice processing module 400 combines the preset number of voice messages into one voice message. That is, during the continuous recording process, several consecutive voice messages are set as one voice message, and the contents of the multiple voice messages are then superimposed and accumulated to obtain a processing result.

[0146] Figure 6 It is a structural diagram of an in-vehicle voice interaction system in another disclosed embodiment.

[0147] Please refer to Figure 6 , Figure 6 For Figure 4Based on the structural diagram, the structure of the vehicle voice interaction system includes:

[0148] A command receiving module 100 is used to obtain voice commands;

[0149] An analysis module 200 is configured to determine a voice system operating mode based on the voice command;

[0150] The voice collection module 300 is used to obtain a number of voice messages;

[0151] The voice processing module 400 is used to combine or independently process multiple voice messages according to the voice system working mode to obtain processing results;

[0152] The result output module 500 is used to output the processing result.

[0153] For detailed description of the command receiving module 100, the analyzing module 200, the voice collecting module 300 and the voice processing module 400, please refer to Figure 4 and Figure 5 , I will not go into details here.

[0154] Please continue to refer to Figure 6 , the result output module 500 is used to output the processing result.

[0155] In one embodiment, the result output module 500 includes: a display unit for displaying the processing result, that is, displaying the result of independent processing or combined processing of multiple voice messages on the vehicle's central control display interface.

[0156] For example, when a user says "play the encyclopedia quiz competition", the independent radio mode of the entire car is turned on. The encyclopedia quiz competition game is set up in a mode where multiple people participate and each person is scored separately. Each car-mounted radio device only receives the voice information of the user in the seat corresponding to the car-mounted radio device, and sends the voice information to the encyclopedia quiz application. The encyclopedia quiz application understands and judges the voice information, and records the scores separately. Finally, the scores of the users corresponding to each car-mounted radio device are displayed on the car's central control display interface.

[0157] In another embodiment, the result output module 500 includes: a voice output unit for outputting the processing result by voice, that is, the result of independently processing or merging multiple voice messages is announced by the vehicle-mounted voice equipment.

[0158] For example, when the user says "book a flight", the independent radio mode of the entire vehicle is turned on. The user states his or her travel time and other requirements under the guidance of the travel ticketing application. The travel ticketing application selects eligible flights based on the travel time and other requirements and announces them by voice. The user then confirms which flight to book based on the announced content and makes a reservation.

[0159] In another embodiment, the result output module 500 includes a processing unit configured to execute the next step according to the content of the processing result, that is, after merging or independently processing a plurality of voice messages, the final user demand obtained requires the corresponding application to perform the next step of processing.

[0160] For example, when a user says "order takeout," the food delivery app opens. With the car's integrated audio system, users in each seat can express their needs. If the driver says "a pizza," "pizza" is added to the order. If the passenger says "fried chicken," "fried chicken" is added to the order. If they want more, they can add the quantity, "order two fried chickens." To cancel, they can say "cancel one fried chicken." Finally, multiple items are ordered into one order. After confirming the order, the user fills in the delivery address and instructs the food delivery app to place the order for delivery.

[0161] Accordingly, the embodiment of the present disclosure further provides a storage medium on which computer instructions are stored, wherein the computer instructions are executed when the computer instructions are executed. Figure 1 and Figure 2 Preferably, the storage medium may include ROM, RAM, magnetic disk or optical disk.

[0162] Accordingly, the embodiment of the present disclosure further provides a storage medium on which computer instructions are stored, wherein the computer instructions are executed when the computer instructions are executed. Figure 3 Preferably, the storage medium may include ROM, RAM, magnetic disk or optical disk.

[0163] Accordingly, an embodiment of the present disclosure further provides a terminal, comprising a memory and a processor, wherein the memory stores computer instructions that can be executed on the processor, and wherein the processor executes the computer instructions. Figure 1 and Figure 2 The steps of the method. Preferably, the terminal may include the vehicle computer.

[0164] Accordingly, an embodiment of the present disclosure further provides a terminal, comprising a memory and a processor, wherein the memory stores computer instructions that can be executed on the processor, and wherein the processor executes the computer instructions. Figure 3The steps of the method. Preferably, the terminal may include the vehicle computer.

[0165] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention should be based on the scope defined by the claims.

Claims

1. A method for in-vehicle voice interaction, characterized in that: include: Get voice commands; Determining a voice system operating mode according to the voice command; Get several voice messages; According to the working mode of the voice system, several voice messages are processed together or independently to obtain the processing results; The method for determining the working mode of the voice system according to the voice command includes: Determining, based on the voice command, an application for executing the voice command; Determine the voice system operating mode based on the service type of the application that executes the voice command; each application has a corresponding voice system operating mode; The voice system working mode includes a full-vehicle merged sound reception mode; in the full-vehicle merged sound reception mode, when the time interval between any two voice messages is less than a preset value, the two voice messages are merged into one voice message; or, in the full-vehicle merged sound reception mode, after obtaining a preset number of voice messages, the preset number of voice messages are merged into one voice message.

2. In the in-vehicle voice interaction method as described in claim 1, several voice messages are collected by different sound receiving devices respectively; or, several voice messages are sent by different people and collected by the same sound receiving device; or, several voice messages are sent by different people and collected by different sound receiving devices.

3. The method for in-vehicle voice interaction as described in claim 2, wherein the voice system operating mode includes a whole-vehicle independent radio mode or a whole-vehicle combined radio mode; in the whole-vehicle independent radio mode, multiple voice messages are processed independently; in the whole-vehicle combined radio mode, multiple voice messages are combined for processing.

4. The in-vehicle voice interaction method of claim 3, wherein the method of independently processing multiple voice messages in the vehicle-wide independent sound reception mode comprises: According to different sound receiving devices, several voice messages are processed separately to obtain independent processing results.

5. The in-vehicle voice interaction method according to claim 4, wherein the method for collecting voice information by the sound receiving device comprises: The sound receiving device collects the voice information with strong signals and filters out the voice information with weak signals according to the signal strength of the collected voice information.

6. The in-vehicle voice interaction method of claim 3, wherein in the vehicle-wide independent sound reception mode, the method for independently processing multiple voice messages comprises: According to the different senders, several voice messages are processed separately to obtain independent processing results.

7. The in-vehicle voice interaction method according to claim 1, further comprising: Output the processing result, and the way of outputting the processing result includes: displaying the processing result, outputting the processing result by voice, or executing the next step of the process according to the content of the processing result.

8. The in-vehicle voice interaction method according to claim 1, wherein the method for merging multiple voice messages comprises: The contents of several voice messages are superimposed and accumulated to obtain a processing result.

9. A vehicle-mounted voice interaction system, characterized in that: include: A command receiving module is used to obtain voice commands; An analysis module, configured to determine a voice system operating mode based on the voice command; Voice collection module, used to obtain several voice messages; The voice processing module is used to combine or independently process several voice messages according to the working mode of the voice system and obtain the processing results; Wherein, the analysis module includes: an analyzing unit, configured to determine, based on the voice command, an application that executes the voice command; A selection unit, configured to determine a voice system operating mode according to a service type of an application executing the voice command; each application has a corresponding voice system operating mode; The voice system working mode includes a vehicle-wide combined radio mode; In the full-vehicle merged reception mode, when the time interval between any two voice messages is less than a preset value, the voice processing module is used to merge the two voice messages into one voice message; or, in the full-vehicle merged reception mode, after obtaining a preset number of voice messages, the voice processing module is used to merge the preset number of voice messages into one voice message.

10. In the in-vehicle voice interaction system as described in claim 9, several pieces of voice information are collected by different sound receiving devices respectively; or, several pieces of voice information are sent by different people and collected by the same sound receiving device; or, several pieces of voice information are sent by different people and collected by different sound receiving devices.

11. The in-vehicle voice interaction system as described in claim 10, wherein the voice system operating mode includes a whole-vehicle independent sound reception mode or a whole-vehicle combined sound reception mode; in the whole-vehicle independent sound reception mode, the voice processing module independently processes multiple voice messages; in the whole-vehicle combined sound reception mode, the voice processing module combines multiple voice messages for processing.

12. The in-vehicle voice interaction system according to claim 9, further comprising: The result output module is used to output the processing result.

13. The in-vehicle voice interaction system according to claim 12, wherein the result output module comprises: A display unit is used to display the processing result.

14. The in-vehicle voice interaction system according to claim 12, wherein the result output module comprises: The voice output unit is used to output the processing result in voice.

15. The in-vehicle voice interaction system according to claim 12, wherein the result output module comprises: The processing unit is used to execute the next step of the process according to the content of the processing result.

16. A storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed, the steps of the method according to any one of claims 1 to 8 are executed.

17. A terminal comprising a memory and a processor, wherein the memory stores computer instructions that can be executed on the processor, characterized in that: When the processor executes the computer instructions, the steps of the method according to any one of claims 1 to 8 are performed.

Citation Information

Patent Citations

  • Voice interaction method and device for vehicle-mounted system, automobile and machine readable medium

    CN110070868A

  • Method for processing voice signals of multiple speakers, and electronic device according thereto

    WO2019124742A1