Voice interaction method, system and related device

By displaying dialogue logos and sub-signs on the display screen of smart cars, the problem that smart cars cannot handle multiple voice commands at the same time is solved, and multiple voice interaction is realized, improving user experience.

CN120071941APending Publication Date: 2025-05-30HUAWEI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311655712.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When multiple users in the car issue different voice commands at the same time, smart cars cannot process multiple voice commands at the same time, affecting the user's user experience.

Method used

Multiple voice interaction is realized by displaying dialogue logos and sub-identities on the vehicle's display screen, allowing the vehicle to have voice interaction with multiple users at the same time.

Benefits of technology

It improves the efficiency of voice interaction, provides better voice interaction services, and meets the needs of multiple users to use simultaneously.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071941A_ABST
    Figure CN120071941A_ABST
Patent Text Reader

Abstract

The invention discloses a voice interaction method and system and a related device, and is applied to a vehicle, and the vehicle comprises a first display screen; the method comprises the following steps: receiving a first voice sent by a first user, wherein the first voice comprises a wake-up word; in response to the first voice, displaying a first dialogue identifier on the first display screen, the first dialogue identifier being used for prompting a user that the first display screen is performing voice interaction with the first user; receiving a second voice sent by a second user; when the second voice comprises the wakeup-free instruction, in response to the second voice, displaying a voice sub-identifier on the first display screen, the voice sub-identifier being used for prompting a user that the first display screen is processing multiple voices. In this way, the vehicle can process the voice instructions sent by the multiple users at the same time, and better voice interaction service is provided for the users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronic technologies, and in particular, to a voice interaction method, system, and related devices. Background Art

[0002] With the continuous development of electronic technologies, the functions of intelligent vehicles have also tended to be diversified. To provide better services to users, more and more intelligent vehicles are equipped with a voice assistant function. In this way, when a user is inside an intelligent vehicle, the intelligent vehicle can receive and respond to a voice command issued by the user, and execute the function indicated by the voice command, such as playing audio, turning on navigation, etc.

[0003] However, if there are multiple users in the vehicle and multiple users issue different voice commands simultaneously, the intelligent vehicle cannot process multiple voice commands at the same time, which will affect the user experience. Summary of the Invention

[0004] This application provides a voice interaction method, system, and related devices, which realizes the simultaneous processing of different voice commands of multiple users and can provide better voice interaction services for users.

[0005] In a first aspect, this application provides a voice interaction method, which is applied to a vehicle, and the vehicle includes a first display screen; the method includes:

[0006] Receiving a first voice issued by a first user, the first voice including a wake-up word; in response to the first voice, displaying a first conversation identifier on the first display screen, the first conversation identifier being used to prompt the user that the first display screen is performing a voice interaction with the first user; receiving a second voice issued by a second user; when the second voice includes a wake-up-free command, in response to the second voice, displaying a voice sub-identifier on the first display screen, the voice sub-identifier being used to prompt the user that the first display screen is processing multiple voices.

[0007] In this way, the vehicle can process the wake-up-free command of the second user through the first display screen while performing a voice interaction between the first display screen and the first user, improving the voice interaction efficiency and providing better voice interaction services for users.

[0008] In a possible implementation, the first voice further includes a first command, and the first command is used to instruct the vehicle to perform a first operation; the method further includes: in response to the first voice, performing the first operation.

[0009] In this way, after receiving the first voice, the vehicle can perform the operation corresponding to the first command in the first voice, such as turning on navigation, playing music, opening a window, etc.

[0010] In a possible implementation, after displaying the first conversation identifier on the first display screen, the method further includes: receiving a third voice emitted by a first user, where the third voice is used to instruct the vehicle to perform a first operation; and in response to the third voice, performing the first operation.

[0011] In this way, after the first user activates the voice assistant function of the first display screen by using a wake-up word, the first user can also issue a voice command. In this case, the first display screen has already started voice interaction with the first user in response to the first user's wake-up word. Therefore, the first display screen can continue to receive the voice commands issued by the user and perform the operations corresponding to the voice commands, such as opening the navigation, playing music, opening the window, etc.

[0012] In a possible implementation, the method further includes: after displaying the first conversation identifier on the first display screen, outputting a first feedback, where the first feedback is used to prompt the user that the vehicle has received the first voice; and after displaying a voice sub-identifier on the first display screen, outputting a second feedback, where the second feedback is used to prompt the user that the vehicle has received the second voice.

[0013] The vehicle can output feedback information (such as the first feedback, the second feedback, etc.) in one or more ways, such as display on the display screen, voice broadcast, indicator light flashing, vibration, etc. The output methods of the first feedback and the second feedback can be different, or the output methods of the first feedback and the second feedback can be the same.

[0014] In this way, the user can be prompted that the vehicle has received the voice issued by the user by outputting the feedback information.

[0015] In a possible implementation, the first feedback is further used to prompt the user whether the operation indicated by the first voice has been performed, and the second feedback is further used to prompt the user whether the operation indicated by the second voice has been performed.

[0016] In this way, the user can be prompted about the execution status of the voice commands issued by the user by the feedback information, such as prompting the user that the execution is completed, etc.

[0017] In a possible implementation, the method further includes: when the second voice includes a wake-up word, stopping the display of the first conversation identifier and displaying a second conversation identifier on the first display screen, where the second conversation identifier is used to prompt the user that the first display screen is performing voice interaction with a second user.

[0018] In this way, when receiving the wake-up word issued by the second user, the vehicle can end the voice interaction between the first display screen and the first user and perform voice interaction with the second user through the first display screen. It can be ensured that at the same moment, the vehicle only needs to process one voice interaction activated by a wake-up word.

[0019] In a possible implementation, the method further includes: when the second voice includes a wake-up word, outputting an interruption prompt, where the interruption prompt is used to prompt that the voice interaction between the first user and the first display screen is interrupted.

[0020] In a possible implementation, the vehicle can output the interruption prompt in one or more ways such as display on the display screen, flashing of the indicator light, voice broadcast, vibration, etc.

[0021] In this way, the first user can be prompted that the voice interaction with the first display screen has ended through the interruption prompt.

[0022] In a possible implementation, the vehicle interior includes a first sound zone and a second sound zone. The first user is located in the first sound zone and the second user is located in the second sound zone; the first display screen is used to respond to voices from the first sound zone and the second sound zone; displaying a first conversation identifier on the first display screen specifically includes: determining that the first voice comes from the first sound zone based on the first voice; displaying the first conversation identifier on the first display screen based on the first sound zone; displaying a voice sub-identifier on the first display screen specifically includes: determining that the second voice comes from the second sound zone based on the second voice; displaying the voice sub-identifier on the first display screen based on the second sound zone.

[0023] In this way, when the vehicle receives the user's voice, it can determine the display screen for displaying the voice identifier (such as the first conversation identifier, voice sub-identifier, etc.) based on the corresponding relationship between the sound zone where the voice comes from and the display screen.

[0024] In a possible implementation, the vehicle interior includes a first sound zone and a second sound zone. The first user is located in the first sound zone and the second user is located in the second sound zone; the first display screen is used to respond to voices from the first sound zone and the second sound zone; displaying a first conversation identifier on the first display screen specifically includes: determining that the first voice comes from the first sound zone based on the first voice; displaying the first conversation identifier on the first display screen based on the first sound zone; displaying a second conversation identifier on the first display screen specifically includes: determining that the second voice comes from the second sound zone based on the second voice; displaying the second conversation identifier on the first display screen based on the second sound zone.

[0025] In this way, when the vehicle receives the user's voice, it can determine the display screen for displaying the voice identifier (such as the first conversation identifier, second conversation identifier, etc.) based on the corresponding relationship between the sound zone where the voice comes from and the display screen.

[0026] In a possible implementation, the method further includes: in response to the first voice, displaying a first indicator on the first display screen, where the first indicator is used to indicate the position of the first sound zone relative to the first display screen; when the second voice includes a hands-free wake-up instruction, in response to the second voice, replacing the first indicator with a second indicator, where the second indicator is used to indicate the positions of the first sound zone and the second sound zone relative to the display screen; when the second voice includes a wake-up word, in response to the second voice, replacing the first indicator with a third indicator, where the third indicator is used to indicate the position of the second sound zone relative to the display screen.

[0027] In this way, the user can be prompted about the source sound zone of the voice currently being processed by the display screen through sound zone indicators (such as the first indicator, the second indicator, the third indicator, etc.).

[0028] In a possible implementation, the voice interaction mode between the first display screen and the first user is full-duplex interaction; when the second voice includes a hands-free wake-up instruction, the voice interaction mode between the first display screen and the second user is single-round interaction; when the second voice includes a wake-up word, the voice interaction mode between the first display screen and the second user is full-duplex interaction.

[0029] In this way, it can be determined that the voice interaction mode enabled by the wake-up word is full-duplex interaction, and the voice interaction mode enabled by the hands-free wake-up instruction is single-round interaction. In this way, it can ensure that timely voice interaction services are provided to more users.

[0030] In a second aspect, the present application provides a voice interaction method applied to a vehicle, where the vehicle includes a first display screen; the method includes: receiving a fourth voice emitted by a first user, where the fourth voice includes a hands-free wake-up instruction; in response to the fourth voice, and displaying a first hands-free wake-up identifier on the first display screen, where the first hands-free wake-up identifier is used to prompt the user that the first display screen is processing a hands-free wake-up instruction; receiving a fifth voice emitted by a second user; when the fifth voice includes a hands-free wake-up instruction, in response to the fifth voice, stopping the display of the first hands-free wake-up identifier, and displaying a composite hands-free wake-up identifier on the first display screen, where the composite hands-free wake-up identifier is used to prompt the user that the first display screen is processing multiple hands-free wake-up instructions.

[0031] In this way, the vehicle can process hands-free wake-up instructions issued by multiple users through the first display screen, improving the voice interaction efficiency and providing better voice interaction services for users.

[0032] In a possible implementation, the method further includes: when the fifth voice includes a wake-up word, in response to the fifth voice, stopping the display of the first hands-free wake-up identifier, and displaying a third conversation identifier and a voice sub-identifier on the first display screen, where the third conversation identifier is used to prompt the user that the first display screen is having a voice interaction with the second user, and the voice sub-identifier is used to prompt the user that the first display screen is processing multiple voices.

[0033] In this way, while the vehicle processes the wake-up-free instruction issued by the first user through the first display screen, it can receive the wake-up word issued by the second user and conduct voice interaction with the second user through the first display screen.

[0034] In a possible implementation, the interior of the vehicle includes a first sound zone and a second sound zone. The first user is located in the first sound zone, and the second user is located in the second sound zone. The first display screen is used to respond to voices from the first sound zone and the second sound zone. Displaying the first wake-up-free identifier on the first display screen specifically includes: determining that the fourth voice comes from the first sound zone based on the fourth voice; displaying the first wake-up-free identifier on the first display screen based on the first sound zone. Displaying the composite wake-up-free identifier on the first display screen specifically includes: determining that the fifth voice comes from the second sound zone based on the fifth voice; displaying the composite wake-up-free identifier on the first display screen based on the second sound zone.

[0035] In this way, when the vehicle receives the wake-up-free instruction from the user, it can determine the display screen for displaying the wake-up-free identifier based on the corresponding relationship between the sound zone where the voice comes from and the display screen.

[0036] In a possible implementation, the interior of the vehicle includes a first sound zone and a second sound zone. The first user is located in the first sound zone, and the second user is located in the second sound zone. The first display screen is used to respond to voices from the first sound zone and the second sound zone. Displaying the first wake-up-free identifier on the first display screen specifically includes: determining that the fourth voice comes from the first sound zone based on the fourth voice; displaying the first wake-up-free identifier on the first display screen based on the first sound zone. Displaying the third dialogue identifier and the voice sub-identifier on the first display screen specifically includes: determining that the fifth voice comes from the second sound zone based on the fifth voice; displaying the third dialogue identifier and the voice sub-identifier on the first display screen based on the second sound zone.

[0037] In this way, when the vehicle receives the wake-up-free instruction or the wake-up word from the user, it can determine the display screen for displaying the wake-up-free identifier based on the corresponding relationship between the sound zone where the voice comes from and the display screen.

[0038] In a possible implementation, the method further includes: when the fifth voice includes a wake-up word, in response to the fifth voice, displaying a fourth indicator on the first display screen, and the fourth indicator is used to indicate the positions of the first sound zone and the second sound zone relative to the display screen.

[0039] In this way, when receiving the wake-up word, the fourth indicator can be displayed on the first display screen.

[0040] In a possible implementation, the method further includes: in response to a fourth voice, displaying a first indicator on the first display screen, where the first indicator is used to indicate the position of the first sound zone relative to the first display screen; when the fifth voice includes a wake-up-free instruction, in response to the fifth voice, replacing the first indicator with a second indicator, where the second indicator is used to indicate the positions of the first sound zone and the second sound zone relative to the display screen; when the fifth voice includes a wake-up word, in response to the fifth voice, replacing the first indicator with a fourth indicator, where the fourth indicator is used to indicate the positions of the first sound zone and the second sound zone relative to the display screen.

[0041] In this way, the source sound zone of the voice currently processed by the display screen can be prompted to the user through sound zone indicators (such as the first indicator, the second indicator, the fourth indicator, etc.).

[0042] In a possible implementation, the voice interaction mode between the first display screen and the first user is single-round interaction; when the fifth voice includes a wake-up-free instruction, the voice interaction mode between the first display screen and the second user is single-round interaction; when the fifth voice includes a wake-up word, the voice interaction mode between the first display screen and the second user is full-duplex interaction.

[0043] In this way, it can be determined that the voice interaction mode enabled by the wake-up word is full-duplex interaction, and the voice interaction mode enabled by the wake-up-free instruction is single-round interaction. In this way, it can be ensured that timely voice interaction services are provided to more users.

[0044] In a third aspect, the present application provides a voice interaction method applied to a vehicle. The vehicle includes a first display screen and a second display screen, and the interior of the vehicle includes a first sound zone and a second sound zone; the first display screen is used to respond to voices from the first sound zone, and the second display screen is used to respond to voices from the second sound zone; the method includes: receiving a sixth voice emitted by a first user from the first sound zone, where the sixth voice includes a wake-up word; in response to the sixth voice, determining based on the sixth voice that the sixth voice comes from the first sound zone; based on the first sound zone, displaying a third conversation identifier on the first display screen, where the third conversation identifier is used to prompt the user that the first display screen is performing a voice interaction with the first user; receiving a seventh voice emitted by a second user from the second sound zone; when the seventh voice includes a wake-up-free instruction, in response to the seventh voice, determining based on the seventh voice that the seventh voice comes from the second sound zone; based on the second sound zone, displaying a second wake-up-free identifier on the second display screen, where the second wake-up-free identifier is used to prompt the user that the second display screen is processing a wake-up-free instruction.

[0045] In this way, based on the correspondence between the sound zone and the display screen, voice interaction can be performed with different users through different display screens, improving the voice interaction efficiency and providing better voice interaction services for users.

[0046] In a possible implementation, the method further includes: when the seventh voice includes a wake-up word, in response to the seventh voice, stopping displaying the third conversation identifier on the first display screen; determining, based on the seventh voice, that the seventh voice comes from the second sound zone; and displaying, on the second display screen based on the second sound zone, a fourth conversation identifier for prompting the user that the second display screen is performing a voice interaction with the second user.

[0047] In this way, when receiving the wake-up word issued by the second user, the vehicle can end the voice interaction between the first display screen and the first user and perform a voice interaction with the second user through the second display screen. It can be ensured that at the same moment, the vehicle only needs to process one voice interaction initiated by the wake-up word.

[0048] In a possible implementation, the voice interaction mode between the first display screen and the first user is full-duplex interaction; when the seventh voice includes a hands-free wake-up instruction, the voice interaction mode between the second display screen and the second user is single-round interaction; when the seventh voice includes a wake-up word, the voice interaction mode between the second display screen and the second user is full-duplex interaction.

[0049] In this way, it can be determined that the voice interaction mode initiated by the wake-up word is full-duplex interaction, and the voice interaction mode initiated by the hands-free wake-up instruction is single-round interaction. In this way, it can be ensured to provide timely voice interaction services for more users.

[0050] In a fourth aspect, the present application provides a voice interaction method applied to a vehicle. The vehicle includes a first display screen and a second display screen, and the vehicle interior includes a first sound zone and a second sound zone; the first display screen is used to respond to voices from the first sound zone, and the second display screen is used to respond to voices from the second sound zone; the method includes: receiving an eighth voice from the first sound zone of the first user, where the eighth voice includes a hands-free wake-up instruction; in response to the eighth voice, determining, based on the eighth voice, that the eighth voice comes from the first sound zone; displaying, on the first display screen based on the first sound zone, a third hands-free wake-up identifier for prompting the user that the first display screen is processing the hands-free wake-up instruction; receiving a ninth voice from the second sound zone of the second user; when the ninth voice includes a hands-free wake-up instruction, in response to the ninth voice, determining, based on the ninth voice, that the ninth voice comes from the second sound zone; and displaying, on the second display screen based on the second sound zone, a fourth hands-free wake-up identifier for prompting the user that the second display screen is processing the hands-free wake-up instruction.

[0051] In this way, based on the corresponding relationship between the sound zone and the display screen, different hands-free wake-up instructions of different users can be processed by different display screens, improving the voice interaction efficiency and providing better voice interaction services for users.

[0052] In a possible implementation, the method further includes: when the ninth voice includes a wake-up word, in response to the ninth voice, determining that the ninth voice comes from the second sound zone based on the ninth voice; and displaying a fifth conversation identifier on the second display screen based on the second sound zone, where the fifth conversation identifier is used to prompt the user that the second display screen is performing voice interaction with the second user.

[0053] In this way, when receiving the wake-up word sent by the second user, the vehicle can, while processing the hands-free instruction of the first user through the first display screen, perform voice interaction with the second user through the second display screen.

[0054] In a possible implementation, the voice interaction mode between the first display screen and the first user is single-round interaction; when the ninth voice includes a hands-free instruction, the voice interaction mode between the second display screen and the second user is single-round interaction; when the ninth voice includes a wake-up word, the voice interaction mode between the second display screen and the second user is full-duplex interaction.

[0055] In this way, it can be determined that the voice interaction mode enabled by the wake-up word is full-duplex interaction, and the voice interaction mode enabled by the hands-free instruction is single-round interaction. In this way, it can ensure that timely voice interaction services are provided for more users.

[0056] In a fifth aspect, the present application provides a vehicle, including one or more processors and one or more memories. The one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer program code, where the computer program code includes computer instructions. When the one or more processors execute the computer instructions, the vehicle executes the voice interaction method in any possible implementation of any of the above aspects.

[0057] In a sixth aspect, an embodiment of the present application provides a computer storage medium, including computer instructions. When the computer instructions run on the vehicle, the vehicle executes the voice interaction method in any possible implementation of any of the above aspects.

[0058] In a seventh aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on the vehicle, the vehicle executes the voice interaction method in any possible implementation of any of the above aspects.

[0059] The beneficial effects of the fifth aspect to the seventh aspect can refer to the beneficial effects of the first aspect to the fourth aspect above. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figures 1A - 1D It is a schematic diagram of four voice interaction modes provided by an embodiment of the present application;

[0061] Figure 2Schematic diagram of the system architecture of a voice interaction system 1000 provided by an embodiment of the present application;

[0062] Figures 3A - 3B Schematic diagrams of the internal forms of two intelligent vehicles 100 provided by an embodiment of the present application;

[0063] Figures 3C - 3D Schematic diagrams of the sound zone distributions inside two intelligent vehicles 100 provided by an embodiment of the present application;

[0064] Figures 3E - 3H Schematic diagrams of the corresponding relationships between multiple sound zones and display screens provided by an embodiment of the present application;

[0065] Figure 4A Schematic diagram of the hardware structure of an intelligent vehicle 100 provided by an embodiment of the present application;

[0066] Figure 4B Schematic diagram of the hardware structure of a mobile phone 200 provided by an embodiment of the present application;

[0067] Figures 5A - 5F Schematic diagram of the interface for a set of central control screens 10 to process multiple voice channels simultaneously provided by an embodiment of the present application;

[0068] Figures 6A - 6D Schematic diagram of the interface for a set of central control screens 10 and a co-pilot screen 20 to process different voices simultaneously respectively provided by an embodiment of the present application;

[0069] Figures 7A - 7D Schematic diagram of the interface for a set of central control screens 10 and a right rear projection screen 30 to process different voices simultaneously respectively provided by an embodiment of the present application;

[0070] Figures 8A - 8C Schematic diagram of the interface for a set of central control screens 10 to process multiple wake-up-free instructions simultaneously provided by an embodiment of the present application;

[0071] Figures 9A - 9D Schematic diagram of the interface for a set of central control screens 10 and a right rear projection screen 30 to process different voices simultaneously respectively provided by an embodiment of the present application;

[0072] Figures 10A - 10C Schematic diagram of the interface for a set of central control screens 10 to process multiple voice channels simultaneously provided by an embodiment of the present application;

[0073] Figures 11A - 11D Schematic diagram of the interface for a set of central control screens 10 and a right rear projection screen 30 to process wake-up words issued by different users successively provided by an embodiment of the present application;

[0074] Figures 12A - 12B Schematic diagram of the interface for a set of central control screens 10 to process wake-up words issued by different users successively provided by an embodiment of the present application;

[0075] Figure 13 Schematic flowchart of a voice interaction method provided by an embodiment of the present application;

[0076] Figure 14 Schematic flowchart after an intelligent vehicle 100 receives a wake-up-free instruction provided by an embodiment of the present application;

[0077] Figure 15A Schematic diagram of an application scenario of a voice interaction method provided by an embodiment of the present application;

[0078] Figures 15B - 15E Schematic diagrams of interfaces of different display screens in a group of application scenarios provided by an embodiment of the present application;

[0079] Figure 16 Schematic diagram of functional modules of an intelligent vehicle 100 provided by an embodiment of the present application;

[0080] Figure 17 Schematic flowchart of a voice interaction method provided by an embodiment of the present application;

[0081] Figure 18 Schematic flowchart of another voice interaction method provided by an embodiment of the present application;

[0082] Figure 19 Schematic flowchart of another voice interaction method provided by an embodiment of the present application;

[0083] Figure 20 Schematic flowchart of another voice interaction method provided by an embodiment of the present application. Detailed implementation manners

[0084] Next, the technical solutions in the embodiments of the present application will be clearly and elaborately described with reference to the accompanying drawings. Among them, in the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may represent A or B; "and / or" in the text is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality of" means two or more than two.

[0085] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more than two.

[0086] In the following embodiments of the present application, the term "user interface (UI)" refers to a media interface for interaction and information exchange between an application or an operating system and a user. It realizes the conversion between the internal form of information and the form acceptable to the user. The user interface is source code written in a specific computer language such as Java or Extensible Markup Language (XML). The interface source code is parsed and rendered on an electronic device and finally presented as content recognizable by the user. The common manifestation form of the user interface is the graphical user interface (GUI), which refers to the user interface related to computer operations presented in a graphical way. It can be visual interface elements such as text, icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and Widgets displayed on the display screen of an electronic device.

[0087] The following introduces some terms related to the embodiments of the present application.

[0088] Wake-up word: A wake-up word is a pre-set word used to trigger an electronic device to turn on (also known as wake up) the voice assistant function (hereinafter referred to as the voice assistant). An electronic device with a voice assistant can be set with one or more wake-up words. When the electronic device detects that the user says the pre-set wake-up word, the electronic device can turn on the voice assistant and start a voice interaction with the user. In this way, it is possible to better distinguish between the user's daily chat scenario and the voice interaction scenario, and it is also possible to avoid the power consumption caused by the long-term activation of the voice assistant.

[0089] Voice command without wake-up: A voice command without wake-up is a pre-set voice command that can trigger an electronic device to perform a pre-set operation. When the electronic device detects that the voice spoken by the user includes a voice command without wake-up, the electronic device can turn on the voice assistant and execute the operation corresponding to the voice command without wake-up (such as opening the window, starting navigation, playing audio, etc.). It should be noted that the difference between a voice command without wake-up and a wake-up word is that the wake-up word is only used to turn on the voice assistant. After turning on the voice assistant, the electronic device needs to determine the operation to be performed based on the voice command spoken by the user later; while a voice command without wake-up can trigger the electronic device to turn on the voice assistant and execute the specified operation when the voice assistant is not turned on, and after the electronic device executes the operation corresponding to the voice command without wake-up, the electronic device usually turns off the voice assistant.

[0090] The following introduces the working modes (also known as voice interaction modes) of various voice assistants involved in the embodiments of the present application.

[0091] Single turn interaction: Single turn interaction means that after the voice assistant is awakened, it executes a complete voice interaction process and then shuts down the working mode of the voice assistant. A complete voice interaction process can include one input and one output, and the input and output cannot be executed simultaneously. That is, at the same moment, the electronic device that activates the voice assistant can only execute the input or the output. Among them, the input refers to receiving the voice issued by the user, and the output refers to the voice feedback for the user's voice input. In the scenario of single turn interaction, the voice assistant needs to be awakened before each voice interaction, and then the voice interaction can be carried out.

[0092] Multi turn interaction: Multi turn interaction means that after the voice assistant is awakened, it can execute multiple complete voice interaction processes, and there is no need to awaken the voice assistant again during these multiple voice interaction processes. It should be noted that in the scenario of multi turn interaction, the input and output cannot be executed simultaneously either.

[0093] Keep listening: Keep listening means that after the voice assistant is awakened, it continuously listens to the user's voice, and at the same time, it can output one or more voice feedbacks based on the listened user's voice. It should be noted that in the scenario of keep listening, the electronic device that activates the voice assistant can execute the input and output at the same moment, or only execute the input or the output.

[0094] Full duplex interaction: Full duplex interaction, abbreviated as full duplex, means that after the voice assistant is awakened, it continuously listens to the user's voice, and at the same time, it continuously outputs voice feedback based on the user's voice. Full duplex is a term in the communication field, and its communication term definition is a real-time, two-way voice information interaction mode. It should be noted that in the scenario of full duplex interaction, the electronic device that activates the voice assistant can conduct real-time, two-way voice interaction with the user.

[0095] It should be noted that in some embodiments, the voice assistant can adopt different working modes for each voice channel.

[0096] Figure 1A The figure shows a schematic diagram of a working mode of single turn interaction provided by an embodiment of the present application.

[0097] As Figure 1A shown, in a one-dimensional coordinate system, the coordinate axis can represent time, and the closer to the positive direction of the coordinate axis (i.e., the arrow direction), the larger the time value. There are three differently colored color blocks shown in the one-dimensional coordinate system, a black color block, a gray color block, and a white color block. Among them, the black color block can represent that the electronic device is in the input stage, the white color block can represent that the electronic device is in the output stage, and the gray color block can represent that the electronic device is in the awakening stage. According to Figure 1AIt can be seen that from time t0 to time t1, the voice assistant of the electronic device is awakened; from time t1 to time t2, the electronic device receives the voice input by the user through the voice assistant; from time t2 to time t3, the electronic device outputs voice through the voice assistant, and the output voice is generated based on the voice input by the user. After the output ends, the voice assistant of the electronic device is turned off. From time t3 to time t4, the voice assistant of the electronic device remains in the off state all the time. From time t4 to time t5, the voice assistant of the electronic device is awakened again; from time t5 to time t6, the electronic device receives the voice input by the user through the voice assistant; from time t6 to time t7, the electronic device outputs voice through the voice assistant, and the output voice is generated based on the voice input by the user. After this output ends, the voice assistant of the electronic device is turned off again.

[0098] It can be understood that Figure 1A The illustrated embodiment is only an example. In the embodiments of the present application, the durations of the wake-up stage, input stage, and output stage of the single-round interaction may also be different from those of the above embodiments, and the present application does not make any limitations here.

[0099] Figure 1B The figure shows a schematic diagram of a working mode of multi-round interaction provided by an embodiment of the present application.

[0100] As Figure 1B shown, in a one-dimensional coordinate system, the coordinate axis can represent time, and the closer to the positive direction of the coordinate axis (i.e., the arrow direction), the larger the time value. There are three color blocks with different colors shown in the one-dimensional coordinate system, a black block, a gray block, and a white block. The stages represented by each block can refer to the relevant descriptions in the above Figure 1A illustrated embodiment. According to Figure 1B It can be seen that from time t10 to time t11, the voice assistant of the electronic device is awakened; from time t11 to time t12, the electronic device receives the voice input by the user through the voice assistant; from time t12 to time t13, the electronic device outputs voice through the voice assistant, and the output voice is generated based on the voice input by the user; from time t13 to time t14, the electronic device receives the voice input by the user through the voice assistant; from time t14 to time t15, the electronic device outputs voice through the voice assistant, and the output voice is generated based on the voice input by the user during the period from time t13 to time t14.

[0101] It can be understood that Figure 1B The illustrated embodiment is only an example. In the embodiments of the present application, the durations of the wake-up stage, input stage, and output stage of the multi-round interaction may also be different from those of the above embodiments, and the multi-round interaction may also include more or fewer voice interaction processes than the above embodiments, and the present application does not make any limitations here.

[0102] Figure 1C It shows a schematic diagram of a working mode of continuous listening provided by an embodiment of the present application.

[0103] As Figure 1C shown, in a one-dimensional coordinate system, the coordinate axis can represent time, and the closer to the positive direction of the coordinate axis (i.e., the arrow direction), the larger the time value. There are three color blocks of different colors shown in the one-dimensional coordinate system, a black color block, a gray color block, and a white color block. The stages represented by each color block can refer to the relevant descriptions in the above Figure 1A shown embodiment. According to Figure 1C it can be known that from time t20 to time t21, the voice assistant of the electronic device is awakened; from time t21 to time t25, the electronic device continuously listens to the voice input by the user through the voice assistant. The voice input by the user can be continuous or intermittent. From time t22 to time t23, the electronic device outputs voice through the voice assistant, and the output voice is generated based on the voice input by the user from time t21 to time t23. After that, from time t24 to time t26, the electronic device can also output voice through the voice assistant, and the voice output this time is generated based on the voice input by the user from time t23 to time t25, where time t25 is earlier than time t26.

[0104] It can be understood that Figure 1C the shown embodiment is just an example. In the embodiments of the present application, the durations of the wake-up stage, input stage, and output stage of continuous listening can also be durations different from those of the above embodiments, and continuous listening can also include more or fewer output stages than the above embodiments. The present application does not make any limitations here.

[0105] Figure 1D It shows a schematic diagram of a working mode of full-duplex interaction provided by an embodiment of the present application.

[0106] As Figure 1D shown, in a one-dimensional coordinate system, the coordinate axis can represent time, and the closer to the positive direction of the coordinate axis (i.e., the arrow direction), the larger the time value. There are three color blocks of different colors shown in the one-dimensional coordinate system, a black color block, a gray color block, and a white color block. The stages represented by each color block can refer to the relevant descriptions in the above Figure 1A shown embodiment. According to Figure 1DIt can be known that from time t30 to time t31, the voice assistant of the electronic device is awakened; from time t31 to time t32, the electronic device continuously listens to the voice input by the user through the voice assistant. The voice input by the user can be continuous or intermittent. In addition, from time t31 to time t32, the electronic device can also perform real-time voice output through the voice assistant, that is, after generating voice for output based on the voice that the user has already input, the electronic device can output the voice in real time.

[0107] It can be understood that Figure 1D The illustrated embodiment is only an example. In the embodiments of the present application, the durations of the full-duplex wake-up stage, input stage, and output stage can also be different from those of the above embodiments, and the present application does not make any limitations here.

[0108] The following introduces a voice interaction system 1000 provided by an embodiment of the present application.

[0109] Figure 2 FIG. shows a schematic diagram of the system architecture of a voice interaction system 1000 provided by an embodiment of the present application.

[0110] As Figure 2 shown, the voice interaction system 1000 may include a smart car 100, and the voice interaction system 1000 may also include any one or more of the following: a mobile phone 200, a headset 300, and a cloud server 400.

[0111] In some embodiments, the smart car 100 and the mobile phone 200 may establish a communication connection. Based on the communication connection between the smart car 100 and the mobile phone 200, the smart car 100 can receive and respond to the user's voice command, and make a call / answer a call through the mobile phone 200.

[0112] In some embodiments, the smart car 100 and the headset 300 may establish a communication connection. The smart car 100 can receive and respond to the user's voice command, and send the specified audio data to the headset 300 through the communication connection between the smart car 100 and the headset 300, and play the audio through the headset 300.

[0113] In other embodiments, the smart car 100 and the mobile phone 200 may establish a communication connection, the smart car 100 and the headset 300 may establish a communication connection, and the mobile phone 200 and the headset 300 may also establish a communication connection. The smart car 100 can receive and respond to the user's voice command, make a call / answer a call through the mobile phone 200, and play the phone audio through the headset 300.

[0114] In some embodiments, the intelligent vehicle 100 may establish a communication connection with the cloud server 400. During the process of the intelligent vehicle 100 executing the operation corresponding to the voice instruction issued by the user, the intelligent vehicle 100 may send a fetch request to the cloud server 400, and the fetch request can be used to request the cloud server 400 to send specified data (such as audio data, video data, picture data, web page data, etc.) to the intelligent vehicle 100. In other embodiments, when the intelligent vehicle 100 cannot accurately recognize the voice issued by the user, the intelligent vehicle 100 may also upload the collected voice to the cloud server 400, and the cloud server 400 analyzes and recognizes the voice and returns the recognition result to the intelligent vehicle 100. The recognition result may include the text content of the voice.

[0115] It can be understood that Figure 2 The illustrated voice interaction system 1000 is just an example. In the embodiments of the present application, the voice interaction system 1000 may further include more or fewer electronic devices such as mobile phones and earphones than those in the above embodiments, and may also include wearable devices such as watches and bracelets. The present application does not make any limitations here.

[0116] The following introduces an internal form of an intelligent vehicle 100 provided by an embodiment of the present application.

[0117] Figure 3A FIG. shows a schematic diagram of an internal form of an intelligent vehicle 100 provided by an embodiment of the present application.

[0118] As Figure 3A shown, multiple seats may be provided inside the intelligent vehicle 100, and one or more display screens may also be provided. Exemplarily, the interior of the intelligent vehicle 100 may include three rows of seats. The first row of seats may include a driver's seat and a passenger seat. The second row of seats may include a left second-row seat and a right second-row seat. The third row of seats may include a left third-row seat and a right third-row seat. One or more display screens inside the intelligent vehicle 100 may include a central control screen 10, a passenger screen 20, and a right rear projection screen 30. Among them, the central control screen 10 may be disposed between the driver's seat and the passenger seat. The passenger screen 20 may be disposed in front of the passenger seat. The right rear projection screen 30 may be disposed directly in front of the right second-row seat, for example, on the backrest of the passenger seat.

[0119] In some embodiments, each display screen may be equipped with a voice assistant, which is used to interact with users through voice, and is also used to display corresponding voice identifiers based on voice commands input by users. One or more microphones and one or more speakers may be provided inside the intelligent vehicle 100, and the microphones and speakers are used to assist the voice assistant of the display screen to achieve voice interaction with users. Among them, the microphones can be used to collect the voices of users, and the speakers can be used to output voices and also to output the audio specified by users.

[0120] In some embodiments, the microphones and speakers may be provided inside the display screen. For example, one or more microphones and one or more speakers may be provided inside each display screen. The display screen can assist the voice assistant to achieve voice interaction with users through the microphones and speakers inside the display screen.

[0121] In some other embodiments, the positions of the microphones and speakers may be outside the display screen. For example, microphones and speakers are provided beside each seat. In this case, the intelligent vehicle 100 can control one or more microphones inside the vehicle to collect the voices of users inside the vehicle, and after recognizing and processing the collected voices, activate the voice assistant of the corresponding display screen, display the voice identifier through the voice assistant of this display screen. Optionally, it can also call the corresponding speaker to output voices or play the audio specified by users.

[0122] In some other embodiments, the microphones and speakers may also be respectively provided inside and outside the display screen. For example, one or more speakers are provided inside each display screen, and one or more microphones are respectively provided at different seats inside the vehicle, etc. Another example is that one or more microphones and one or more speakers are provided inside each display screen, and one or more microphones are respectively provided at different seats inside the vehicle, etc. In the above cases, the intelligent vehicle 100 can also assist the voice assistant of the display screen to achieve voice interaction with users by controlling the microphones and speakers inside the vehicle. It can be understood that the above-mentioned multiple embodiments are only examples, and the present application does not limit the installation positions of the microphones and speakers.

[0123] Figure 3B Fig. shows a schematic diagram of the internal form of another intelligent vehicle 100 provided by an embodiment of the present application.

[0124] As Figure 3AAs shown, multiple seats can be provided inside the intelligent vehicle 100, and one or more display screens can also be provided. Exemplarily, the interior of the intelligent vehicle 100 can include three rows of seats. The first row of seats can include a driver's seat and a passenger seat. The second row of seats can include a second-row left seat and a second-row right seat. The third row of seats can include a third-row left seat and a third-row right seat. One or more display screens inside the intelligent vehicle 100 can include a central control screen 10, a co-pilot screen 20, and a laser curtain 40. Among them, the central control screen 10 can be set between the driver's seat and the passenger seat. The co-pilot screen 20 can be set in front of the passenger seat. The laser curtain 40 can be set directly in front of the second row of seats, for example, spanning across the backrests of the driver's seat and the passenger seat.

[0125] In some embodiments, each display screen can be equipped with a voice assistant. The voice assistant is used for voice interaction with users and is also used to display corresponding voice identifiers based on voice commands input by users. One or more microphones and one or more speakers can be provided inside the intelligent vehicle 100. The microphones and speakers are used to assist the voice assistants of the display screens to achieve voice interaction with users. Among them, the microphones can be used to collect the voices of users, and the speakers can be used to output voices and also to output the audio specified by users.

[0126] The specific setting methods of the microphones and speakers inside the intelligent vehicle 100 can refer to the relevant descriptions in the above Figure 3A illustrated embodiments and will not be elaborated here.

[0127] It can be understood that Figures 3A to 3B These are just two examples. In the embodiments of the present application, the interior of the intelligent vehicle 100 can also include more, fewer seats or seats with an arrangement different from the above embodiments, or can include more, fewer display screens or display screens with a distribution position different from the above embodiments. The present application does not make any limitations here.

[0128] In the embodiments of the present application, the intelligent vehicle 100 can divide different sound zones according to the seats. Each sound zone can correspond to a display screen, and the display screen can be used to display the voice interaction status between the users in the sound zone and the intelligent vehicle 100. The following introduces the sound zone distribution inside the intelligent vehicle 100.

[0129] Figure 3C Shows a sound zone distribution provided by the embodiments of the present application.

[0130] As Figure 3CAs shown in the figure, the interior of the intelligent vehicle 100 can be divided into multiple sound zones according to the seat distribution. For example, the driver's sound zone, the co-driver's sound zone, the left sound zone in the second row, the right sound zone in the second row, the left sound zone in the third row, the right sound zone in the third row, etc. Among them, each sound zone can include one seat. For example, the driver's sound zone can include the driver's seat, the co-driver's sound zone can include the co-driver's seat, the left sound zone in the second row can include the seat on the left side of the second row, the right sound zone in the second row can include the seat on the right side of the second row, the left sound zone in the third row can include the seat on the left side of the third row, the right sound zone in the third row can include the seat on the right side of the third row, and so on.

[0131] Figure 3D Another sound zone distribution provided by the embodiment of the present application is shown.

[0132] As Figure 3D shown in the figure, the interior of the intelligent vehicle 100 can be divided into multiple sound zones according to the seat distribution, such as the driver's sound zone, the co-driver's sound zone, and the rear row sound zone. Among them, each sound zone can include one or more seats. For example, the driver's sound zone can include the driver's seat, the co-driver's sound zone can include the co-driver's seat, and the rear row sound zone can include the other seats except the driver's seat and the co-driver's seat.

[0133] It can be understood that Figure 3C and Figure 3D are only two examples. In the embodiment of the present application, the intelligent vehicle 100 may also include more, fewer, or different sound zones from the above embodiments. The present application does not make any limitations here.

[0134] In an application scenario, if the internal form of the intelligent vehicle 100 is the same as that in the Figure 3A shown embodiment, and the internal sound zone distribution is the same as that in the Figure 3C shown embodiment, then when the display screens inside the intelligent vehicle 100 are all in the on state, the corresponding relationship between the sound zones inside the intelligent vehicle 100 and the display screens can refer to the following Figure 3E shown embodiment.

[0135] As Figure 3E shown in the figure, when the central control screen 10, the co-driver screen 20, and the right rear projection screen 30 inside the intelligent vehicle 100 are all in the on state, the central control screen 10 can correspond to the following multiple sound zones: the driver's sound zone, the left sound zone in the second row, the left sound zone in the third row, and the right sound zone in the third row. That is, the central control screen 10 can process the voices from the driver's sound zone, the left sound zone in the second row, the left sound zone in the third row, and the right sound zone in the third row; the co-driver screen 20 can correspond to the co-driver's sound zone, that is, the co-driver screen 20 only needs to process the voices from the co-driver's sound zone; the right rear projection screen 30 can correspond to the right sound zone in the second row, that is, the right rear projection screen 30 only needs to process the voices from the right sound zone in the second row.

[0136] It can be understood that Figure 3EThe illustrated embodiments are merely examples. In the embodiments of the present application, the correspondence relationship between the internal display screen and the sound zones in the intelligent vehicle 100 may also be a correspondence relationship different from that of the above embodiments, and the present application does not make any limitation here.

[0137] In an application scenario, if the internal form of the intelligent vehicle 100 is the same as that of the Figure 3A illustrated embodiment, and the internal sound zone distribution is the same as that of the Figure 3C illustrated embodiment, then when the central control screen 10 and the co-pilot screen 20 inside the intelligent vehicle 100 are in the on state, and the right rear projection screen 30 is in the off state, the central control screen 10 can take over the sound zone originally corresponding to the right rear projection screen 30. At this time, the correspondence relationship between the sound zones and the display screens inside the intelligent vehicle 100 can refer to the following Figure 3F illustrated embodiment.

[0138] As Figure 3F shown, when the central control screen 10 and the co-pilot screen 20 inside the intelligent vehicle 100 are both in the on state, and the right rear projection screen 30 is in the off state, the central control screen 10 can correspond to the following multiple sound zones: the driver's sound zone, the left second-row sound zone, the right second-row sound zone, the left third-row sound zone, and the right third-row sound zone. That is, the central control screen 10 can process the voices from the driver's sound zone, the left second-row sound zone, the right second-row sound zone, the left third-row sound zone, and the right third-row sound zone; the co-pilot screen 20 can correspond to the co-pilot sound zone, that is, the co-pilot screen 20 only needs to process the voices from the co-pilot sound zone.

[0139] It can be understood that Figures 3E to 3F this is only an exemplary illustration. When other display screens inside the intelligent vehicle 100 are in the off state, the central control screen 10 can take over the sound zone corresponding to the display screen in the off state and process the voices of that sound zone. In the embodiments of the present application, the display screen in the off state can also be the co-pilot screen 20. At this time, the central control screen 10 can also take over the co-pilot sound zone corresponding to the co-pilot screen 20; in addition, if all display screens except the central control screen 10 are in the off state, or if there is only the central control screen 10 in the intelligent vehicle 100, then the central control screen 10 can process the voices of all sound zones in the entire intelligent vehicle 100, and the present application does not make any limitation here.

[0140] In an application scenario, if the internal form of the intelligent vehicle 100 is the same as that of the Figure 3B illustrated embodiment, and the internal sound zone distribution is the same as that of the Figure 3D illustrated embodiment, then when all the display screens inside the intelligent vehicle 100 are in the on state, the correspondence relationship between the sound zones and the display screens inside the intelligent vehicle 100 can refer to the following Figure 3G illustrated embodiment.

[0141] As Figure 3GAs shown, when the central control screen 10, the co-pilot screen 20, and the laser curtain 40 inside the intelligent vehicle 100 are all in the on state, the central control screen 10 can correspond to the driver's audio zone, that is, the central control screen 10 can process the voice from the driver's audio zone; the co-pilot screen 20 can correspond to the co-pilot's audio zone, that is, the co-pilot screen 20 only needs to process the voice from the co-pilot's audio zone; the laser curtain 40 can correspond to the rear row's audio zone, that is, the laser curtain 40 only needs to process the voice from the rear row's audio zone.

[0142] It can be understood that Figure 3G The illustrated embodiment is only an example. In the embodiments of the present application, the corresponding relationship between the display screen and the audio zone inside the intelligent vehicle 100 can also be a corresponding relationship different from the above embodiment, and the present application does not make a limitation here.

[0143] In an application scenario, if the internal form of the intelligent vehicle 100 is the same as Figure 3B the illustrated embodiment, and the internal audio zone distribution is the same as Figure 3D the illustrated embodiment, then when the central control screen 10 and the co-pilot screen 20 inside the intelligent vehicle 100 are in the on state, and the laser curtain 40 is in the off state, the central control screen 10 can take over the audio zone originally corresponding to the laser curtain 40. At this time, the corresponding relationship between the audio zone and the display screen inside the intelligent vehicle 100 can refer to the following Figure 3H illustrated embodiment.

[0144] As Figure 3H shown, when the central control screen 10 and the co-pilot screen 20 inside the intelligent vehicle 100 are in the on state, and the laser curtain 40 is in the off state, the central control screen 10 can correspond to the driver's audio zone and the rear row's audio zone, that is, the central control screen 10 can process the voice from the driver's audio zone and the rear row's audio zone; the co-pilot screen 20 can correspond to the co-pilot's audio zone, that is, the co-pilot screen 20 only needs to process the voice from the co-pilot's audio zone.

[0145] It can be understood that Figures 3G to 3H This is only an exemplary illustration. When other display screens inside the intelligent vehicle 100 are in the off state, the central control screen 10 can take over the audio zone corresponding to the display screen in the off state and process the voice of this audio zone. In the embodiments of the present application, the display screen in the off state can also be the co-pilot screen 20. At this time, the central control screen 10 can also take over the co-pilot's audio zone corresponding to the co-pilot screen 20; in addition, if all display screens except the central control screen 10 are in the off state, or there is only the central control screen 10 in the intelligent vehicle 100, then the central control screen 10 can process the voice of all audio zones in the entire intelligent vehicle 100, and the present application does not make a limitation here.

[0146] Next, a hardware structure of an intelligent vehicle 100 provided by the embodiments of the present application will be introduced.

[0147] Figure 4AThe figure shows a schematic diagram of the hardware structure of an intelligent vehicle 100 provided by an embodiment of the present application.

[0148] As Figure 4A shown, the intelligent vehicle 100 includes: a controller area network (CAN) bus 11, multiple electronic control units (ECUs), an engine 13, a telematics box (T-box) 14, a transmission 15, a driving recorder 16, an antilock brake system (ABS) 17, a sensor system 18, a camera system 19, a microphone 20, and so on.

[0149] The CAN bus 11 is a serial communication network that supports distributed control or real-time control and is used to connect the various components of the intelligent vehicle 100. Any component on the CAN bus 11 can monitor all the data transmitted on the CAN bus 11. The frames transmitted by the CAN bus 11 can include data frames, remote frames, error frames, and overload frames, and different frames transmit different types of data. In the embodiments of the present application, the CAN bus 11 can be used to transmit the data involved in the control method based on voice instructions by each component, and the specific implementation of this method can refer to the detailed description of the method embodiments hereinafter.

[0150] Not limited to the CAN bus 11, in some other embodiments, the various components of the intelligent vehicle 100 can also be connected and communicate in other ways. For example, the various components can also communicate through in-vehicle Ethernet, a local interconnect network (LIN) bus, FlexRay, and a media-oriented systems (MOST) bus, etc., and the embodiments of the present application do not limit this. The following embodiments will be described with the various components communicating through the CAN bus 11.

[0151] The ECU is equivalent to the processor or brain of the intelligent vehicle 100 and is used to instruct the corresponding component to perform the corresponding action according to the instructions obtained from the CAN bus 11 or according to the operations input by the user. The ECU can be composed of a security chip, a microcontroller unit (MCU), a random access memory (RAM), a read-only memory (ROM), an input / output interface (I / O), an analog / digital converter (A / D converter), and large-scale integrated circuits such as input, output, shaping, and driving.

[0152] There are many types of ECUs, and different types of ECUs can be used to achieve different functions.

[0153] Multiple ECUs in the intelligent vehicle 100 may include, for example: an engine ECU 121, an ECU 122 of a telematics box (T-box), a transmission ECU 123, a driving recorder ECU 124, an antilock brake system (ABS) ECU 125, etc.

[0154] The engine ECU 121 is used to manage the engine and coordinate various functions of the engine. For example, it can be used to start the engine, shut down the engine, etc. The engine is a device that provides power for the intelligent vehicle 100. The engine is a machine that converts one form of energy into mechanical energy. The intelligent vehicle 100 can be used to convert the chemical energy of liquid or gas combustion, or convert electrical energy into mechanical energy and output power externally. The components of the engine can include two major mechanisms: a crankshaft connecting rod mechanism and a valve train mechanism, as well as five major systems: cooling, lubrication, ignition, energy supply, and starting systems. The main components of the engine are a cylinder block, a cylinder head, a piston, a piston pin, a connecting rod, a crankshaft, a flywheel, etc.

[0155] The T-box ECU 122 is used to manage the T-box 14.

[0156] The T-box 14 is mainly responsible for communicating with the Internet, providing a remote communication interface for the intelligent vehicle 100, and providing services including navigation, entertainment, driving data collection, driving trajectory recording, vehicle fault monitoring, vehicle remote query and control (such as unlocking and locking, air conditioning control, window control, engine torque limitation, engine start and stop, seat adjustment, querying battery power, fuel level, door status, etc.), driving behavior analysis, wireless hotspot sharing, road rescue, anomaly reminder, etc.

[0157] The T-box 14 can be used to communicate with a telematics service provider (TSP) and the user (such as the driver)'s side electronic device to achieve vehicle status display and control on the electronic device. When the user sends a control command through the vehicle management application on the electronic device, the TSP will send a request instruction to the T-box 14. After obtaining the control command, the T-box 14 sends a control message through the CAN bus and realizes the control of the intelligent vehicle 100, and finally feeds back the operation result to the vehicle management application on the user side electronic device. That is to say, the data read by the T-box 14 through the CAN bus 11, such as vehicle condition reports, driving reports, fuel consumption statistics, traffic violation inquiries, location trajectories, driving behaviors, etc., can be transmitted to the TSP back-end system through the network and forwarded by the TSP back-end system to the user side electronic device for the user to view.

[0158] T-box 14 may specifically include a communication module and a display screen.

[0159] Among them, the communication module can be used to provide wireless communication functions, supporting the intelligent vehicle 100 to communicate with other devices through wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), ultra-wideband (UWB) and other wireless communication technologies. The communication module can also be used to provide mobile communication functions, supporting the intelligent vehicle 100 to communicate with other devices through communication technologies such as global system for mobile communications (GSM), universal mobile telecommunications system (UMTS), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), 5G and future emerging 6G.

[0160] The communication module can establish a connection and communicate with other devices such as servers and user-side electronic devices through vehicle-to-everything (V2X) communication technology (cellular V2X, C-V2X) based on cellular networks. C-V2X can include, for example, V2X based on long term evolution (LTE) (LTE-V2X), 5G-V2X, etc.

[0161] The display screen is used to provide a visual interface for the driver. One or more display screens may be included in the intelligent vehicle 100. For example, it may include an in-vehicle display screen set in front of the driver's seat, a display screen set above the seat for displaying the surrounding conditions, and may also include a head up display (HUD) that projects information onto the windshield, etc. The display screen for displaying the user interface in the intelligent vehicle 100 provided in the subsequent embodiments may be an in-vehicle display screen set beside the seat, or a display screen set above the seat, or a HUD, etc., which is not limited here. The user interface displayed on the display screen in the intelligent vehicle 100 can be specifically referred to the detailed description of the subsequent embodiments, and will not be elaborated here for the time being.

[0162] T-box 14 can also be referred to as a vehicle infotainment system, a telematics unit, a vehicle gateway, etc., and the embodiments of the present application do not limit this.

[0163] The transmission ECU 123 is used to manage the transmission.

[0164] The transmission 15 is a mechanism used to change the engine speed and torque. It can fix or shift gears to change the transmission ratio between the output shaft and the input shaft. The components of the transmission 15 may include a transmission mechanism, a control mechanism, and a power output mechanism, etc. The main function of the transmission mechanism is to change the numerical value and direction of the torque and speed; the main function of the control mechanism is to control the transmission mechanism to achieve the change of the transmission ratio of the transmission, that is, to achieve gear shifting, so as to achieve speed change and torque change.

[0165] The driving recorder ECU 124 is used to manage the driving recorder 16.

[0166] The components of the driving recorder 16 may include a host, a vehicle speed sensor, data analysis software, etc. The driving recorder 16 refers to an instrument that records the images and sounds during the vehicle's driving, including relevant information such as driving time, speed, and location. In the embodiments of the present application, when the vehicle is driving, the vehicle speed sensor collects the wheel speed and sends the vehicle speed information to the driving recorder 16 through the CAN bus.

[0167] The ABS ECU 125 is used to manage the ABS 17.

[0168] The ABS 17 automatically controls the magnitude of the braking force of the brake when the vehicle brakes, so that the wheels are not locked and are in a state of rolling and sliding, so as to ensure that the adhesion between the wheels and the ground is the maximum value. During the braking process, when the electronic control device determines that there is a wheel tending to lock according to the wheel speed signal input by the wheel speed sensor, the ABS enters the anti-lock braking pressure regulation process.

[0169] The sensor system 18 may include: an acceleration sensor, a vehicle speed sensor, a vibration sensor, a gyroscope sensor, a radar sensor, a signal transmitter, a signal receiver, and so on. The acceleration sensor and the vehicle speed sensor are used to detect the speed of the intelligent vehicle 100. The vibration sensor can be arranged under the seat, on the seat belt, on the backrest, on the operation panel, in the airbag or other positions, and is used to detect whether the intelligent vehicle 100 is collided and the position where the user is located. The gyroscope sensor can be used to determine the motion posture of the intelligent vehicle 100. The radar sensor may include a lidar, an ultrasonic radar, a millimeter wave radar, etc. The radar sensor is used to emit electromagnetic waves to irradiate the target and receive its echo, thereby obtaining information such as the distance from the target to the electromagnetic wave emission point, the rate of change of distance (radial velocity), azimuth, altitude, etc., so as to identify other vehicles, pedestrians or roadblocks near the intelligent vehicle 100. The signal transmitter and the signal receiver are used to transmit and receive signals, and the signal can be used to detect the position where the user is located. The signal can be, for example, ultrasonic wave, millimeter wave, laser, etc.

[0170] The camera system 19 may include a plurality of cameras, and the cameras are used to capture static images or videos. The cameras in the camera system 19 can be arranged in front of the vehicle, behind the vehicle, on the side, inside the vehicle, etc., so as to facilitate functions such as assisted driving, driving record, panoramic surround view, in-vehicle monitoring, etc.

[0171] The sensor system 18 and the camera system 19 can be used to detect the surrounding environment, so as to facilitate the intelligent vehicle 100 to make corresponding decisions to cope with environmental changes. For example, it can be used to complete the task of paying attention to the surrounding environment in the automatic driving stage.

[0172] The microphone 20, also called "microphone" and "transmitter", is used to convert sound signals into electrical signals. When making a call or outputting a voice command, the user can make a sound close to the microphone 20 with the mouth to input the sound signal into the microphone 20. The intelligent vehicle 100 can be provided with at least one microphone 20. In some other embodiments, the intelligent vehicle 100 can be provided with two microphones 20. In addition to collecting sound signals, it can also achieve a noise reduction function. In some other embodiments, the intelligent vehicle 100 can also be provided with three, four or more microphones 20 to form a microphone array, which can achieve functions such as collecting sound signals, noise reduction, identifying the sound source, and realizing the function of directional recording.

[0173] In addition, the intelligent vehicle 100 can also include a plurality of interfaces, such as USB interfaces, RS-232 interfaces, RS485 interfaces, etc., and can be externally connected to cameras, microphones, headphones and user-side electronic devices, such as the mobile phone 200, etc.

[0174] In an embodiment of the present application, the microphone 20 can be used to detect a voice command input by a user. The sensor system 18, the camera system 19, the T-box 14, etc. can be used to obtain the role information of the user who inputs the voice command. The manner in which each component in the intelligent vehicle 100 obtains the role information of the user can refer to the relevant descriptions in the subsequent method embodiments. The T-box ECU 122 can be used to determine whether the current user has the permission corresponding to the voice command according to the role information, and only when the permission is available, the T-box ECU 122 schedules the corresponding components in the intelligent vehicle 100 to respond to the voice command.

[0175] In some embodiments, the sensor system 18, the camera system 19, the T-box 14, etc. are not only used to obtain the role information of the user who inputs the voice command, but also can be used to obtain the role information of other users. The T-box ECU 122 can be used to combine the role information of the user who inputs the voice command and the role information of other users to determine whether the current user has the permission corresponding to the voice command.

[0176] In some embodiments, the sensor system 18, the camera system 19, the T-box 14, etc. can be used to obtain the vehicle state of the intelligent vehicle 100. The T-box ECU 122 can be used to combine the vehicle state and the role information of the user to determine whether the current user has the permission corresponding to the voice command.

[0177] In some embodiments, the memory in the intelligent vehicle 100 can be used to store the binding relationship between the vehicle and the user.

[0178] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the vehicle system. The intelligent vehicle 100 may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0179] For example, the intelligent vehicle 100 may further include a separate memory, a battery, vehicle lights, windshield wipers, an instrument panel, a sound system, a transmission control unit (TCU), an auxiliary control unit (ACU), a passive entry passive start (PEPS) system, an on board unit (OBU), a body control module (BCM), a charging interface, and so on. Among them, the memory can be used to store the permission information for different roles in the intelligent vehicle 100, and this permission information indicates the usage permissions that the role has or does not have for the intelligent vehicle 100. In some embodiments, the memory can be used to store the permission information of different roles in the intelligent vehicle 100 under different vehicle states. In some embodiments, the memory can be used to store the permission information of different roles in the intelligent vehicle 100 when there are different other roles.

[0180] For the specific functions of each component of the intelligent vehicle 100, reference can also be made to the description of the subsequent method embodiments, which will not be elaborated here.

[0181] Next, a hardware structure of a mobile phone 200 provided in an embodiment of the present application will be introduced.

[0182] Figure 4B The figure shows a hardware structure of a mobile phone 200 provided in an embodiment of the present application.

[0183] The mobile phone 200 may include a processor 210, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, antenna 1, antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a sensor module 280, buttons 290, a motor 291, an indicator 292, and a display screen 294, etc. Among them, the sensor module 280 may include a touch sensor 280K. Optionally, the sensor module 280 may further include any one or more of the following: a pressure sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, an ambient light sensor, a bone conduction sensor, etc.

[0184] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on the mobile phone 200. In other embodiments of the present application, the mobile phone 200 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0185] The processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0186] The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.

[0187] A memory may also be provided in the processor 210 for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory may save the instructions or data that the processor 210 has just used or recycled. If the processor 210 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.

[0188] In some embodiments, the processor 210 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0189] The USB interface 230 is an interface that complies with the USB standard specification. Specifically, it can be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 230 can be used to connect a charger to charge the mobile phone 200, and can also be used to transfer data between the mobile phone 200 and peripheral devices. It can also be used to connect headphones to play audio. This interface can also be used to connect other electronic devices, such as AR devices, etc.

[0190] It can be understood that the interface connection relationship between the modules schematically shown in the embodiments of the present invention is only a schematic illustration and does not constitute a structural limitation on the mobile phone 200. In other embodiments of the present application, the mobile phone 200 can also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.

[0191] The charging management module 240 is used to receive a charging input from a charger. Among them, the charger can be a wireless charger or a wired charger. In some embodiments of wired charging, the charging management module 240 can receive the charging input of the wired charger through the USB interface 230. In some embodiments of wireless charging, the charging management module 240 can receive the wireless charging input through the wireless charging coil of the mobile phone 200. While charging the battery 242, the charging management module 240 can also supply power to the electronic device through the power management module 241.

[0192] The power management module 241 is used to connect the battery 242, the charging management module 240 and the processor 210. The power management module 241 receives the input from the battery 242 and / or the charging management module 240 and supplies power to the processor 210, the internal memory 221, the display screen 294, the wireless communication module 260, etc. The power management module 241 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some other embodiments, the power management module 241 can also be provided in the processor 210. In some other embodiments, the power management module 241 and the charging management module 240 can also be provided in the same device.

[0193] The wireless communication function of the mobile phone 200 can be implemented through antenna 1, antenna 2, the mobile communication module 250, the wireless communication module 260, the modulation and demodulation processor, and the baseband processor, etc.

[0194] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the mobile phone 200 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example: Antenna 1 can be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.

[0195] The mobile communication module 250 can provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the mobile phone 200. The mobile communication module 250 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 250 can receive electromagnetic waves through the antenna 1, filter, amplify, and perform other processing on the received electromagnetic waves, and then transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 250 can also amplify the signal modulated by the modulation and demodulation processor and convert it into electromagnetic waves through the antenna 1 for radiation. In some embodiments, at least some functional modules of the mobile communication module 250 may be provided in the processor 210. In some embodiments, at least some functional modules of the mobile communication module 250 and at least some modules of the processor 210 may be provided in the same device.

[0196] The modulation and demodulation processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 270A, receiver 270B, etc.), or displays an image or video through the display screen 294. In some embodiments, the modulation and demodulation processor may be an independent device. In some other embodiments, the modulation and demodulation processor may be independent of the processor 210 and be provided in the same device as the mobile communication module 250 or other functional modules.

[0197] The wireless communication module 260 can provide solutions for wireless communications applied to the mobile phone 200, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSSs), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. The wireless communication module 260 can be one or more devices integrating at least one communication processing module. The wireless communication module 260 receives electromagnetic waves via the antenna 2, demodulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 210. The wireless communication module 260 can also receive the signals to be sent from the processor 210, perform frequency modulation and amplification on them, and convert them into electromagnetic waves through the antenna 2 for radiation.

[0198] In some embodiments, the antenna 1 of the mobile phone 200 is coupled to the mobile communication module 250, and the antenna 2 is coupled to the wireless communication module 260, so that the mobile phone 200 can communicate with the network and other devices through wireless communication technologies. The wireless communication technologies may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include Global Positioning System (GPS), Global Navigation Satellite System (GLONASS), Beidou Navigation Satellite System (BDS), Quasi-Zenith Satellite System (QZSS), and / or Satellite Based Augmentation Systems (SBAS).

[0199] The mobile phone 200 implements the display function through the GPU, the display screen 294, and the application processor, etc. The GPU is a microprocessor for image processing, and is connected to the display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 210 may include one or more GPUs, which execute program instructions to generate or change display information.

[0200] The display screen 294 is used to display images, videos, etc. The display screen 294 includes a display panel. The display panel can be a liquid crystal display (LCD), or it can also be an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the mobile phone 200 may include one or N display screens 294, where N is a positive integer greater than 1.

[0201] The internal memory 221 may include one or more random access memories (RAM) and one or more non-volatile memories (NVM). The random access memory can be directly read and written by the processor 210, and can be used to store the operating system or executable programs of other running programs (such as machine instructions), and can also be used to store data of users and application programs, etc. The non-volatile memory can also store executable programs and store data of users and application programs, etc., and can be pre-loaded into the random access memory for the processor 210 to directly read and write.

[0202] The audio module 270 may include a speaker 270A, a receiver 270B, and a microphone 270C. The mobile phone 200 can implement audio functions through the audio module 270 and the application processor, etc. For example, music playback, recording, etc.

[0203] The audio module 270 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 270 can also be used to encode and decode audio signals. In some embodiments, the audio module 270 can be disposed in the processor 210, or some functional modules of the audio module 270 can be disposed in the processor 210.

[0204] The speaker 270A, also known as the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The mobile phone 200 can listen to music or hands-free calls through the speaker 270A.

[0205] The receiver 270B, also known as the "earpiece", is used to convert audio electrical signals into sound signals. When the mobile phone 200 answers a call or a voice message, the voice can be received by bringing the receiver 270B close to the human ear.

[0206] The microphone 270C, also known as the "microphone" or "transmitter", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak close to the microphone 270C, and the sound signal is input into the microphone 270C. The mobile phone 200 can be provided with at least one microphone 270C. In some other embodiments, the mobile phone 200 can be provided with two microphones 270C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the mobile phone 200 can also be provided with three, four or more microphones 270C, which can collect sound signals, reduce noise, identify the sound source, and implement functions such as directional recording.

[0207] The touch sensor 280K, also known as the "touch device". The touch sensor 280K can be disposed on the display screen 294, and the touch sensor 280K and the display screen 294 together form a touch screen, also known as the "touch panel". The touch sensor 280K is used to detect touch operations acting thereon or nearby. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual outputs related to the touch operation can be provided through the display screen 294. In some other embodiments, the touch sensor 280K can also be disposed on the surface of the mobile phone 200, at a different position from the display screen 294.

[0208] The keys 290 include a power-on key, volume keys, etc. The keys 290 can be mechanical keys or touch keys. The mobile phone 200 can receive key inputs and generate key signal inputs related to the user settings and function controls of the mobile phone 200.

[0209] The motor 291 can generate vibration prompts. The motor 291 can be used for incoming call vibration prompts or touch vibration feedback. For example, touch operations on different applications (such as taking pictures, audio playing, etc.) can correspond to different vibration feedback effects. Touch operations on different regions of the display screen 294 can also correspond to different vibration feedback effects for the motor 291. Different application scenarios (such as time reminders, receiving messages, alarms, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effects can also support customization.

[0210] The indicator 292 can be an indicator light, which can be used to indicate the charging status, power change, and can also be used to indicate messages, missed calls, notifications, etc.

[0211] It should be noted that the hardware structure of the earphone 300 can also refer to the above Figure 4BThe hardware structure of the mobile phone 200 shown. Moreover, the headset 300 may include fewer, more, or different components than the mobile phone 200, and this application does not make any limitations here.

[0212] A voice interaction method provided by an embodiment of this application. The intelligent vehicle 100 may include one or more display screens, such as display screen A and display screen B, and the intelligent vehicle 100 may carry multiple users (such as user A, user B, etc.). User A is located in sound zone a, and user B is located in sound zone b. During the process of user A interacting with display screen A by voice, the intelligent vehicle 100 may receive and respond to voice 1 issued by user B. When it is determined that voice 1 belongs to a hands-free wake-up instruction or voice 1 includes a wake-up word, the intelligent vehicle 100 may control the display screen B corresponding to sound zone b to interact with user B by voice. Among them, display screen A may be the same as or different from display screen B.

[0213] In this way, the intelligent vehicle 100 can interact with multiple users by voice simultaneously. Moreover, in a multi-user voice interaction scenario, users can see in real time whether the voice command is recognized without waiting, bringing a better user experience.

[0214] In some application scenarios, the intelligent vehicle 100 carries user A and user B. User A is located within sound zone a, and user B is located within sound zone b. When interface 1 is displayed on display screen A of the intelligent vehicle 100, the intelligent vehicle 100 receives the wake-up word issued by user A. The intelligent vehicle 100 may respond to the wake-up word issued by user A from sound zone a. After determining that the display screen corresponding to sound zone a is display screen A, it may determine display screen A as the main display screen and display a voice dialogue identifier 1 on interface 1 of display screen A. The voice dialogue identifier 1 is used to prompt the user that the voice assistant of display screen A has been activated through the wake-up word. When the voice dialogue identifier 1 is displayed on display screen A, the intelligent vehicle 100 may receive and respond to the hands-free wake-up instruction issued by the user from sound zone b. When it is determined that the display screen corresponding to sound zone b is display screen A, the display screen A of the intelligent vehicle 100 may split out a voice sub-identifier from the voice dialogue identifier 1 on interface 1. The voice sub-identifier is used to prompt the user that display screen A is processing multiple voices.

[0215] In this way, during the process of user A interacting with display screen A through the voice assistant, user B can also interact with display screen A through the hands-free wake-up instruction. The intelligent vehicle 100 can interact with multiple users by voice simultaneously, providing more convenient services for users.

[0216] Exemplarily, such as Figure 5AAs shown, the central control screen 10 can display a main interface 500. One or more application icons can be displayed in the main interface 500. For example, a music application icon, a navigation application icon, a settings application icon, etc.

[0217] The intelligent vehicle 100 can receive and respond to the wake-up word issued by user A from sound zone a. If the intelligent vehicle 100 determines that the display screen corresponding to sound zone a is the central control screen 10, the intelligent vehicle 100 can determine the central control screen 10 as the main display screen and display the Figure 5B voice dialogue identifier 501 as shown.

[0218] As Figure 5B shown, the main interface 500 includes a voice dialogue identifier 501. The voice dialogue identifier 501 can include a rounded rectangle frame 502 and a dialogue symbol 503. Optionally, it can also include text 504. The voice dialogue identifier 1 is used to prompt the user that the voice assistant of the central control screen 10 has been awakened by the wake-up word. In the voice dialogue identifier 501, the dialogue symbol 503 and the text 504 can be displayed in the rounded rectangle frame 502. The dialogue symbol 503 can be a symbol composed of multiple vertical lines of different lengths, and the length of each vertical line in the multiple vertical lines can change with time, used to simulate the change of sound volume. The dialogue symbol 503 can be used to prompt the user that the central control screen 10 is performing a voice interaction with the user. The text 504 can be used to prompt the user to start issuing a voice command. Exemplarily, the text 504 can be "You can talk at any time". In the embodiments of the present application, Figure 5B The rounded rectangle-shaped voice dialogue identifier 501 shown can also be called a voice capsule. It can be understood that the voice dialogue identifier 501 can also adopt different shapes such as a rectangle, a circle, a heart shape, etc. as in the above embodiments. The dialogue symbol 503 can also adopt a symbol different from that in the above embodiments, and the text inside the voice dialogue identifier 501 can also be text different from that in the above embodiments. The present application does not make any limitations here. Further optionally, a sound zone indicator 505 can also be displayed in the main interface 500. The sound zone indicator 505 can be used to indicate the sound zone where the user (i.e., user A) who is performing a voice interaction with the central control screen 10 is located. Figure 5B The display area of the sound zone indicator 505 shown includes the lower left area of the voice dialogue identifier 501. The sound zone indicator 505 can be used to indicate that the sound zone a where user A is located is the sound zone on the left side of the central control screen 10 (such as the driver's sound zone). It can be understood that the sound zone indicator 505 here is only an example. In the embodiments of the present application, the sound zone indicator can also adopt a shape, color or symbol different from that of the above sound zone indicator 505. The present application does not make any limitations here.

[0219] After the voice assistant is turned on on the central control screen 10, the intelligent vehicle 100 can receive and respond to a voice command issued by user A from sound zone a (such as "The weather is nice today, open the window"). As Figure 5C shown, the central control screen 10 can replace the text 504 displayed in the voice conversation identifier 501 with the text 506.

[0220] As Figure 5C shown, the voice conversation identifier 501 can display the text 506, and the content of the text 506 can be the same or similar to the voice command issued by the user. Exemplarily, the text 506 can include "The weather is nice today, open the window". It should be noted that due to the limited size of the voice conversation identifier 501, if the number of characters included in the text 506 is too large and exceeds the display range of the voice conversation identifier 501, the central control screen 10 can adopt a scrolling display method to scroll and display the text 506 in the voice conversation identifier 501. The text 506 can be used to prompt the user that the central control screen 10 is receiving and processing the voice command issued by the user.

[0221] While the intelligent vehicle 100 is receiving the voice command issued by user A (such as "The weather is nice today, open the window"), the intelligent vehicle 100 can also receive a wake-up-free command issued by user B from sound zone b. In response to the wake-up-free command issued by user B from sound zone b, when it is determined that the display screen corresponding to sound zone b is the central control screen 10, the central control screen 10 of the intelligent vehicle 100 can split the voice conversation identifier 501 and display the split identifier 510 as Figure 5D shown.

[0222] As Figure 5D shown, the split identifier 510 is displayed in the main interface 500. The split identifier 510 can include two voice identifiers that are not completely split. One of the voice identifiers is the voice conversation identifier 520, and the other is the voice sub-identifier 530. There can also be a connection channel between the voice conversation identifier 520 and the voice sub-identifier 530, indicating that the two voice identifiers are still in the process of splitting and have not been completely split. The split identifier 510 can be used to indicate that the central control screen 10 is processing the voices of multiple users simultaneously. Among them, the specific content of the voice conversation identifier 520 can refer to the above Figure 5CThe related description of the voice dialogue identifier 501 in the illustrated embodiment will not be elaborated here. The voice sub-identifier 530 may include a rounded rectangle box and a sub-symbol 531, and the sub-symbol 531 may be used to prompt the user that the voice sub-identifier 530 where the sub-symbol 531 is located is a split voice identifier. In some embodiments, the sub-symbol 531 may be composed of multiple points of the same size (or line segments of the same length). In other embodiments, the sub-symbol 531 may also adopt other symbols, which are not limited in this application. Optionally, a voice area indicator 511 may also be displayed on the main interface 500, and the display area of the voice area indicator 511 is the area below the split identifier 510 (including the lower left area, the directly below area, and the lower right area). The voice area indicator 511 is used to indicate that the voice areas of multiple channels being processed by the central control screen 10 are located in different orientations of the central control screen 10, that is, the voice area a where user A is located is the voice area on the left side of the central control screen 10, and the voice area b where user B is located is the voice area on the right side of the central control screen 10 (such as the right voice area in the third row).

[0223] After the split identifier 510 is completely split, as Figure 5E shown, a completely split voice dialogue identifier 520 and a voice sub-identifier 530 may be displayed on the main interface 500 of the central control screen 10, that is, there is no connection channel between the voice dialogue identifier 520 and the voice sub-identifier 530.

[0224] After the intelligent vehicle 100 executes the hands-free instruction issued by user B, as Figure 5F shown, the central control screen 10 may stop displaying the voice sub-identifier 530. Optionally, a voice area indicator 521 may also be displayed on the main interface 500, and the voice area indicator 521 may be the same as the voice area indicator 505 shown above Figure 5B shown.

[0225] It should be noted that after the central control screen 10 executes the voice instruction issued by user A after the wake-up word, when it is detected that the closing condition is met (for example, the duration of user A not issuing a voice is greater than the preset duration, etc.), the central control screen 10 may close the voice dialogue identifier 520 (and the voice area indicator 521) displayed on the main interface 500, that is, the central control screen 10 may redisplay the main interface 500 shown above Figure 5A shown.

[0226] In some embodiments, while executing the voice instruction issued by the user (or after executing the voice instruction), the intelligent vehicle 100 may also output feedback information in one or more ways such as voice broadcast, indicator light flashing, vibration, display screen display, etc. The feedback information is used to prompt the user that the intelligent vehicle 100 has executed the operation corresponding to the voice instruction. The specific content of the feedback information may refer to the related description in the embodiment shown below Figure 6D shown, which will not be elaborated here for the time being.

[0227] It is understandable that the above Figures 5A to 5F illustrated embodiments are merely examples. In the embodiments of the present application, the user can also wake up the voice assistants of other display screens (such as the co-pilot screen 20, the right rear projection screen 30, the laser curtain 40, etc.) through wake-up words. Moreover, the interface 1 displayed on the display screen A can also be other interfaces other than the main interface. In addition, the voice dialogue identifier and the voice sub-identifier can also adopt shapes and / or symbols different from those of the above embodiments, or include contents different from those of the above embodiments. In addition, the voice command issued by the user A can also be a voice command different from that of the above embodiments. The present application does not make any limitations here.

[0228] In some embodiments, the same display screen (such as the central control screen 10, the laser curtain 40, etc.) in the intelligent vehicle 100 can also receive and process the wake-up-free commands issued by two or more other users while processing non-wake-up commands (such as the voice command issued by the user A after saying the wake-up word). In this case, the display screen can also display a voice dialogue identifier and a voice sub-identifier. The voice dialogue identifier is used to prompt the user about the voice interaction situation between the display screen and the user A, and the voice sub-identifier is used to prompt that the display screen is also processing one or more other wake-up-free commands at the same time. The specific interface can also refer to the relevant descriptions in the above Figure 5E illustrated embodiments, which will not be elaborated here.

[0229] In some application scenarios, the intelligent vehicle 100 carries the user A and the user B. The user A is located in the sound zone a, and the user B is located in the sound zone b. When the interface 1 is displayed on the display screen A of the intelligent vehicle 100 and the interface 2 is displayed on the display screen B, the intelligent vehicle 100 receives the wake-up word issued by the user A. The intelligent vehicle 100 can respond to the wake-up word issued by the user A from the sound zone a. After determining that the display screen corresponding to the sound zone a is the display screen A, the display screen A can be determined as the main display screen, and the voice dialogue identifier 1 is displayed on the interface 1 of the display screen A. The voice dialogue identifier 1 is used to prompt the user that the voice assistant of the display screen A has been activated through the wake-up word. When the voice dialogue identifier 1 is displayed on the display screen A, the intelligent vehicle 100 can receive and respond to the wake-up-free command issued by the user from the sound zone b. After determining that the display screen corresponding to the sound zone b is the display screen B, the wake-up-free identifier is displayed on the interface 2 of the display screen B. The wake-up-free identifier is used to prompt the user that the display screen B is receiving and processing the wake-up-free command of the user.

[0230] In this way, during the process of the user A having a voice interaction with the display screen A through the voice assistant, the user B can also have a voice interaction with the display screen B through the wake-up-free command. The intelligent vehicle 100 can conduct voice interactions with multiple users simultaneously, providing more convenient services for the users.

[0231] Exemplarily, when the intelligent vehicle 100 receives the wake-up word sent by user A from sound zone a and determines that the display screen corresponding to sound zone a is the central control screen 10, the central control screen 10 of the intelligent vehicle 100 can display as follows Figure 6A The main interface 500 shown, and the co-pilot screen 20 can display as follows Figure 6B The main interface 600 shown.

[0232] As Figure 6A shown, the central control screen 10 of the intelligent vehicle 100 can display the main interface 500, and the main interface 500 can display a voice dialogue identifier 501, which is used to prompt the user that the voice assistant of the central control screen 10 has been activated through the wake-up word. Among them, the specific content of the main interface 500 and the voice dialogue identifier 501 can refer to the relevant descriptions in the above Figure 5C shown embodiment. In addition, the specific process of the central control screen 10 displaying the voice dialogue identifier 501 in response to the wake-up word sent by user A can also refer to the relevant content in the above Figures 5A to 5C shown embodiment, which will not be elaborated here. Optionally, a sound zone indicator 505 can also be displayed in the main interface 500, and the relevant content of the sound zone indicator 505 can refer to the relevant content in the above Figure 5B shown embodiment, which will not be elaborated here.

[0233] As Figure 6B shown, the co-pilot screen 20 of the intelligent vehicle 100 displays the main interface 600, and the main interface 600 can display a wallpaper, such as a landscape picture.

[0234] While the central control screen 10 is performing voice interaction with user A, the intelligent vehicle 100 can also receive a hands-free wake-up instruction sent by user B from sound zone b. The intelligent vehicle 100 can respond to the hands-free wake-up instruction sent by user B from sound zone b. When it is determined that the display screen corresponding to sound zone b is the co-pilot screen 20, as Figure 6C shown, the co-pilot screen 20 of the intelligent vehicle 100 can display a hands-free wake-up identifier 601 on the main interface 600.

[0235] As Figure 6C shown, the hands-free wake-up identifier 601 can include a rounded rectangle frame 602 and a hands-free wake-up symbol 603. Optionally, text 604 can also be displayed in the hands-free wake-up identifier 601. The hands-free wake-up identifier 601 can be used to prompt the user that the display screen B is receiving and processing the user's hands-free wake-up instruction. In the hands-free wake-up identifier 601, the hands-free wake-up symbol 603 and the text 604 can be located inside the rounded rectangle frame 602. The hands-free wake-up symbol 603 can be Figure 6C the spherical symbol shown, or other symbols of different shapes and colors. This application does not make any limitations here. Moreover, the hands-free wake-up symbol 603 and the above Figure 5Bis different from the dialogue symbol 503 shown. The text 604 can be the text content of the voice-free instruction issued by user B, such as "turn on the air conditioner". In the embodiments of the present application, Figure 6C The voice-free identifier 601 shown can also be referred to as a voice-free capsule.

[0236] It should be noted that since the sound zone corresponding to the co-pilot screen 20 is only the co-pilot sound zone, that is, the co-pilot screen 20 only needs to process the voice from the co-pilot seat, therefore, in some embodiments, the co-pilot screen 20 may not display a sound zone indicator.

[0237] In some embodiments, after the intelligent vehicle 100 recognizes the voice-free instruction (such as "turn on the air conditioner") issued by user B, the intelligent vehicle 100 can, while (or after) executing the operation corresponding to the voice-free instruction, such as Figure 6D shown, change the text 604 in the voice-free identifier 601 to the text 605.

[0238] Such as Figure 6D shown, on the main interface 600, there is a voice-free identifier 601 displayed, and the text 605 is displayed in the voice-free identifier 601. The text 605 can be used to prompt the user that the voice-free instruction responded by the co-pilot screen 20 has been executed. For example, the text 605 can be "The air conditioner has been turned on for you".

[0239] In some embodiments, while the co-pilot screen 20 displays the text 605 in the voice-free identifier 601, the co-pilot screen 20 can also play a voice, which is used to prompt the user that the voice-free instruction responded by the co-pilot screen 20 has been executed. For example, the content of the played voice can be the content of the above Figure 6D shown text 605.

[0240] It can be understood that Figure 6D the embodiments shown are only exemplary illustrations that the feedback information can be output in the form of display on the display screen. In the embodiments of the present application, the intelligent vehicle 100 can also output the feedback information in one or more ways such as voice broadcast, indicator light flashing (such as the indicator light on the display screen flashing, the door panel atmosphere light flashing, etc.), display screen display, vibration, etc. In another possible implementation manner, executing the operation corresponding to the user's voice instruction can also be regarded as a way of outputting feedback information, and the present application does not make a limitation here.

[0241] After the intelligent vehicle 100 executes the voice-free instruction of user B, or after the intelligent vehicle 100 outputs the feedback information (such as playing audio, outputting the above text 605), the intelligent vehicle 100 can control the co-pilot screen 20 to stop displaying the voice-free identifier 601 and redisplay the above Figure 6B shown main interface 600.

[0242] It should be noted that during the process of the co-pilot screen 20 receiving and processing the wake-up-free instruction of user B, the central control screen 10 can perform voice interaction with user A through the voice assistant. Before the central control screen 10 meets the closing condition, the central control screen 10 can always display a voice dialogue identifier 501 in the main interface 500. Whether the co-pilot screen 20 responds to the wake-up-free instruction will not affect the display content of the central control screen 10.

[0243] It can be understood that the above Figures 6A to 6D illustrated embodiments are just examples. In the embodiments of the present application, the user can also wake up the voice assistants of other display screens (such as the co-pilot screen 20, the right rear projection screen 30, the laser curtain 40, etc.) through wake-up words. The display screen B can also be other display screens other than the co-pilot screen 20 (such as the right rear projection screen 30, the laser curtain 40, etc.). In addition, the interface 1 displayed on the display screen A can also be other interfaces other than the main interface 500, and the display screen 2 displayed on the display screen B can also be other interfaces other than the main interface 600. The voice dialogue identifier and the wake-up-free identifier can also adopt shapes, symbols, etc. different from those in the above embodiments. The voice dialogue identifier and the wake-up-free identifier can also include contents different from those in the above embodiments. The present application does not make any limitations here. In addition, the voice instructions issued by user A and the wake-up-free instructions issued by user B can also be voice instructions different from those in the above embodiments. The present application does not make any limitations here.

[0244] In some application scenarios, the intelligent vehicle 100 carries user A and user B. User A is located in sound zone a, and user B is located in sound zone b. When the interface 1 is displayed on the display screen A of the intelligent vehicle 100 and the interface 2 is displayed on the display screen B, the intelligent vehicle 100 receives a wake-up-free instruction issued by user A from sound zone a. In response to the wake-up-free instruction issued by user A from sound zone a, the intelligent vehicle 100 can display a wake-up-free identifier 1 on the interface 1 of the display screen A. The wake-up-free identifier 1 is used to prompt the user that the display screen A is receiving and processing the user's wake-up-free instruction. When the wake-up-free identifier 1 is displayed on the display screen A, the intelligent vehicle 100 can receive and respond to the wake-up-free instruction issued by the user from sound zone b. When it is determined that the display screen corresponding to sound zone b is the display screen B, the intelligent vehicle 100 can display a wake-up-free identifier 2 on the interface 2 of the display screen B. The wake-up-free identifier 2 is used to prompt the user that the display screen B is receiving and processing the user's wake-up-free instruction.

[0245] In this way, the intelligent vehicle 100 can simultaneously receive and process the wake-up-free instructions of different users, perform voice interaction with multiple users through different display screens, and provide more convenient services for the users.

[0246] Exemplarily, the central control screen 10 of the intelligent vehicle 100 can display a main interface 700 as Figure 7A shown, and the right rear projection screen 30 can display asFigure 7B The main interface 710 shown.

[0247] As Figure 7A shown, the central control screen 10 can display the main interface 700. One or more application icons can be displayed in the main interface 700. For example, a music application icon, a navigation application icon, a settings application icon, etc.

[0248] As Figure 7B shown, the right rear projection screen 30 of the intelligent vehicle 100 displays the main interface 710, and a wallpaper such as a landscape picture can be displayed in the main interface 710.

[0249] The intelligent vehicle 100 can receive and respond to the hands-free wake-up instruction issued by user A from sound zone a, such as "open the window". If the intelligent vehicle 100 determines that the display screen corresponding to sound zone a is the central control screen 10, the intelligent vehicle 100 can display the hands-free wake-up identifier 701 as Figure 7C shown in the main interface 700 of the central control screen 10.

[0250] As Figure 7C shown, the main interface 700 includes a hands-free wake-up identifier 701. The hands-free wake-up identifier 701 can include a rounded rectangle frame 702 and a hands-free wake-up symbol 703. Optionally, it can also include text 704. The hands-free wake-up identifier 701 is used to prompt the user that the central control screen 10 is receiving and processing the user's hands-free wake-up instruction. In the hands-free wake-up identifier 701, the hands-free wake-up symbol 703 and the text 704 can be displayed in the rounded rectangle frame 702. The specific content of the hands-free wake-up symbol 703 can refer to the relevant description of the hands-free wake-up identifier 601 shown above Figure 6C and will not be elaborated here. The text 704 can include the text content of the hands-free wake-up instruction issued by user A, such as "open the window". Further optionally, a sound zone indicator 705 can also be displayed in the main interface 700. The sound zone indicator 705 can be used to indicate the sound zone where the user (i.e., user A) who is having a voice interaction with the central control screen 10 is located. Figure 7C The display area of the sound zone indicator 705 shown includes the lower left area of the hands-free wake-up identifier 701. The sound zone indicator 705 can be used to indicate that the sound zone a where user A is located is the sound zone on the left side of the central control screen 10 (such as the driver's sound zone).

[0251] During the process of the central control screen 10 processing the hands-free wake-up instruction of user A, the intelligent vehicle 100 can also receive the hands-free wake-up instruction issued by user B from sound zone b, such as "turn on the air conditioner". In response to the hands-free wake-up instruction issued by user B from sound zone b, when it is determined that the display screen corresponding to sound zone b is the right rear projection screen 30, as Figure 7D shown, the intelligent vehicle 100 can display the hands-free wake-up identifier 711 on the main interface 710 of the right rear projection screen 30.

[0252] AsFigure 7D As shown, the main interface 710 displays a wake-up-free mark 711, which may include a rounded rectangular frame 712 and a wake-up-free symbol 713, and optionally, may also include text 714. The specific content of the wake-up-free mark 711 may refer to the above Figure 6C The relevant description of the wake-up-free mark 601 is not repeated here. The wake-up-free mark 711 can be used to indicate that the right rear projection screen 30 is receiving and processing the user's wake-up-free instruction.

[0253] It should be noted that since the sound zone corresponding to the right rear projection screen 30 only has the second row right sound zone, that is, the right rear projection screen 30 only needs to process the voice from the second row right seats, therefore, in some embodiments, the right rear projection screen 30 may not display the sound zone indicator.

[0254] In some embodiments, after executing the operation corresponding to the wake-up-free instruction, the smart car 100 can control the central control screen 10 and the right rear projection screen 30 to output feedback information. The specific content of the feedback information can refer to the above Figure 6D The relevant description in the illustrated embodiment will not be repeated here.

[0255] In addition, after executing the operation corresponding to the wake-up-free instruction, the smart car 100 can also control the display screen that responds to the wake-up-free instruction (such as the central control screen 10, the right rear projection screen 30, etc.) to stop displaying the wake-up-free logo, which is not limited in this application.

[0256] It is understandable that the above Figures 7A to 7D The embodiment shown is only an example. In the embodiment of the present application, the user can also perform voice interaction with other display screens (such as the co-pilot screen 20, the laser curtain 40, etc.) through the wake-up-free command, and the interface 1 displayed on the display screen A can also be other interfaces other than the main interface 700, and the interface 2 displayed on the display screen B can also be other interfaces other than the main interface 710. The present application does not limit this. In addition, the wake-up-free logo can also adopt a shape and / or symbol different from that of the above embodiment, or include content different from that of the above embodiment, and the wake-up-free command issued by user A and user B can also be a wake-up-free command different from that of the above embodiment, which is not limited in the present application.

[0257] In some application scenarios, the intelligent vehicle 100 carries user A and user B. User A is located within sound zone a, and user B is located within sound zone b. When interface 1 is displayed on display screen A of the intelligent vehicle 100, the intelligent vehicle 100 receives a hands-free wake-up instruction issued by user A from sound zone a. In response to the hands-free wake-up instruction issued by user A from sound zone a, the intelligent vehicle 100 can display a hands-free wake-up identifier 1 on interface 1 of display screen A. The hands-free wake-up identifier 1 is used to prompt the user that display screen A is receiving and processing the user's hands-free wake-up instruction. When the hands-free wake-up identifier 1 is displayed on display screen A, the intelligent vehicle 100 can receive and respond to a hands-free wake-up instruction issued by the user from sound zone b. When it is determined that the display screen corresponding to sound zone b is display screen A, the intelligent vehicle 100 can display a composite hands-free wake-up identifier on interface 1 of display screen A. The composite hands-free wake-up identifier is used to prompt the user that display screen A is receiving and processing hands-free wake-up instructions issued by multiple users.

[0258] In this way, the intelligent vehicle 100 can simultaneously receive and process hands-free wake-up instructions from different users, conduct voice interactions with multiple users through the same display screen, and provide more convenient services to the users.

[0259] Exemplarily, the intelligent vehicle 100 can receive and respond to a hands-free wake-up instruction issued by user A from sound zone a. When it is determined that the display screen corresponding to sound zone a is the central control screen 10, as Figure 8A shown, a hands-free wake-up identifier 801 can be displayed in the main interface 800 of the central control screen 10. The hands-free wake-up identifier 801 is used to prompt the user that the central control screen 10 is receiving and processing the hands-free wake-up instruction issued by the user. The specific contents of the main interface 800 and the hands-free wake-up identifier 801 can refer to the relevant descriptions in the above Figure 7C shown embodiments. In addition, the specific process of the central control screen 10 receiving and responding to the hands-free wake-up instruction issued by user A and displaying the hands-free wake-up identifier 801 can also refer to the relevant descriptions in the above Figure 7A and Figure 7C shown embodiments, which will not be elaborated here. Optionally, a sound zone indicator 802 can also be displayed in the main interface 800. The sound zone indicator 802 can be used to indicate the sound zone where the user (i.e., user A) who is currently conducting a voice interaction with the central control screen 10 is located. Figure 8A As shown, the display area of the sound zone indicator 802 includes the lower left area of the composite hands-free wake-up identifier 810. The sound zone indicator 802 can be used to indicate that the sound zone a where user A is located is the sound zone on the left side of the central control screen 10 (such as the second row left sound zone, etc.).

[0260] During the process of the central control screen 10 processing the hands-free wake-up instruction of user A, the intelligent vehicle 100 can also receive a hands-free wake-up instruction issued by user B from sound zone b, such as "turn on the air conditioner". In response to the hands-free wake-up instruction issued by user B from sound zone b, when it is determined that the display screen corresponding to sound zone b is the central control screen 10, asFigure 8B As shown, the intelligent vehicle 100 can display a composite wake-up-free identifier 810 on the main interface 800 of the center control screen 10.

[0261] As Figure 8B shown, a composite wake-up-free identifier 810 is displayed in the main interface 800. The composite wake-up-free identifier 810 is used to prompt the user that the center control screen 10 is processing wake-up-free instructions of multiple users. The composite wake-up-free identifier 810 may include a rounded rectangle frame 811 and a composite wake-up-free symbol 812. Optionally, it may further include text 813 and / or a digital indicator 814. Among them, the composite wake-up-free symbol 812, the text 813, and the digital indicator 814 are all displayed inside the rounded rectangle frame 811. The composite wake-up-free symbol 812 can be used to prompt the user that the center control screen 10 is processing wake-up-free instructions of multiple users. The composite wake-up-free symbol 812 can be two overlapping wake-up-free composites, or symbols of other shapes and colors. This application does not make a limitation here. The text 813 can be used to prompt the user that the center control screen 10 is processing multiple wake-up-free instructions. For example, the text 813 can be "Multiple executions in progress". In some other embodiments, the text 813 can also rotate and display multiple wake-up-free instructions being processed. The digital indicator 814 can be used to prompt the number of wake-up-free instructions being processed by the center control screen 10. For example, the digital indicator 814 can be "2", used to prompt the user that the current center control screen 10 is simultaneously processing two wake-up-free instructions. Further optionally, a sound zone indicator 815 can also be displayed in the main interface 800. The sound zone indicator 815 can be used to indicate the sound zones where the users (i.e., user A and user B) who are having a voice interaction with the center control screen 10 are located. Figure 8B The display area of the sound zone indicator 815 shown includes the lower left area of the composite wake-up-free identifier 810. The sound zone indicator 815 can be used to indicate that the sound zone a where user A is located and the sound zone b where user B is located are both sound zones on the left side of the center control screen 10 (such as the second row left sound zone, the third row left sound zone, etc.).

[0262] In some embodiments, after the intelligent vehicle 100 executes the wake-up-free instruction issued by user A and the wake-up-free instruction issued by user B (for example, the intelligent vehicle 100 turns on the air conditioner and opens the window), the intelligent vehicle 100 can control the center control screen 10 to stop displaying the composite wake-up-free identifier 810 and redisplay the main interface 800 as Figure 8C shown. The main interface 800 does not include a wake-up-free identifier, nor a composite wake-up-free identifier, nor a sound zone indicator.

[0263] It can be understood that the above Figures 8A to 8CThe illustrated embodiments are merely examples. In the embodiments of the present application, the user can also perform voice interaction with other display screens (such as the laser curtain 40, etc.) through the wake-up-free instruction, and the interface 1 displayed on the display screen A can also be other interfaces other than the main interface 800. The present application does not make any limitations here. In addition, the wake-up-free identifier can also adopt a shape and / or symbol different from those of the above embodiments, or include content different from those of the above embodiments. Moreover, the wake-up-free instructions issued by user A and user B can also be wake-up-free instructions different from those of the above embodiments. The present application does not make any limitations here.

[0264] In some other embodiments, one or more display screens of the intelligent vehicle 100 (such as the central control screen 10, the laser curtain 40, etc.) can also receive and process three or more wake-up-free instructions at the same time. In this case, the display screen can indicate the number of wake-up-free instructions processed by the display screen at the same time through the digital indicator displayed in the composite wake-up-free identifier. The present application does not make any limitations here.

[0265] In some application scenarios, the intelligent vehicle 100 carries user A and user B. User A is located in sound zone a, and user B is located in sound zone b. When the interface 1 is displayed on the display screen A of the intelligent vehicle 100 and the interface 2 is displayed on the display screen B, the intelligent vehicle 100 receives the wake-up-free instruction issued by user A from sound zone a. In response to the wake-up-free instruction issued by user A from sound zone a, the intelligent vehicle 100 can display the wake-up-free identifier 1 on the interface 1 of the display screen A. The wake-up-free identifier 1 is used to prompt the user that the display screen A is receiving and processing the user's wake-up-free instruction. When the wake-up-free identifier 1 is displayed on the display screen A, the intelligent vehicle 100 can receive and respond to the wake-up word issued by the user from sound zone b. When it is determined that the display screen corresponding to sound zone b is the display screen B, the intelligent vehicle 100 can determine the display screen B as the main display screen and display the voice dialogue identifier 1 on the interface 2 of the display screen B. The voice dialogue identifier 1 is used to prompt the user that the voice assistant of the display screen B has been activated through the wake-up word.

[0266] In this way, during the process of receiving and processing the user's wake-up-free instruction, another user can wake up the voice assistant of another display screen through the wake-up word. The intelligent vehicle 100 can perform voice interaction with multiple users through multiple display screens, providing more convenient services for the users.

[0267] Exemplarily, when the intelligent vehicle 100 receives the wake-up-free instruction issued by user A from sound zone a and determines that the display screen corresponding to sound zone a is the right rear projection screen 30, the right rear projection screen 30 can display the following Figure 9A shown main interface 900, and the central control screen 10 can display the following Figure 9B shown main interface 910.

[0268] As Figure 9AAs shown, the right-back projection screen 30 of the intelligent vehicle 100 displays the main interface 900. The main interface 900 can display a wallpaper, such as a landscape picture. The main interface 900 can display a wake-up-free identifier 901. The specific content of the wake-up-free identifier 901 can refer to the relevant description in the above Figure 7C illustrated embodiment. In addition, for the specific process of the intelligent vehicle 100 receiving and responding to the wake-up-free instruction issued by user A and displaying the wake-up-free identifier 901, it can also refer to the relevant content in the above Figure 7A and Figure 7C illustrated embodiment, which will not be elaborated here.

[0269] As Figure 9B shown, the center control screen 10 can display the main interface 910. The main interface 910 can display one or more application icons, for example, a music application icon, a navigation application icon, a settings application icon, etc.

[0270] During the process of the right-back projection screen 30 processing the wake-up-free instruction of user A, the intelligent vehicle 100 can also receive the wake-up word issued by user B from sound zone b. In response to the wake-up word issued by user B from sound zone b, when it is determined that the display screen corresponding to sound zone b is the center control screen 10, the intelligent vehicle 100 can determine the center control screen 10 as the main display screen and display the voice dialogue identifier 911 as shown in Figure 9C on the main interface 910 of the center control screen 10.

[0271] As Figure 9C shown, the voice dialogue identifier 911 is displayed on the main interface 910. The voice dialogue identifier 911 is used to prompt the user that the voice assistant of the center control screen 10 has been activated by the wake-up word. The specific content of the voice dialogue identifier 911 can refer to the relevant description of the voice dialogue identifier 501 in the above Figure 5B illustrated embodiment, which will not be elaborated here. Optionally, the main interface 910 can also display a sound zone indicator 912. The specific content of the sound zone indicator 912 can refer to the relevant content in the above Figure 5B illustrated embodiment, which will not be elaborated here.

[0272] After the center control screen 10 activates the voice assistant, the intelligent vehicle 100 can receive and respond to the voice instruction (such as "Play a song") issued by user B from sound zone b. As shown in Figure 9D the center control screen 10 can display the text 913 in the voice dialogue identifier 911. The text 913 can be the text content of the voice instruction issued by the user, such as "Play a song". The text 913 can be used to prompt the user that the center control screen 10 is receiving and processing the voice instruction issued by the user.

[0273] It should be noted that after the intelligent vehicle 100 executes the hands-free wake-up instruction issued by user A, for example, after opening the window, the intelligent vehicle 100 can control the right rear projection screen 30 to stop displaying the hands-free wake-up logo 901. Optionally, it can also control the right rear projection screen 30 to output feedback information. For the specific content of the feedback information, reference can be made to the relevant description in the above Figure 6D illustrated embodiment, which will not be elaborated here. In addition, when the intelligent vehicle 100 detects that the center control screen 10 meets the shutdown condition, it can also turn off the voice assistant of the center control screen 10 and stop displaying the voice dialogue logo.

[0274] It can be understood that the above Figures 9A to 9D illustrated embodiment is just an example. In the embodiments of the present application, the user can also perform voice interaction with other display screens (such as the center control screen 10, etc.) through the hands-free wake-up instruction, and the interface 1 displayed on the display screen A can also be other interfaces other than the main interface 900. The present application does not make any limitations here. In addition, the hands-free wake-up logo and the voice dialogue logo can also adopt shapes and / or symbols different from those in the above embodiments, or include contents different from those in the above embodiments. Moreover, the hands-free wake-up instructions issued by user A and user B can also be hands-free wake-up instructions different from those in the above embodiments. The present application does not make any limitations here.

[0275] In some application scenarios, the intelligent vehicle 100 carries user A and user B. User A is located in sound zone a, and user B is located in sound zone b. When the interface 1 is displayed on the display screen A of the intelligent vehicle 100, the intelligent vehicle 100 receives the hands-free wake-up instruction issued by user A from sound zone a. In response to the hands-free wake-up instruction issued by user A from sound zone a, the intelligent vehicle 100 can display the hands-free wake-up logo 1 on the interface 1 of the display screen A. The hands-free wake-up logo 1 is used to prompt the user that the display screen A is receiving and processing the user's hands-free wake-up instruction. When the hands-free wake-up logo 1 is displayed on the display screen A, the intelligent vehicle 100 can receive and respond to the wake-up word issued by the user from sound zone b. When it is determined that the display screen corresponding to sound zone b is the display screen A, the intelligent vehicle 100 can split the hands-free wake-up logo 1 on the interface 1 into a voice dialogue logo and a voice sub-logo. The voice dialogue logo and the voice sub-logo can be used to prompt the user that the display screen A is simultaneously processing the voices of multiple users, and one of the users has activated the voice assistant of the display screen A through the wake-up word.

[0276] In this way, the intelligent vehicle 100 can simultaneously receive and process the voice instructions of different users, and conduct voice interaction with multiple users through the same display screen, providing more convenient services for the users.

[0277] Exemplarily, when the intelligent vehicle 100 receives the hands-free wake-up instruction issued by user A from sound zone a and determines that the display screen corresponding to sound zone a is the center control screen 10, the intelligent vehicle 100 controls the center control screen 10 to display as Figure 10AThe main interface 1010 shown.

[0278] As Figure 10A shown, the central control screen 10 can display the main interface 1010. In the main interface 1010, a wake-up-free identifier 1011 can be displayed. The wake-up-free identifier 1011 is used to prompt the user that the central control screen 10 is receiving and processing the user's wake-up-free instruction. Among them, the relevant content of the main interface 1010 and the wake-up-free identifier 1011 can refer to the relevant content in the above Figure 5A shown embodiment and the above Figure 7C shown embodiment respectively, which will not be elaborated here. In addition, the specific process of the central control screen 10 displaying the wake-up-free identifier 1011 in response to the wake-up-free instruction issued by user A can also refer to the relevant content in the above Figures 7A to 7C shown embodiment, which will not be elaborated here. Optionally, the main interface 1010 can also display a sound area indicator 1012. The sound area indicator 1012 is used to indicate the positional relationship between the sound area a where user A is located and the central control screen 10. For example, Figure 10A shown, the display area of the sound area indicator 1012 includes the lower left of the wake-up-free identifier 1011, indicating that the sound area a is located on the left side of the central control screen 10 (such as the left sound area in the second row, etc.).

[0279] During the process of the central control screen 10 processing the wake-up-free instruction of user A, the intelligent vehicle 100 can also receive the wake-up word issued by user B from the sound area b. In response to the wake-up word issued by user B from the sound area b, when it is determined that the display screen corresponding to the sound area b is the central control screen 10, the intelligent vehicle 100 can determine the central control screen 10 as the main display screen and display a split identifier 1020 as shown in Figure 10B on the main interface 1010 of the central control screen 10.

[0280] As Figure 10B shown, the split identifier 1020 is displayed in the main interface 1010. The split identifier 1020 can include two voice identifiers that are not completely split. One of the voice identifiers is a voice conversation identifier 1021, and the other is a voice sub-identifier 1022. And there can also be a connection channel between the voice conversation identifier 1021 and the voice sub-identifier 1022, indicating that the two voice identifiers are still in the process of splitting and have not been completely split. The split identifier 1020 can be used to indicate that the central control screen 10 is processing the voices of multiple users simultaneously. Among them, the specific content of the voice conversation identifier 1021 can refer to the relevant description of the voice conversation identifier 501 in the above Figure 5B shown embodiment, which will not be elaborated here. The voice sub-identifier 1022 can include a rounded rectangle frame and a sub-symbol 1024. The sub-symbol 1024 can be the same as the above Figure 5DIt is the same as the sub-symbol 531 shown. Optionally, a voice area indicator 1023 may also be displayed on the main interface 1010. The display area of the voice area indicator 1023 is the area below the split identifier 1020 (including the lower left area, the directly below area, and the lower right area). The voice area indicator 1023 is used to indicate that the voice areas of the multiple channels of voice being processed by the central control screen 10 are located in different orientations of the central control screen 10. That is, the voice area a where user A is located is the voice area on the left side of the central control screen 10, and the voice area b where user B is located is the voice area on the right side of the central control screen 10 (such as the right voice area in the third row).

[0281] After the split identifier 1020 is completely split, as Figure 10C shown, a completely split voice conversation identifier 1021 and a voice sub-identifier 1022 may be displayed on the main interface 1010 of the central control screen 10, that is, there is no connection channel between the voice conversation identifier 1021 and the voice sub-identifier 1022.

[0282] It should be noted that after the intelligent vehicle 100 executes the wake-up-free instruction issued by user A, for example, after opening the window, the intelligent vehicle 100 can control the central control screen 10 to stop displaying the wake-up-free identifier 901. Optionally, it can also control the central control screen 10 to output feedback information. The specific content of the feedback information can refer to the relevant description in the above Figure 6D shown embodiments and will not be elaborated here. In addition, when the intelligent vehicle 100 detects that the central control screen 10 meets the closing conditions, it can also close the voice assistant of the central control screen 10 and stop displaying the voice conversation identifier.

[0283] It can be understood that the above Figures 10A to 10C shown embodiments are only examples. In the embodiments of the present application, the user can also wake up the voice assistants of other display screens (such as the co-pilot screen 20, the right rear projection screen 30, the laser curtain 40, etc.) through wake-up words, and the interface 1 displayed on the display screen A can also be other interfaces other than the main interface. In addition, the wake-up-free identifier, the voice conversation identifier, and the voice sub-identifier can also adopt shapes and / or symbols different from those in the above embodiments, or include contents different from those in the above embodiments. In addition, the voice commands issued by user A and user B can also be voice commands different from those in the above embodiments, and the present application does not make any limitations here.

[0284] In some application scenarios, the intelligent vehicle 100 carries user A and user B. User A is located within sound zone a, and user B is located within sound zone b. When interface 1 is displayed on display screen A of the intelligent vehicle 100 and interface 2 is displayed on display screen B, the intelligent vehicle 100 receives a wake-up word issued by user A from sound zone a. In response to the wake-up word issued by user A from sound zone a, the intelligent vehicle 100 can display a voice dialogue identifier 1 on interface 1 of display screen A. The voice dialogue identifier 1 is used to prompt the user that the voice assistant of display screen A has been activated through the wake-up word. When the voice dialogue identifier 1 is displayed on display screen A, the intelligent vehicle 100 can receive and respond to a wake-up word issued by the user from sound zone b. When it is determined that the display screen corresponding to sound zone b is display screen B, the intelligent vehicle 100 can turn off the voice dialogue identifier 1 displayed on display screen A, determine display screen B as the new main display screen, and display a voice dialogue identifier 2 on display screen B. The voice dialogue identifier 2 is used to prompt the user that the voice assistant of display screen B has been activated through the wake-up word.

[0285] In this way, when the intelligent vehicle 100 receives a wake-up word from another user, it can interrupt the voice interaction between the user who previously used the wake-up word and the display screen, ensuring that at the same moment, among the voices that need to be processed in the intelligent vehicle 100, only one voice is not a hands-free wake-up command.

[0286] Exemplarily, when the intelligent vehicle 100 receives a wake-up word issued by user A from sound zone a and determines that the display screen corresponding to sound zone a is the central control screen 10, the intelligent vehicle 100 controls the central control screen 10 to display the main interface 1010 as shown in Figure 10A At the same time, the right rear projection screen 30 can display the main interface 1110 as shown in Figure 11B below.

[0287] As shown in Figure 11A below, the central control screen 10 can display a main interface 1100. In the main interface 1100, a voice dialogue identifier 1101 can be displayed. The voice dialogue identifier 1101 is used to prompt the user that the voice assistant of the central control screen 10 has been activated through the wake-up word and that the user is currently having a voice interaction with the central control screen 10 through the voice assistant. Text 1103 can also be displayed in the voice dialogue identifier 1101. The text 1103 can be part or all of the text content of the voice command issued by user A after issuing the wake-up word, such as "The weather is nice today. Open the window." Optionally, a sound zone indicator 1102 can also be displayed in the main interface 1100. For the specific content of the sound zone indicator 1102, reference can be made to the relevant content in the embodiment shown in Figure 5B below, and details will not be elaborated here.

[0288] As shown in Figure 11BAs shown, the right rear projection screen 30 of the intelligent vehicle 100 displays the main interface 1110, and a wallpaper, such as a landscape picture, can be displayed in the main interface 1110.

[0289] During the process of the central control screen 10 processing the voice command issued by user A after the wake-up word, the intelligent vehicle 100 can also receive the wake-up word sent by user B from the sound zone b. In response to the wake-up word sent by user B from the sound zone b, when it is determined that the display screen corresponding to the sound zone b is the right rear projection screen 30, the intelligent vehicle 100 can determine the right rear projection screen 30 as the new main display screen. As Figure 11C shown, control the central control screen 10 to stop displaying the voice dialogue identifier 1101, and control the right rear projection screen 30 to display the voice dialogue identifier 1111 as Figure 11D shown.

[0290] As Figure 11C shown, the central control screen 10 displays the main interface 1100, and the main interface 1100 does not include the voice dialogue identifier 1101. Optionally, while stopping the display of the voice dialogue identifier 1101, the central control screen 10 can also display an interruption prompt 1103 in the main interface 1100. The interruption prompt 1103 can be used to prompt the user that the current voice interaction is interrupted. For example, the text "Voice interaction interrupted" can be displayed in the interruption prompt 1103. In some other embodiments, the central control screen 10 can also output the interruption prompt in one or more ways such as voice broadcast, vibration, and indicator light flashing. This application does not make any limitations here.

[0291] As Figure 11D shown, the right rear projection screen 30 displays the main interface 1110, and the voice dialogue identifier 1111 is displayed in the main interface 1110. The voice dialogue identifier 1111 can be used to prompt the user that the voice assistant of the right rear projection screen 30 has been activated through the wake-up word. The specific content of the voice dialogue identifier 1111 can refer to the relevant description in the above Figure 5B shown embodiments and will not be elaborated here.

[0292] It should be noted that after detecting that the display duration of the interruption prompt 1103 is greater than the preset duration (such as 5 seconds, etc.), the intelligent vehicle 100 can stop displaying the interruption prompt 1103. In addition, when the intelligent vehicle 100 detects that the right rear projection screen 30 meets the closing condition, it can also close the voice assistant of the right rear projection screen 30 and stop displaying the voice dialogue identifier.

[0293] It can be understood that the above Figures 11A to 11DThe illustrated embodiments are merely examples. In the embodiments of the present application, the user can also activate the voice assistants of other display screens (such as the co-pilot screen 20, the laser curtain 40, etc.) through the wake-up word. Moreover, the interface 1 displayed on the display screen A and the interface 2 displayed on the display screen B can also be other interfaces other than the main interface. In addition, the voice dialogue identifier and the interruption prompt can also adopt shapes and / or symbols different from those of the above embodiments, or include contents different from those of the above embodiments. In addition, the voice commands issued by user A and user B can also be voice commands different from those of the above embodiments. The present application does not make any limitations here.

[0294] In some application scenarios, the intelligent vehicle 100 carries user A and user B. User A is located in the sound zone a, and user B is located in the sound zone b. When the interface 1 is displayed on the display screen A of the intelligent vehicle 100 and the interface 2 is displayed on the display screen B, the intelligent vehicle 100 receives the wake-up word issued by user A from the sound zone a. In response to the wake-up word issued by user A from the sound zone a, the intelligent vehicle 100 can display the voice dialogue identifier 1 on the interface 1 of the display screen A. The voice dialogue identifier 1 is used to prompt that the user has activated the voice assistant of the display screen A through the wake-up word. When the voice dialogue identifier 1 is displayed on the display screen A, the intelligent vehicle 100 can receive and respond to the wake-up word issued by the user from the sound zone b. When it is determined that the display screen corresponding to the sound zone b is the display screen A, the intelligent vehicle 100 can turn off the voice dialogue identifier 1 displayed on the display screen A and display the voice dialogue identifier 2 on the display screen A. The voice dialogue identifier 2 is used to prompt that the user has activated the voice assistant of the display screen A through the wake-up word.

[0295] In this way, when the intelligent vehicle 100 receives the wake-up word of another user, it can interrupt the voice interaction between the user who previously used the wake-up word and the display screen, and it can ensure that at the same moment, among the voices that need to be processed in the intelligent vehicle 100, only one voice is not a hands-free wake-up command.

[0296] Exemplarily, when the intelligent vehicle 100 receives the wake-up word issued by user A from the sound zone a and determines that the display screen corresponding to the sound zone a is the central control screen 10, the intelligent vehicle 100 controls the central control screen 10 to display as Figure 10A the shown main interface 1010.

[0297] As Figure 12AAs shown, the central control screen 10 can display a main interface 1200. In the main interface 1200, a voice dialogue identifier 1201 can be displayed. The voice dialogue identifier 1201 is used to prompt the user that the voice assistant of the central control screen 10 has been activated by a wake-up word and that a voice interaction is being carried out with the central control screen 10 through the voice assistant. Text 1203 can also be displayed in the voice dialogue identifier 1201. The text 1203 can be part or all of the text content of the voice command issued by user A after issuing the wake-up word. For example, "The weather is nice today. Open the window." Optionally, a sound zone indicator 1202 can also be displayed in the main interface 1200. The display area of the sound zone indicator 1202 can be the lower left of the voice dialogue identifier 1201, and is used to prompt that the sound zone a where user A is located is on the left side of the central control screen 10.

[0298] During the process of the central control screen 10 processing the voice command issued by user A after the wake-up word, the intelligent vehicle 100 can also receive a wake-up word issued by user B from sound zone b. In response to the wake-up word issued by user B from sound zone b, when it is determined that the display screen corresponding to sound zone b is the central control screen 10, as Figure 12B shown, the intelligent vehicle 100 can control the central control screen 10 to stop displaying the voice dialogue identifier 1201 and display a voice dialogue identifier 1211 in the main interface 1200. Alternatively, the intelligent vehicle 100 can also control the central control screen 10 to update the content (such as text) in the voice dialogue identifier 1201 and change the voice dialogue identifier 1201 to Figure 12B the voice dialogue identifier 1211 shown.

[0299] As Figure 12B shown, the central control screen 10 displays the main interface 1200, and the main interface 1200 displays a voice dialogue identifier 1211. Text 1213 can be displayed in the voice dialogue identifier 1211. The text 1213 is different from the text in the Figure 12A voice dialogue identifier 1201 shown. For example, the text 1213 can be "You can talk at any time." Optionally, a sound zone indicator 1212 can also be displayed in the main interface 1200. The display area of the sound zone indicator 1212 can be the lower right of the voice dialogue identifier 1211, and is used to prompt that the sound zone b where user B is located is on the right side of the central control screen 10. Further optionally, the central control screen 10 can also output an interruption prompt in one or more ways such as display on the screen, voice broadcast, vibration, indicator light flashing, etc. The specific content of the interruption prompt can refer to the relevant description in the Figure 11C embodiment shown above. This application will not elaborate further here.

[0300] It should be noted that when the intelligent vehicle 100 detects that the central control screen 10 meets the closing conditions, the voice assistant of the central control screen 10 can also be closed, and the voice dialogue identifier display can be stopped.

[0301] It can be understood that the above Figures 12A to 12B illustrated embodiments are merely examples. In the embodiments of the present application, the user can also activate the voice assistants of other display screens (such as the co-pilot screen 20, the right-back projection screen 30, the laser curtain 40, etc.) through the wake-up word, and the interface 1 displayed on the display screen A can also be other interfaces other than the main interface. In addition, the voice dialogue identifier and the interruption prompt can also adopt shapes and / or symbols different from those of the above embodiments, or include contents different from those of the above embodiments. In addition, the voice commands issued by user A and user B can also be voice commands different from those of the above embodiments, and the present application does not make any limitations here.

[0302] The following introduces the process of a voice interaction method provided by the embodiments of the present application.

[0303] As Figure 13 shown, the specific process for the intelligent vehicle 100 to execute the voice interaction method may include the following steps:

[0304] S1301, the intelligent vehicle 100 receives the voice 1 issued by the user.

[0305] The intelligent vehicle 100 can receive the voice issued by the user through one or more microphones set inside.

[0306] S1302, the intelligent vehicle 100 determines the sound zone 1 corresponding to the voice 1.

[0307] In some embodiments, the intelligent vehicle 100 can respectively set microphones in different sound zones inside the vehicle, and the intelligent vehicle 100 can store one or more sound source localization algorithms. When a user inside the vehicle (such as user A) issues a voice, since the azimuth and distance of this user A relative to each microphone are different, the moments when the voice reaches each microphone are different, and the arrival directions are also different. The intelligent vehicle 100 can determine the sound zone where the user A is located based on the sound source localization algorithm according to the voices collected by multiple microphones. In this way, after collecting the voice 1, the intelligent vehicle 100 can determine the sound zone where the user who issued the voice 1 is located based on the sound source localization algorithm, and this sound zone is the sound zone 1 corresponding to the voice 1.

[0308] In some other embodiments, the intelligent vehicle 100 may be provided with microphones in different sound zones inside the vehicle, and each microphone may have its own device identifier, such as the identity document (ID) of the microphone, etc. When a vehicle user (such as user A) emits a voice, the microphone in the same sound zone as the user A can collect the voice emitted by the user A and report the voice carrying the device identifier of the microphone to the intelligent vehicle 100. The intelligent vehicle 100 can determine the microphone that reports the voice based on the device identifier of the microphone, and thus determine the sound zone corresponding to the voice based on the sound zone where the microphone is located. In this way, after collecting voice 1, the intelligent vehicle 100 can determine the microphone that reports voice 1 based on voice 1, and further determine the sound zone 1 corresponding to voice 1.

[0309] It should be noted that, in some embodiments, when multiple users in the vehicle emit voices, the microphone can determine whether the collected voice comes from its own sound zone (i.e., the sound zone where the microphone is located) based on any one or more of the parameters such as the angle of arrival, volume, time delay, voiceprint, and frequency of the voice. After that, the microphone can filter the collected signal and only retain the voice in its own sound zone, and filter out the voices in other sound zones.

[0310] Exemplarily, in the intelligent vehicle 100, microphone 1 is located in sound zone 1, microphone 2 is located in sound zone 2, user A often sits in the area where sound zone 1 is located, and user B often sits in the area where sound zone 2 is located. After the intelligent vehicle 100 starts, user A speaks voice a, and at the same time, user B also speaks voice b. Microphone 1 can collect voice a emitted by user A and voice b emitted by user B, and microphone 2 can collect voice a emitted by user A and voice B emitted by user B. In this case, microphone 1 can identify the voiceprint feature of user A based on the voice data collected by microphone 1 in the past, and determine to retain voice a and filter voice b based on the voiceprint feature of user A. Microphone 2 can also use a similar method to determine to retain voice b and filter voice a based on the voiceprint feature of user B. After that, microphone 1 can report voice a carrying the device identifier of microphone 1 to the intelligent vehicle 100, and microphone 2 can report voice b carrying the device identifier of microphone 2 to the intelligent vehicle 100. The intelligent vehicle 100 can determine that the sound zone corresponding to voice a is sound zone 1 and the sound zone corresponding to voice b is sound zone 2 based on the device identifiers carried by voice a and voice b.

[0311] It can be understood that the above embodiments only exemplarily illustrate two ways to determine the sound zone corresponding to the voice. In the embodiments of the present application, the intelligent vehicle 100 can also use more or different ways from the above embodiments to determine the sound zone corresponding to the voice, and the present application does not make any limitation here.

[0312] S1303, the intelligent vehicle 100 determines whether the voice 1 contains a wake-up word.

[0313] The intelligent vehicle 100 may store one or more wake-up words, such as "Xia A", etc.

[0314] If the voice 1 includes a wake-up word, the intelligent vehicle 100 may perform the following step S1304.

[0315] If the voice 1 does not include a wake-up word, the intelligent vehicle 100 may perform the following step S1307.

[0316] S1304, the intelligent vehicle 100 determines whether there is a main display screen, and the main display screen is the display screen that is processing non-wake-up-free instructions through the voice assistant.

[0317] The main display screen refers to the display screen that is processing non-wake-up-free instructions through the voice assistant. A non-wake-up-free instruction is a voice instruction that does not belong to a wake-up-free instruction, and a non-wake-up-free instruction may include a wake-up word. According to the noun explanation in the embodiments of the present application, only when the user activates the voice assistant of the display screen through the wake-up word, can the display screen process the user's non-wake-up-free instructions through the voice assistant. Therefore, in the embodiments of the present application, the main display screen may also be the display screen whose voice assistant is activated by the user through the wake-up word, or the display screen whose voice dialogue identifier is triggered by the user through the wake-up word.

[0318] At the same moment, among one or more display screens of the intelligent vehicle 100, there is at most one main display screen. It should be noted that compared with non-wake-up-free instructions, the execution process of wake-up-free instructions is shorter and consumes fewer resources. The intelligent vehicle 100 may be set such that among the multiple voices processed at the same moment, at most one voice is a non-wake-up-free instruction. In this way, the intelligent vehicle 100 can ensure that each user's voice instruction can be promptly feedback. Therefore, at the same moment, there is at most one main display screen in the intelligent vehicle 100, and among the one or more voices processed by the main display screen, at most one voice is a non-wake-up-free instruction.

[0319] In some embodiments, when the intelligent vehicle 100 receives a wake-up word, it may determine the display screen corresponding to the sound area based on the sound area corresponding to the wake-up word voice, and determine this display screen as the main display screen, and store the device identifier of this display screen as the device identifier of the main display screen. In this case, when performing step S1304, the intelligent vehicle 100 may determine whether there is a main display screen currently based on whether there is a stored device identifier of the main display screen.

[0320] In some embodiments, the intelligent vehicle 100 may determine whether each display screen is processing a non-wake-up-free instruction. If there is a display screen that is processing a non-wake-up-free instruction, then that display screen is the main display screen, and the intelligent vehicle 100 may determine that there is a main display screen currently; if there is no display screen that is processing a non-wake-up-free instruction, then the intelligent vehicle 100 may determine that there is no main display screen currently.

[0321] In some embodiments, the intelligent vehicle 100 may determine whether a voice dialogue identifier is displayed on each display screen. If there is a display screen on which a voice dialogue identifier is displayed, then the intelligent vehicle 100 may determine that that display screen is the main display screen, that is, there is a main display screen currently; if there is no display screen on which a voice dialogue identifier is displayed, then the intelligent vehicle 100 may determine that there is no main display screen currently.

[0322] If the intelligent vehicle 100 determines that there is a main display screen, the intelligent vehicle 100 may perform the following step S1305.

[0323] If the intelligent vehicle 100 determines that there is no main display screen, the intelligent vehicle 100 may perform the following step S1306.

[0324] S1305, the intelligent vehicle 100 turns off the voice dialogue identifier of the main display screen.

[0325] When the intelligent vehicle 100 receives a wake-up word and there is already a main display screen currently, the intelligent vehicle 100 may turn off the voice dialogue opened by the previous user (such as user C) through the wake-up word, that is, turn off the voice dialogue identifier of the main display screen and stop the voice interaction between the original main display screen and user C.

[0326] Exemplarily, referring to the above Figure 11A and Figure 11C shown embodiments, the voice 1 may be the wake-up word issued by user B in the above embodiments. When the central control screen 10 already displays the voice dialogue identifier 1101 and is receiving and processing the non-wake-up-free instruction of the user through the voice assistant, when the intelligent vehicle 100 receives a wake-up word issued by another user, the intelligent vehicle 100 may turn off the voice dialogue identifier 1101 displayed on the central control screen 10.

[0327] Again exemplarily, referring to the above Figure 12A and Figure 12B shown embodiments, the voice 1 may be the wake-up word issued by user B in the above embodiments. When the central control screen 10 already displays the voice dialogue identifier 1201 and is receiving and processing the non-wake-up-free instruction of the user through the voice assistant, when the intelligent vehicle 100 receives a wake-up word issued by another user, the intelligent vehicle 100 may turn off the voice dialogue identifier 1201 displayed on the central control screen 10.

[0328] S1306, the intelligent vehicle 100 displays a voice dialogue identifier on the display screen A corresponding to sound zone 1.

[0329] The intelligent vehicle 100 may store the correspondence between sound zones and display screens. Exemplarily, the correspondence between sound zones and display screens may refer to the relevant descriptions in the above Figures 3E to 3F illustrated embodiments.

[0330] In some embodiments, the state of the display screen (such as the on state and the off state) also affects the correspondence between sound zones and display screens. The intelligent vehicle 100 may also store the correspondence between the display screen state, sound zones, and display screens. Exemplarily, when the display screen is in different states, the correspondence between sound zones and display screens may also refer to the relevant descriptions in the above Figures 3E to 3F illustrated embodiments, which will not be elaborated here.

[0331] It can be understood that the above Figures 3E to 3H are only some examples. In the embodiments of the present application, the division method of sound zones, the number of display screens, the setting method of display screens, and the correspondence between sound zones and display screens may all be different from the above embodiments, and the present application does not make any limitations here.

[0332] The intelligent vehicle 100 may determine the display screen A corresponding to sound zone 1 from the correspondence between sound zones and display screens based on sound zone 1. In some embodiments, the intelligent vehicle 100 may also determine the display screen A corresponding to sound zone 1 based on the current display screen state, the correspondence between sound zones, display screen states, and display screens.

[0333] When the intelligent vehicle 100 receives a wake-up word and there is a main display screen currently, the intelligent vehicle 100 may, while turning off the voice dialogue identifier of the main display screen, display the voice dialogue identifier on the display screen A corresponding to sound zone 1.

[0334] Exemplarily, referring to the above Figures 11A to 11D illustrated embodiments, in the above embodiments, the display screen A may be the right rear projection screen 30, and the voice dialogue identifier may be Figure 11D the shown voice dialogue identifier 1111. Another example, referring to the above Figures 12A to 12B illustrated embodiments, in the above embodiments, the display screen A may be the center control screen 10, and the newly displayed voice dialogue identifier may be Figure 12B the shown voice dialogue identifier 1211.

[0335] When the intelligent vehicle 100 receives a wake-up word and there is no main display screen currently, the intelligent vehicle 100 may display the voice dialogue identifier on the display screen A corresponding to sound zone 1. Exemplarily, referring to the above Figures 5A to 5BIn the illustrated embodiment, the display screen A may be the central control screen 10, and the voice dialogue identifier may be the above-mentioned Figure 5B voice dialogue identifier 501 shown above.

[0336] In some embodiments, if the display screen A is the same as the previous main display screen, the main display screen may also change the text displayed in the voice dialogue identifier, and the newly displayed text may be used to prompt the user that the original voice interaction has been interrupted and a new round of voice interaction has been started.

[0337] In some embodiments, when the intelligent vehicle 100 displays the voice dialogue identifier on the display screen A, it may also display a sound zone indicator on the display screen A. The sound zone indicator is used to indicate the relative position of the source sound zone of the voice being processed by the display screen A with respect to the display screen A.

[0338] It should be noted that in a possible implementation manner, the intelligent vehicle 100 may only display the sound zone indicator on the main display screen, and the sound zone indicator is used to indicate the relative position of the source sound zone of one or more voices being processed by the main display screen with respect to the main display screen. In this case, the sound zone indicator may also be used to prompt the user to continue issuing voice commands for voice interaction with the main display screen. Moreover, the sound zone indicator does not need to be displayed on the secondary display screen.

[0339] In another possible implementation manner, the intelligent vehicle 100 may also determine whether to display the sound zone indicator on the display screen based on whether the display screen has the ability to process multiple voices. For example, when the display screen A is a display screen such as the central control screen 10 or the laser curtain 40 that can process multiple voices, the display screen A may display the sound zone indicator on the display screen A in response to a wake-up word or a hands-free wake-up instruction issued from sound zone 1 (or other corresponding sound zones); when the display screen A is a display screen such as the co-pilot screen 20 or the right rear projection screen 30 that only needs to process one voice, the display screen A may not display the sound zone indicator. Exemplarily, the sound zone indicator may be the above-mentioned Figure 5B sound zone indicator 505 shown above, or may also be Figure 9C sound zone indicator 912 shown above, etc.

[0340] In another possible implementation manner, the intelligent vehicle 100 may also display the sound zone indicator on the corresponding display screen when receiving a wake-up word or a hands-free wake-up instruction. In another possible implementation manner, the intelligent vehicle 100 may also not display the sound zone indicator on all display screens.

[0341] After executing step S1306, the intelligent vehicle 100 may execute the following step S1309.

[0342] S1307, the intelligent vehicle 100 determines whether the voice 1 belongs to a hands-free wake-up instruction.

[0343] One or more wake - up - free instructions can be stored in the intelligent vehicle 100. Among them, the wake - up - free instructions can be pre - set, or can be determined based on the usage frequency of different voice instructions by the user during the user's usage process. For example, the intelligent vehicle 100 can record the voice instructions used by the user, and set the voice instructions with the usage times greater than a preset number of times (such as 5 times, etc.) within a period of time (such as within a week) by the user as wake - up - free instructions.

[0344] Exemplarily, Table 1 shows a plurality of wake - up - free instructions stored in the intelligent vehicle 100 provided by an embodiment of the present application.

[0345] Table 1

[0346] Wake - up - free command Turn on the air conditioner Open the window Start navigation, go home Play music

[0347] As shown in Table 1, the intelligent vehicle 100 can store one or more wake - up - free instructions. For example, "turn on the air conditioner", "open the window", "turn on the navigation and go home", and "play music", etc.

[0348] It can be understood that the embodiment shown in Table 1 only exemplarily illustrates that the intelligent vehicle 100 can store one or more wake - up - free instructions. In the embodiment of the present application, the intelligent vehicle 100 can store more or fewer wake - up - free instructions than those in Table 1 above, and the present application does not make any limitations here.

[0349] After the intelligent vehicle 100 recognizes the text content of Voice 1, based on the text content of Voice 1 and the stored wake - up - free instructions, it can judge whether Voice 1 belongs to a wake - up - free instruction. If the text content of Voice 1 includes any one of the wake - up - free instructions, the intelligent vehicle 100 can determine that Voice 1 belongs to a wake - up - free instruction; if the text content of Voice 1 does not include any one of the wake - up - free instructions, the intelligent vehicle 100 can determine that Voice 1 does not belong to a wake - up - free instruction.

[0350] If Voice 1 belongs to a wake - up - free instruction, the intelligent vehicle 100 can execute the following step S1308. Optionally, if the intelligent vehicle 100 stores the feedback information corresponding to this wake - up - free instruction, the intelligent vehicle 100 can also output this feedback information.

[0351] If Voice 1 does not belong to a wake - up - free instruction, the intelligent vehicle 100 can execute the following step S1309.

[0352] S1308, the intelligent vehicle 100 displays a wake - up - free identifier, a voice sub - identifier, or a composite wake - up - free identifier on the display screen A corresponding to Zone 1.

[0353] The intelligent vehicle 100 can determine the display screen A corresponding to the sound zone 1 based on the correspondence between the sound zone 1 and the display screen. The specific determination method can refer to the relevant description in the above step S1306.

[0354] When the voice 1 is a hands-free wake-up command, the intelligent vehicle 100 can control the display screen A to display any one of the following voice identifiers: hands-free wake-up identifier, voice sub-identifier, and composite hands-free wake-up identifier. Among them, in different scenarios, the voice identifiers displayed on the display screen A can be different. The specific relationship between the scenario and the voice identifier, as well as the method for judging the scenario, can all refer to the Figure 14 illustrated embodiments below and will not be elaborated here for the time being.

[0355] S1309, the intelligent vehicle 100 determines whether there is a main display screen and whether the voice 1 is issued by user A, where user A is the user who activates the voice assistant of the main display screen through a wake-up word.

[0356] The specific method for the intelligent vehicle 100 to determine whether there is a main display screen can refer to the relevant content in the above step S1304 and will not be elaborated here.

[0357] In the case where there is no main display screen, the intelligent vehicle 100 can stop processing the voice 1.

[0358] In the case where there is a main display screen, the intelligent vehicle 100 can further determine whether the voice 1 is issued by user A, where user A is the user who activates the voice assistant of the main display screen through a wake-up word, that is, user A is the user who is currently having a voice interaction with the original main display screen and the voice command issued is a non-hands-free wake-up command.

[0359] The intelligent vehicle 100 can determine whether the voice 1 is issued by user A based on whether the voiceprint information of the voice 1 is consistent with the voiceprint information of user A. The voiceprint information of user A can be obtained based on the voice commands (including wake-up words, etc.) collected during the previous voice interaction process. If the voiceprint information of the voice 1 is consistent with that of user A, the intelligent vehicle 100 can determine that the voice 1 is issued by user A; if the voiceprint information of the voice 1 is inconsistent with the voiceprint information of user A, the intelligent vehicle 100 can determine that the voice 1 is not issued by user A.

[0360] If the intelligent vehicle 100 determines that there is no main display screen, then the intelligent vehicle 100 can not process the voice 1.

[0361] If the intelligent vehicle 100 determines that there is a main display screen and the voice 1 is not issued by user A, then the intelligent vehicle 100 can also stop processing the voice 1.

[0362] If the intelligent vehicle 100 determines that there is a main display screen and the voice 1 is issued by user A, the intelligent vehicle 100 may perform the following step S1310.

[0363] In some embodiments, after step S1301, the intelligent vehicle 100 may perform step S1309. In this case, if the intelligent vehicle 100 determines that there is a main display screen and the voice 1 is issued by user A, the intelligent vehicle 100 may perform the following step S1310; if the intelligent vehicle 100 determines that there is no main display screen or determines that the voice 1 is not issued by user A, the intelligent vehicle 100 may perform the above step S1302.

[0364] S1310, the intelligent vehicle 100 performs the operation corresponding to the voice 1.

[0365] When the voice 1 belongs to a hands-free command, the intelligent vehicle 100 may perform the operation corresponding to the hands-free command. Exemplarily, if the voice 1 includes the hands-free command "turn on the air conditioner", the intelligent vehicle 100 may turn on the air conditioner; and again exemplarily, if the voice 1 includes the hands-free command "play music", the intelligent vehicle 100 may play music, etc.

[0366] When the voice 1 includes a wake-up word and other voice commands, the intelligent vehicle 100 may perform the voice command. Exemplarily, if the voice command is "play a song", the intelligent vehicle 100 may play a song, etc.

[0367] When the voice 1 includes a wake-up word and does not include other voice commands, the intelligent vehicle 100 may prompt the user to output a voice command through the display screen A or ask the user for the next voice command through voice. The operation of the intelligent vehicle 100 to prompt the user to output a voice command may be the operation corresponding to the voice 1 in this scenario.

[0368] When the voice 1 includes a non-hands-free command and the display screen A is the main display screen, the intelligent vehicle 100 may perform the voice command. Exemplarily, if the voice command is "go home", the intelligent vehicle 100 may turn on the navigation and set the destination to "home", etc.

[0369] In some embodiments, when performing the operation corresponding to the voice 1, the intelligent vehicle 100 may also interact with other electronic devices (such as the cloud server 400, the mobile phone 200, the earphone 300, etc.).

[0370] Exemplarily, if the voice 1 is "Play movie X", the intelligent vehicle 100 can obtain the video data of movie X through the communication connection with the cloud server 400, and play movie X through the display screen A. Also exemplarily, if the voice 1 is "Answer the phone", then the intelligent vehicle 100 can send an answer instruction to the mobile phone 200 through the communication connection with the mobile phone 200, instructing the mobile phone 200 to answer the phone. In this case, the intelligent vehicle 100 can also obtain the call data of the mobile phone 200 through the communication connection between the intelligent vehicle 100 and the mobile phone 200, and play the audio of the call through the speaker of the intelligent vehicle 100 based on the call data. Also exemplarily, if the voice 1 is "Play music through headphones", then the intelligent vehicle 100 can establish a communication connection with the headphones 300, and send the audio data to the headphones 300 through the communication connection with the headphones 300, and instruct the headphones 300 to play the audio based on the audio data.

[0371] S1311, the intelligent vehicle 100 outputs feedback information.

[0372] Step S1311 is an optional step.

[0373] While (or after) the intelligent vehicle 100 is performing the operation corresponding to the voice 1, the intelligent vehicle 100 can also output feedback information in any one or more of the ways such as display on the display screen, voice broadcast, indicator light flashing, vibration, etc. In another possible implementation, performing the operation corresponding to the voice 1 can also be regarded as a way of outputting feedback information, which is not limited in this application.

[0374] In some embodiments, when the intelligent vehicle 100 receives and recognizes the voice command (such as a hands-free wake-up command, wake-up word, non-hands-free wake-up command, etc.) included in the voice 1, it can output specified feedback information, such as "Received", "Okay", "Executing for you immediately".

[0375] In other embodiments, the intelligent vehicle 100 can also determine to output different feedback information based on the voice command in the voice 1 during or after performing the operation corresponding to the voice 1. The intelligent vehicle 100 can store the feedback information corresponding to the voice command. The voice command can include a hands-free wake-up command, a wake-up word, and a non-hands-free wake-up command. When receiving the voice command, the intelligent vehicle 100 can not only activate the voice assistant, but also output the feedback information corresponding to the voice command in a way such as display on the display screen and / or audio playback.

[0376] Exemplarily, Table 2 shows the corresponding relationship between the voice commands stored in the intelligent vehicle 100 provided in the embodiments of the present application and the feedback information.

[0377] Table 2

[0378] Voice command Feedback information Start navigation, go home Navigation has been started, destination: home Play music About to play music for you Open the window Received, will open the window for you right away Xiaoa You can talk at any time

[0379] As shown in Table 2, the intelligent vehicle 100 may store one or more voice commands, and may also store feedback information corresponding to the voice commands. For example, the feedback information corresponding to the voice command "Turn on the navigation and go home" may be "Navigation has been turned on, destination: home"; the feedback information corresponding to the voice command "Play music" may be "Music will be played for you soon"; the feedback information corresponding to the voice command "Open the window" may be "Received, will open the window for you immediately"; the feedback information corresponding to the wake word "Xiaoa" may be "You can talk at any time", etc.

[0380] It can be understood that the embodiments shown in Table 2 only exemplarily illustrate the corresponding relationship between the voice commands and the feedback information that the intelligent vehicle 100 can store. In the embodiments of the present application, the intelligent vehicle 100 may store more or fewer voice commands and feedback information than those in Table 2 above, and may also store a corresponding relationship different from that of the embodiments shown in Table 2. The present application does not make any limitations here.

[0381] By using the voice interaction method provided by the embodiments of the present application, the intelligent vehicle 100 can process voice commands of multiple users simultaneously and perform voice interaction with multiple users through one or more display screens. In addition, when the intelligent vehicle 100 processes multiple paths of voice, since there is at most one path of voice that is a non-wake-up command, and the remaining one or more paths of voice are wake-up-free commands, the intelligent vehicle 100 can ensure that each user can obtain timely feedback and improve the user experience.

[0382] It should be noted that the above Figure 13 shown embodiments only exemplarily illustrate that the intelligent vehicle 100 can perform different operations in different scenarios. In the embodiments of the present application, when the intelligent vehicle 100 receives the voice 1 of the user, it may also execute the voice interaction method in an execution order different from that of the above Figure 13 shown embodiments. For example, first determine whether there is a main display screen, then determine whether the voice 1 belongs to a wake-up-free word, then determine whether the voice 1 contains a wake-up word, and then determine the sound area a corresponding to the voice 1, or simultaneously determine whether there is a main display screen, whether the voice 1 belongs to a wake-up-free word, whether the voice 1 contains a wake-up word, and the sound area a corresponding to the voice 1, etc. The present application does not limit the specific execution order of each judgment step.

[0383] In some embodiments, when a voice dialogue identifier is displayed on display screen A, intelligent vehicle 100 may stop displaying the voice dialogue identifier when it detects that a closing condition is met. The closing condition may include, but is not limited to, any one or more of the following: receiving a wake word, the duration during which the user currently having a voice interaction with display screen A through a non-wake-free instruction does not utter a voice is greater than a preset duration, receiving a closing instruction, receiving a closing operation by the user, etc. Among them, the closing instruction may be a preset voice instruction, such as "end the conversation", etc. The closing operation by the user may be an operation by the user on a specified button, or an operation on a control in display screen A, etc.

[0384] In some embodiments, when a wake-free identifier is displayed on display screen A, intelligent vehicle 100 may stop displaying the wake-free identifier after the wake-free instruction that triggers the display of the wake-free identifier is executed.

[0385] In some embodiments, when a composite wake-free identifier is displayed on display screen A, intelligent vehicle 100 may change the digital indicator in the composite wake-free identifier after one or more of the multiple wake-free instructions being processed are executed.

[0386] In some embodiments, when a voice dialogue identifier and a voice sub-identifier are displayed on display screen A, intelligent vehicle 100 may stop displaying the voice sub-identifier after all the currently executed wake-free indicators are executed.

[0387] The following introduces the specific process of an intelligent vehicle 100 provided by an embodiment of the present application for executing the above step S1308.

[0388] As Figure 14 shown, the specific process of intelligent vehicle 100 for executing step S1308 may include the following steps:

[0389] S1401, intelligent vehicle 100 determines whether there is a main display screen, and the main display screen is the display screen that is currently processing non-wake-free instructions through a voice assistant.

[0390] For the specific content of step S1401, reference may be made to the relevant content in step S1304 as Figure 13 shown above, which will not be elaborated here.

[0391] If intelligent vehicle 100 determines that there is currently a main display screen, then intelligent vehicle 100 may execute the following step S1402.

[0392] If intelligent vehicle 100 determines that there is currently no main display screen, then intelligent vehicle 100 may execute the following step S1404.

[0393] S1402, the intelligent vehicle 100 determines whether the display screen A corresponding to sound zone 1 is the main display screen.

[0394] When there is a main display screen, the intelligent vehicle 100 can determine the display screen A corresponding to sound zone 1 based on the correspondence between the sound zone and the display screen, and determine whether the display screen A is the main display screen. Among them, the correspondence between the sound zone and the display screen can refer to the relevant content in the above step S1306, which will not be elaborated here.

[0395] If the intelligent vehicle 100 determines that the display screen A corresponding to sound zone 1 is the main display screen, the intelligent vehicle 100 can execute the following step S1403.

[0396] If the intelligent vehicle 100 determines that the display screen A corresponding to sound zone 1 is not the main display screen, the intelligent vehicle 100 can execute the following step S1404.

[0397] S1403, the intelligent vehicle 100 displays a voice dialogue identifier and a voice sub-identifier on the main display screen.

[0398] When there is a main display screen currently and the display screen A is the main display screen, the main display screen can process multiple voices simultaneously, and one of the voices is a non-wake-up instruction, and the other one or more voices are wake-up-free instructions.

[0399] In some embodiments, before receiving voice 1, the main display screen only processes one voice, and this voice is a non-wake-up instruction. In this case, a voice dialogue identifier can be displayed on the main display screen, and the voice dialogue identifier is used to prompt the user that the current display screen A is the main display screen and the main display screen is processing a non-wake-up instruction. When receiving voice 1, the intelligent vehicle 100 can split a voice sub-identifier from the voice dialogue identifier, and the voice sub-identifier is used to prompt the user that the current display screen A (i.e., the main display screen) is processing a non-wake-up instruction and a wake-up-free instruction.

[0400] Exemplarily, the voice dialogue identifier can be the voice dialogue identifier 501 shown above Figure 5C and the voice sub-identifier can be the voice sub-identifier 530 shown above Figure 5E The specific process of splitting the voice sub-identifier from the voice dialogue identifier can refer to the relevant description in the above Figures 5C to 5E shown embodiment.

[0401] In some other embodiments, before receiving Voice 1, the main display screen can process multiple voices, where one of the voices is a non-wake-up-free instruction, and the other one or more voices are wake-up-free instructions. In this case, the main display screen can display a voice conversation identifier and a voice sub-identifier. When receiving Voice 1, the intelligent vehicle 100 can keep displaying the voice conversation identifier and the voice sub-identifier. Optionally, in the above case, a number can also be displayed in the voice sub-identifier, and the number is used to prompt the user of the number of wake-up-free instructions being processed by the current display screen A.

[0402] S1404, the intelligent vehicle 100 determines whether the display screen A is processing other wake-up-free instructions.

[0403] In some embodiments, the intelligent vehicle 100 can determine whether the display screen A is processing other wake-up-free instructions based on whether the display screen A displays a wake-up-free identifier or conforms to a wake-up-free identifier.

[0404] In some other embodiments, the intelligent vehicle 100 can also determine whether the display screen A is processing other wake-up-free instructions based on whether the display screen A has the voice assistant turned on.

[0405] It can be understood that these two embodiments are just two examples here. In the embodiments of the present application, the intelligent vehicle 100 can also determine whether the display screen A is processing other wake-up-free instructions based on other methods.

[0406] If the intelligent vehicle 100 determines that the display screen A is processing other wake-up-free instructions, then the intelligent vehicle 100 can execute the following step S1405.

[0407] If the intelligent vehicle 100 determines that the display screen A is not currently processing other wake-up-free instructions, then the intelligent vehicle 100 can execute the following step S1406.

[0408] S1405, the intelligent vehicle 100 displays a composite wake-up-free identifier on the display screen A corresponding to Zone 1, and the composite wake-up-free identifier is used to prompt the user that the display screen A is processing multiple wake-up-free instructions.

[0409] In some embodiments, before receiving Voice 1, there is only one wake-up-free instruction being processed by the display screen A. In this case, a wake-up-free identifier can be displayed on the display screen A. When receiving Voice 1, the display screen A can respond to the wake-up-free instruction in Voice 1 and display a composite wake-up-free identifier on the display screen A. The composite wake-up-free identifier is used to prompt the user that the display screen A is processing multiple wake-up-free instructions.

[0410] Exemplarily, referring to the above Figures 8A to 8B illustrated embodiment, in the above embodiment, Voice 1 can be a wake-up-free instruction issued by User B, the display screen A can be the central control screen 10, and the wake-up-free identifier can be the aboveFigure 8B The composite no-wake-up identifier 810 shown.

[0411] In some other embodiments, before receiving Voice 1, there are multiple (two or more) no-wake-up instructions being processed by Display Screen A. In this case, a composite no-wake-up identifier may be displayed on Display Screen A. When Voice 1 is received, Display Screen A can update the digital indicator in the composite no-wake-up identifier in response to the no-wake-up instruction in Voice 1. The digital indicator can be used to prompt the user of the number of no-wake-up instructions currently being executed by Display Screen A. Exemplarily, the digital indicator can refer to the digital indicator 814 in the embodiment shown above Figure 8B shown in the embodiment.

[0412] S1406, the intelligent vehicle 100 displays a no-wake-up identifier on Display Screen A corresponding to Sound Zone 1.

[0413] Exemplarily, the no-wake-up identifier can be the no-wake-up identifier 601 shown above Figure 6C or the no-wake-up identifier 701 shown above Figure 7C or the no-wake-up identifier 711 shown above Figure 7D and so on.

[0414] By using the voice interaction method provided in the embodiments of the present application, the intelligent vehicle 100 can perform voice interaction with multiple users through one or more display screens, and supports simultaneous processing of multiple no-wake-up instructions, improving the voice interaction efficiency and being able to provide better voice interaction services for users.

[0415] In the embodiments of the present application, the voice assistants in different display screens can adopt the same working mode or different working modes. Moreover, the voice assistant of the same display screen can also adopt different working modes for voices initiated by different users.

[0416] In some embodiments, the voice assistant in the display screen can adopt different working modes for voices of different users based on the wake-up method of the voice assistant. For example, when the voice assistant is woken up by the user with a wake-up word, the voice assistant can adopt a full-duplex interaction working mode to perform voice interaction with the user who issued the wake-up word. When the voice assistant is woken up by a no-wake-up instruction, the voice assistant can adopt a single-round interaction working mode to perform voice interaction with the user who issued the no-wake-up instruction. By adopting the above method, at the same moment, the intelligent vehicle 100 can ensure that at most one voice adopts the full-duplex interaction working mode, and the other one or more voices all adopt the single-round interaction working mode. Since the single-round interaction requires lower power consumption than other working modes and has lower requirements for voice interaction, in the above situation, the intelligent vehicle 100 can ensure that the voice commands of multiple users can all be promptly feedback, providing better voice interaction services for users.

[0417] It can be understood that the above embodiments are only examples. In the embodiments of the present application, when the voice assistant is awakened by the user with a wake-up word, the voice assistant can also adopt other working modes (such as continuous listening, multi-round interaction, etc.) to conduct voice interaction with the user, and the present application does not make any limitations here.

[0418] The following introduces a specific application scenario of a voice interaction method provided by the embodiments of the present application.

[0419] Exemplarily, Figure 15A FIG. shows a schematic diagram of an application scenario provided by the embodiments of the present application.

[0420] As Figure 15A shown, inside the intelligent vehicle 100, the father is sitting in the driver's seat, the mother is sitting in the passenger seat, the daughter is sitting in the left position in the second row, the son is sitting in the right position in the second row, the grandfather is sitting in the left position in the third row, and the grandmother is sitting in the right position in the third row. Among them, the father is located in the driver's voice zone, the mother is located in the passenger voice zone, the daughter is located in the left voice zone of the second row, the son is located in the right voice zone of the second row, the grandfather is located in the left voice zone of the third row, and the grandmother is located in the right voice zone of the third row.

[0421] The correspondence between the display screen and the voice zone inside the intelligent vehicle 100 is as follows: The central control screen 10 corresponds to the following multiple voice zones: the driver's voice zone, the left voice zone of the second row, the left voice zone of the third row, and the right voice zone of the third row; the passenger screen 20 corresponds to the passenger voice zone; the right rear projection screen 30 corresponds to the right voice zone of the second row.

[0422] The intelligent vehicle 100 can receive the voice containing the wake-up word issued by the father from the driver's voice zone, such as "Xiaoa, navigate home", activate the voice assistant of the central control screen 10, determine the central control screen 10 as the main display screen, and display a voice dialogue identifier on the central control screen 10. For the specific interface, reference can be made to the above Figure 5B shown main interface 500 and voice dialogue identifier 501.

[0423] While the intelligent vehicle 100 receives the voice command of the father, the intelligent vehicle 100 can also receive the voice commands issued by other users, such as the mother's non-wake-up command "Turn on the air conditioner", the grandfather's non-wake-up command "Close the window", and the son's non-wake-up command "I want to watch cartoon XX".

[0424] According to the correspondence between the voice zone and the display screen, it can be known that the central control screen 10 can be responsible for processing the following voice commands: the voice command "Xiaoa, navigate home" issued by the father from the driver's voice zone and the non-wake-up command "Close the window" issued by the grandfather from the left voice zone of the third row. In the above situation, the central control screen 10 has to process one non-wake-up command and one non-wake-up command at the same time. At this time, the central control screen 10 can display asFigure 15B The main interface 1500 shown. At the same time, the intelligent vehicle 100 can also execute the voice command "close the window" sent by grandpa without waking up.

[0425] As Figure 15B shown, in the main interface 1500 of the central control screen 10, a voice dialogue identifier 1501 can be displayed, as well as a voice sub-identifier 1502 split from the voice dialogue identifier 1501. Optionally, a sound zone indicator 1503 can also be displayed. The voice dialogue identifier 1501 and the voice sub-identifier 1502 can be used to prompt the user that the central control screen 10 is processing multiple voice channels simultaneously, and one of the voice channels is a non-wake-up command. The voice dialogue identifier 1501 can also be used to prompt the user that the current main display screen is the central control screen 10. The specific content and display process of the main interface 1500, the voice dialogue identifier 1501, and the voice sub-identifier 1502 can refer to the relevant descriptions in the above Figures 5A to 5F shown embodiment, which will not be elaborated here. Moreover, the text 1504 displayed in the voice dialogue identifier 1501 can be a voice command sent by father, such as "Xiaoa, navigate home".

[0426] According to the relationship between the sound zone and the display screen, the co-pilot screen 20 can be responsible for processing the voice command "turn on the air conditioner" sent by mother without waking up. In the above case, the co-pilot screen 20 has to process a voice command without waking up. At this time, the co-pilot screen 20 can display the main interface 1510 as Figure 15C shown.

[0427] As Figure 15C shown, in the main interface 1510 of the co-pilot screen 20, a non-wake-up identifier 1511 can be displayed. The non-wake-up identifier 1511 can be used to prompt the user that the co-pilot screen 20 is receiving and processing a non-wake-up command. The specific content and display process of the main interface 1510 and the non-wake-up identifier 1511 can refer to the relevant descriptions in the above Figures 6B to 6D shown embodiment, which will not be elaborated here. Moreover, the text 1512 displayed in the non-wake-up identifier 1511 can be a non-wake-up command sent by mother, such as "turn on the air conditioner".

[0428] According to the relationship between the sound zone and the display screen, the right rear projection screen 30 can be responsible for processing the voice command "I want to watch cartoon XX" sent by son without waking up. In the above case, the right rear projection screen 30 has to process a voice command without waking up. At this time, the right rear projection screen 30 can display the main interface 1520 as Figure 15D shown.

[0429] As Figure 15DAs shown, the main interface 1520 of the right rear projection screen 30 may display a hands-free wake-up identifier 1521. The hands-free wake-up identifier 1521 can be used to prompt the user that the right rear projection screen 30 is receiving and processing a hands-free wake-up instruction. For the specific content and display process of the main interface 1520 and the hands-free wake-up identifier 1521, reference can be made to the relevant descriptions in the above Figures 6B to 6D illustrated embodiments, which will not be elaborated here. Moreover, the text 1512 displayed in the hands-free wake-up identifier 1511 can be feedback information based on the voice instruction issued by the mother, such as "I want to watch cartoon XX".

[0430] After recognizing the hands-free wake-up instruction "I want to watch cartoon XX" of the son and obtaining the video data of cartoon XX, the right rear projection screen 30 can stop displaying the hands-free wake-up identifier 1521 and display, based on the video data, Figure 15E as shown, the video playback interface 1530. The video playback interface 1530 can be used to play cartoon XX.

[0431] It can be understood that the above Figures 15A to 15E illustrated embodiments are only a scenario example, and the voice interaction method provided by the embodiments of the present application can also be applied to more application scenarios different from the above embodiments, which are not limited herein.

[0432] Next, a functional module of an intelligent vehicle 100 provided by the embodiments of the present application will be introduced.

[0433] Figure 16 The figure shows a schematic diagram of the functional modules of an intelligent vehicle 100 provided by the embodiments of the present application.

[0434] As Figure 16 shown, the intelligent vehicle 100 may include a voice receiving module 1601, a sound zone recognition module 1602, a voice recognition module 1603, a voice assistant module 1604, and an execution module 1605. Optionally, the intelligent vehicle 100 may further include any one or more of the following: a sound zone locking module 1606 and a voiceprint recognition module 1607.

[0435] Among them, the voice receiving module 1601 can collect the voices emitted by the users inside the intelligent vehicle 100, including but not limited to voice instructions such as wake-up words, hands-free wake-up instructions, and non-hands-free wake-up instructions. The voice receiving module 1601 may include one or more sub-modules, and each sub-module may correspond to a sound zone for collecting the voices in that sound zone. After receiving the voice of the user, the voice receiving module 1601 can send the collected voice to the voice recognition module 1603 and the sound zone recognition module 1602.

[0436] The sound zone recognition module 1602 can determine the sound zone corresponding to the voice based on the voice. The specific method for the sound zone recognition module 1602 to determine the sound zone corresponding to the voice can be referred to the aboveFigure 13 The relevant description in step S1302 shown above will not be elaborated here. After determining the voice zone, the voice zone recognition module 1602 can send the voice zone corresponding to the voice to the voice assistant module 1604.

[0437] The voice recognition module 1603 can recognize the text content of the voice. After receiving the voice sent by the voice receiving module 1601, the voice recognition module 1603 can recognize the text content of the voice, and make a judgment based on the text content of the voice to determine whether the voice contains a wake-up word, whether it belongs to a hands-free instruction, and whether it contains a non-hands-free instruction, etc.

[0438] When the voice recognition module 1603 determines that the user's voice contains a wake-up word, the voice recognition module 1603 can send instruction 1 to the voice assistant module 1604. Instruction 1 is used to instruct the voice assistant module 1604 to determine the main display screen and turn on the voice assistant function of the main display screen.

[0439] When the voice recognition module 1603 determines that the user's voice contains a hands-free instruction, the voice recognition module 1603 can send instruction 2 to the voice assistant module 1604. Instruction 2 can include the text content of the voice (or the hands-free instruction contained in the voice). Instruction 2 can be used to instruct the voice assistant module 1604 to process this hands-free instruction.

[0440] In some embodiments, the voice recognition module 1603 can also receive the judgment result sent by the voiceprint recognition module 1607. When it is determined that the user's voice contains a non-hands-free instruction, and based on the judgment result sent by the voiceprint recognition module 1607, it is determined that the user is a user who turns on the voice assistant function of the main display screen through the wake-up word, then the voice recognition module 1603 can send instruction 3 to the voice assistant module 1604. Instruction 3 is used to instruct the voice assistant module 1604 to process this non-hands-free instruction.

[0441] The voice assistant module 1604 can include one or more sub-modules. Each sub-module can correspond to a display screen in the intelligent vehicle 100, and each display screen can use the corresponding sub-module function to implement the voice assistant function of the display screen. When receiving instruction 1 sent by the voice recognition module 1603, the voice assistant module 1604 can determine the display screen corresponding to the voice zone based on the voice zone sent by the voice zone recognition module 1602, determine this display screen as the main display screen, and turn on the voice assistant function of this display screen through the sub-module corresponding to the voice assistant module 1604 on this display screen, and control the main display screen to display the voice conversation identifier. Optionally, corresponding feedback information can also be output, etc.

[0442] When the voice assistant module 1604 receives the instruction 2 sent by the speech recognition module 1603, the voice assistant module 1604 can determine the display screen corresponding to the sound zone based on the sound zone sent by the sound zone recognition module 1602, and through the sub-module corresponding to the voice assistant module 1604 on this display screen, activate the voice assistant function of this display screen, and control this display screen to display the no-wake-up logo. Optionally, corresponding feedback information can also be output, etc.

[0443] When the voice assistant module 1604 receives the instruction 3 sent by the speech recognition module 1603, the voice assistant module 1604 can process this non-no-wake-up instruction through the main display screen via the sub-module corresponding to the voice assistant module 1604. For example, control the main display screen to output the text content of this non-no-wake-up instruction, or output corresponding feedback information, etc.

[0444] In some embodiments, when the voice assistant module 1604 receives the text content of the voice instruction sent by the speech recognition module 1604, such as a no-wake-up instruction or a non-no-wake-up instruction, it can determine the operation to be performed based on the received text content of the voice instruction, and send an execution instruction to the execution module 1605. The execution instruction is used to instruct the execution module 1605 to perform a specified operation (such as closing the window, turning on the navigation, turning on the air conditioner, etc.).

[0445] The execution module 1605 can receive the execution instruction sent by the voice assistant module 1604 and perform the operation specified by this execution instruction (such as closing the window, turning on the navigation, turning on the air conditioner, etc.).

[0446] The sound zone locking module 1606 can assist the voice receiving module 1601 in performing voice collection work. When the voice receiving module 1601 collects the voice of the user in a sound zone through one of its sub-modules, the sound zone locking module 1606 can suppress the voice in non-this sound zone and enhance the voice in this sound zone at the same time. In this way, it can be ensured that the voice receiving module 1601 can respectively collect the voices in different sound zones.

[0447] The voiceprint recognition module 1607 can extract voiceprint information based on the voice, and judge whether two voices come from the same user based on the voiceprint information. In the embodiments of the present application, the voiceprint recognition module 1607 can receive the voice sent by the voice receiving module 1601, and judge whether this voice comes from the user who activates the voice assistant function of the main display screen through the wake-up word, and send the judgment result to the speech recognition module 1603.

[0448] It can be understood that the above Figure 16The illustrated embodiments are merely examples. In the embodiments of the present application, the intelligent vehicle 100 may further include more, fewer, or different functional modules than those in the above embodiments, and the present application does not make any limitations in this regard.

[0449] The following introduces the voice interaction method provided by the embodiments of the present application.

[0450] Figure 17 The flowchart of a voice interaction method provided by the embodiments of the present application is shown.

[0451] As Figure 17 shown, the specific process of the vehicle executing the voice interaction method may include the following steps:

[0452] S1701, the vehicle receives a first voice issued by a first user, and the first voice includes a wake-up word.

[0453] The vehicle includes a first display screen. The vehicle may be the intelligent vehicle 100 in the above embodiments.

[0454] Exemplarily, reference may be made to the above Figures 5A to 5F shown embodiments. In the above embodiments, the first display screen may be the central control screen 10, the first user may be user A, and the first voice may include the wake-up word issued by user A.

[0455] Again exemplarily, reference may be made to the above Figures 12A to 12B shown embodiments. In the above embodiments, the first display screen may be the central control screen 10, the first user may be user A, and the first voice may include the wake-up word issued by user A.

[0456] In a possible implementation manner, the first voice further includes a first instruction for instructing the vehicle to perform a first operation; the method further includes: in response to the first voice, performing the first operation. In this way, after receiving the first voice, the vehicle can perform the operation corresponding to the first instruction in the first voice, such as turning on the navigation, playing music, opening the window, etc.

[0457] In a possible implementation manner, after displaying a first conversation identifier on the first display screen, the method further includes: receiving a third voice issued by the first user for instructing the vehicle to perform a first operation; in response to the third voice, performing the first operation. In this way, the first user can also issue a voice command after activating the voice assistant function of the first display screen through the wake-up word. In this case, the first display screen has already started voice interaction with the first user in response to the wake-up word of the first user. Therefore, the first display screen can continue to receive the voice commands issued by the user and perform the operations corresponding to the voice commands, such as turning on the navigation, playing music, opening the window, etc.

[0458] S1702. The vehicle, in response to the first voice, displays a first conversation identifier on the first display screen, and the first conversation identifier is used to prompt the user that the first display screen is performing voice interaction with the first user.

[0459] Exemplarily, the first conversation identifier may be Figure 5B the voice conversation identifier 501 shown in the figure.

[0460] Also exemplarily, the first conversation identifier may be Figure 12A the voice conversation identifier 1201 shown in the figure.

[0461] In a possible implementation manner, the method further includes: after displaying the first conversation identifier on the first display screen, outputting a first feedback, where the first feedback is used to prompt the user that the vehicle has received the first voice; after displaying a voice sub-identifier on the first display screen, outputting a second feedback, where the second feedback is used to prompt the user that the vehicle has received the second voice.

[0462] The vehicle may output feedback information (such as the first feedback, the second feedback, etc.) in one or more ways such as display on the display screen, voice broadcast, indicator light flashing, vibration, etc. The output manners of the first feedback and the second feedback may be different, or the output manners of the first feedback and the second feedback may be the same. In this way, the user can be prompted that the vehicle has received the voice issued by the user by outputting the feedback information.

[0463] In a possible implementation manner, the first feedback is further used to prompt whether the operation indicated by the first voice has been executed, and the second feedback is further used to prompt whether the operation indicated by the second voice has been executed. In this way, the user can be prompted about the execution status of the voice command issued by the user by the feedback information, such as prompting the user that the execution is completed, etc.

[0464] For the specific content and output manner of the feedback information, reference may also be made to the relevant descriptions in the above Figure 6D and Figure 13 shown in step S1311, which will not be elaborated here.

[0465] S1703. The vehicle receives a second voice issued by a second user.

[0466] Exemplarily, reference may be made to the above Figures 5A to 5B shown embodiment, in the above embodiment, the second user may be user B.

[0467] S1704. When the second voice includes a hands-free wake-up instruction, the vehicle, in response to the second voice, displays a voice sub-identifier on the first display screen, and the voice sub-identifier is used to prompt the user that the first display screen is processing multiple voices.

[0468] Exemplarily, reference may be made to the above Figures 5A to 5FIn the illustrated embodiment, in the above embodiment, the second voice may be a hands-free instruction issued by User B, and the voice sub-identifier may be the above-mentioned Figure 5E shown voice sub-identifier 530.

[0469] In this way, while the vehicle is performing voice interaction with the first user on the first display screen, it can process the hands-free instruction of the second user through the first display screen, improving the voice interaction efficiency and providing better voice interaction services for the user.

[0470] In a possible implementation manner, the vehicle interior includes a first sound zone and a second sound zone. The first user is located in the first sound zone, and the second user is located in the second sound zone; the first display screen is used to respond to voices from the first sound zone and the second sound zone; displaying a first conversation identifier on the first display screen specifically includes: determining that the first voice comes from the first sound zone based on the first voice; displaying the first conversation identifier on the first display screen based on the first sound zone; displaying a voice sub-identifier on the first display screen specifically includes: determining that the second voice comes from the second sound zone based on the second voice; displaying the voice sub-identifier on the first display screen based on the second sound zone. In this way, when the vehicle receives the voice of the user, it can determine the display screen for displaying the voice identifier (such as the first conversation identifier, the voice sub-identifier, etc.) based on the corresponding relationship between the sound zone where the voice comes from and the display screen.

[0471] Exemplarily, referring to the illustrated embodiment shown in 3E above, if the first display screen is the central control screen 10, the first sound zone may be the driver's sound zone, and the second sound zone may be the left sound zone in the second row.

[0472] S1705, when the second voice includes a wake-up word, the vehicle stops displaying the first conversation identifier and displays a second conversation identifier on the first display screen. The second conversation identifier is used to prompt the user that the first display screen is performing voice interaction with the second user.

[0473] Exemplarily, reference may be made to the above Figures 12A to 12B illustrated embodiment. The first conversation identifier may be Figure 12A the shown voice conversation identifier 1201, and the second conversation identifier may be Figure 12B the shown voice conversation identifier 1211.

[0474] In this way, when receiving the wake-up word issued by the second user, the vehicle can end the voice interaction between the first display screen and the first user and perform voice interaction with the second user through the first display screen. It can ensure that at the same moment, the vehicle only needs to process one voice interaction initiated by the wake-up word.

[0475] In a possible implementation, the interior of the vehicle includes a first sound zone and a second sound zone. A first user is located in the first sound zone, and a second user is located in the second sound zone. The first display screen is configured to respond to voices from the first sound zone and the second sound zone. Displaying a first conversation identifier on the first display screen specifically includes: determining that the first voice comes from the first sound zone based on the first voice; displaying the first conversation identifier on the first display screen based on the first sound zone. Displaying a second conversation identifier on the first display screen specifically includes: determining that the second voice comes from the second sound zone based on the second voice; displaying the second conversation identifier on the first display screen based on the second sound zone. In this way, when the vehicle receives the voice of a user, it can determine the display screen for displaying the voice identifier (such as the first conversation identifier, the second conversation identifier, etc.) based on the correspondence between the sound zone where the voice comes from and the display screen.

[0476] Exemplarily, referring to the embodiment shown in 3E above, if the first display screen is the central control screen 10, the first sound zone may be the driver's sound zone, and the second sound zone may be the left sound zone in the third row.

[0477] In a possible implementation, the method further includes: when the second voice includes a wake-up word, outputting an interruption prompt, where the interruption prompt is used to prompt that the voice interaction between the first user and the first display screen is interrupted. In a possible implementation, the vehicle can output the interruption prompt in one or more ways such as display on the display screen, flashing of the indicator light, voice broadcast, vibration, etc. In this way, the first user can be prompted that the voice interaction with the first display screen has ended through the interruption prompt. Exemplarily, the interruption prompt may be the interruption prompt 1103 shown above. Figure 11C as shown.

[0478] In a possible implementation, the method further includes: in response to the first voice, displaying a first indicator on the first display screen, where the first indicator is used to indicate the position of the first sound zone relative to the first display screen; when the second voice includes a hands-free wake-up instruction, in response to the second voice, replacing the first indicator with a second indicator, where the second indicator is used to indicate the positions of the first sound zone and the second sound zone relative to the display screen; when the second voice includes a wake-up word, in response to the second voice, replacing the first indicator with a third indicator, where the third indicator is used to indicate the position of the second sound zone relative to the display screen.

[0479] In this way, the user can be prompted of the source sound zone of the voice currently processed by the display screen through the sound zone indicator (such as the first indicator, the second indicator, the third indicator, etc.).

[0480] Exemplarily, the first indicator may be the sound zone indicator 505 shown above. Figure 5B The second indicator may be the sound zone indicator 511 shown above. Also exemplarily, the first indicator may be the one shown above. Figure 5D as shown. Figure 12AThe shown pitch range indicator 1202, the third indicator can be the above-mentioned Figure 12B shown pitch range indicator 1212.

[0481] In a possible implementation, the voice interaction mode between the first display screen and the first user is full-duplex interaction; when the second voice includes a hands-free wake-up instruction, the voice interaction mode between the first display screen and the second user is single-round interaction; when the second voice includes a wake-up word, the voice interaction mode between the first display screen and the second user is full-duplex interaction.

[0482] In this way, it can be determined that the voice interaction mode enabled by the wake-up word is full-duplex interaction, and the voice interaction mode enabled by the hands-free wake-up instruction is single-round interaction. In this way, it can ensure that timely voice interaction services are provided for more users.

[0483] Figure 18 The flowchart shows another voice interaction method provided by the embodiments of the present application.

[0484] As Figure 18 shown, the specific process of the vehicle executing the voice interaction method may include the following steps:

[0485] S1801, the vehicle receives the fourth voice issued by the first user, and the fourth voice includes a hands-free wake-up instruction.

[0486] The vehicle can be the intelligent vehicle 100 in the above-mentioned embodiment.

[0487] Exemplarily, reference can be made to the above-mentioned Figures 8A to 8C shown embodiment. In the above-mentioned embodiment, the first user can be user A, and the fourth voice can be the hands-free wake-up instruction issued by user A.

[0488] Another example is that reference can be made to the above-mentioned Figures 10A to 10C shown embodiment. In the above-mentioned embodiment, the first user can be user A, and the fourth voice can be the hands-free wake-up instruction issued by user A.

[0489] S1802, the vehicle responds to the fourth voice and displays a first hands-free wake-up identifier on the first display screen, and the first hands-free wake-up identifier is used to prompt the user that the first display screen is processing the hands-free wake-up instruction.

[0490] Exemplarily, the first display screen can be the central control screen 10 in the above-mentioned Figures 8A to 8C shown embodiment, or it can also be the central control screen 10 in the above-mentioned Figures 10A to 10C shown embodiment. The first hands-free wake-up identifier can be the hands-free wake-up identifier 801 in the above-mentioned Figure 8A shown, or it can also be the hands-free wake-up identifier 1011 in the above-mentioned Figure 10A shown.

[0491] S1803, The vehicle receives the fifth voice sent by the second user.

[0492] Exemplarily, reference may be made to the above Figures 8A to 8C , or Figures 10A to 10C illustrated embodiment. In the above embodiment, the second user may be User B.

[0493] S1804, When the fifth voice includes a hands-free wake-up instruction, the vehicle responds to the fifth voice, stops displaying the first hands-free wake-up identifier, and displays a composite hands-free wake-up identifier on the first display screen. The composite hands-free wake-up identifier is used to prompt the user that the first display screen is processing multiple hands-free wake-up instructions.

[0494] Exemplarily, the composite hands-free wake-up identifier may be the above Figure 8B illustrated composite hands-free wake-up identifier 810.

[0495] In this way, the vehicle can process hands-free wake-up instructions sent by multiple users through the first display screen, improving the voice interaction efficiency and providing better voice interaction services for users.

[0496] In a possible implementation, the vehicle interior includes a first sound zone and a second sound zone. The first user is located in the first sound zone, and the second user is located in the second sound zone; the first display screen is used to respond to voices from the first sound zone and the second sound zone; displaying the first hands-free wake-up identifier on the first display screen specifically includes: determining that the fourth voice comes from the first sound zone based on the fourth voice; displaying the first hands-free wake-up identifier on the first display screen based on the first sound zone; displaying the composite hands-free wake-up identifier on the first display screen specifically includes: determining that the fifth voice comes from the second sound zone based on the fifth voice; displaying the composite hands-free wake-up identifier on the first display screen based on the second sound zone. In this way, when the vehicle receives a hands-free wake-up instruction from a user, it can determine the display screen for displaying the hands-free wake-up identifier based on the corresponding relationship between the sound zone where the voice comes from and the display screen.

[0497] S1805, When the fifth voice includes a wake-up word, the vehicle responds to the fifth voice, stops displaying the first hands-free wake-up identifier, and displays a third conversation identifier and a voice sub-identifier on the first display screen. The third conversation identifier is used to prompt the user that the first display screen is having a voice interaction with the second user, and the voice sub-identifier is used to prompt the user that the first display screen is processing multiple voices.

[0498] Exemplarily, reference may be made to the above Figures 10A to 10C illustrated embodiment. In the above embodiment, the first hands-free wake-up identifier may be Figure 10A illustrated hands-free wake-up identifier 1011, the third conversation identifier may be Figure 10C illustrated voice conversation identifier 1021, and the voice sub-identifier may be Figure 10C illustrated voice sub-identifier 1022.

[0499] In this way, while the vehicle processes the hands-free wake-up instruction issued by the first user through the first display screen, it can receive the wake-up word issued by the second user and conduct voice interaction with the second user through the first display screen.

[0500] In a possible implementation, the interior of the vehicle includes a first sound zone and a second sound zone. The first user is located in the first sound zone, and the second user is located in the second sound zone. The first display screen is used to respond to voices from the first sound zone and the second sound zone. Displaying a first hands-free wake-up identifier on the first display screen specifically includes: determining that the fourth voice comes from the first sound zone based on the fourth voice; displaying the first hands-free wake-up identifier on the first display screen based on the first sound zone. Displaying a third conversation identifier and a voice sub-identifier on the first display screen specifically includes: determining that the fifth voice comes from the second sound zone based on the fifth voice; displaying the third conversation identifier and the voice sub-identifier on the first display screen based on the second sound zone. In this way, when the vehicle receives the hands-free wake-up instruction or the wake-up word from the user, it can determine the display screen for displaying the hands-free wake-up identifier based on the corresponding relationship between the sound zone where the voice source is located and the display screen.

[0501] In a possible implementation, the method further includes: when the fifth voice includes a wake-up word, in response to the fifth voice, displaying a fourth indicator on the first display screen, where the fourth indicator is used to indicate the positions of the first sound zone and the second sound zone relative to the display screen. In this way, the fourth indicator can be displayed on the first display screen when the wake-up word is received.

[0502] Exemplarily, the fourth indicator may be the sound zone indicator 1023 shown above Figure 10B as shown.

[0503] In a possible implementation, the method further includes: in response to the fourth voice, displaying a fifth indicator on the first display screen, where the fifth indicator is used to indicate the position of the first sound zone relative to the first display screen; when the fifth voice includes a hands-free wake-up instruction, in response to the fifth voice, replacing the fifth indicator with a sixth indicator, where the sixth indicator is used to indicate the positions of the first sound zone and the second sound zone relative to the display screen; when the fifth voice includes a wake-up word, in response to the fifth voice, replacing the fifth indicator with a seventh indicator, where the seventh indicator is used to indicate the positions of the first sound zone and the second sound zone relative to the display screen.

[0504] In this way, the user can be prompted about the source sound zone of the voice currently processed by the display screen through sound zone indicators (such as the first indicator, the second indicator, the fourth indicator, etc.).

[0505] Exemplarily, the fifth indicator may be the sound zone indicator 802 shown above Figure 8A as shown, the sixth indicator may be the sound zone indicator 805 shown above Figure 8B as shown, and the seventh indicator may be the sound zone indicator 805 shown above Figure 10BThe shown pitch indicator 1023.

[0506] In a possible implementation, the voice interaction mode between the first display screen and the first user is single-round interaction; when the fifth voice includes a wake-up-free instruction, the voice interaction mode between the first display screen and the second user is single-round interaction; when the fifth voice includes a wake-up word, the voice interaction mode between the first display screen and the second user is full-duplex interaction.

[0507] In this way, it can be determined that the voice interaction mode enabled by the wake-up word is full-duplex interaction, and the voice interaction mode enabled by the wake-up-free instruction is single-round interaction. In this way, it can ensure that timely voice interaction services are provided to more users.

[0508] Figure 19 The flowchart of another voice interaction method provided by an embodiment of the present application is shown.

[0509] As Figure 19 shown, the specific process of the vehicle executing the voice interaction method may include the following steps:

[0510] S1901, the vehicle receives the sixth voice from the first user in the first pitch area, the sixth voice includes a wake-up word, the vehicle includes a first display screen and a second display screen, and the vehicle interior includes a first pitch area and a second pitch area; the first display screen is used to respond to the voice from the first pitch area, and the second display screen is used to respond to the voice from the second pitch area.

[0511] The vehicle may be the intelligent vehicle 100 in the above embodiment.

[0512] Exemplarily, reference may be made to the above Figures 6A to 6D shown embodiment. In the above embodiment, the first display screen may be the central control screen 10, the second display screen may be the co-pilot screen 20. The first user may be user A, the first pitch area may be pitch area a (such as the driver's pitch area, the left pitch area in the second row, etc.), the second user may be user B, and the second pitch area may be pitch area b (such as the co-pilot pitch area). The sixth voice may include the wake-up word issued by user A.

[0513] Again exemplarily, reference may be made to the above Figures 11A to 11D shown embodiment. In the above embodiment, the first display screen may be the central control screen 10, the second display screen may be the right rear projection screen 30. The first user may be user A, the first pitch area may be pitch area a (such as the driver's pitch area, the left pitch area in the second row, etc.), the second user may be user B, and the second pitch area may be pitch area b (such as the right pitch area in the second row).

[0514] S1902, the vehicle responds to the sixth voice and determines that the sixth voice comes from the first pitch area based on the sixth voice.

[0515] S1903, The vehicle displays a third conversation identifier on the first display screen based on the first sound zone. The third conversation identifier is used to prompt the user that the first display screen is performing voice interaction with the first user.

[0516] Exemplarily, the third conversation identifier may be the voice conversation identifier 501 shown above Figure 6A or the voice conversation identifier 1101 shown above. Figure 11A shown above.

[0517] S1904, The vehicle receives the seventh voice sent by the second user from the second sound zone.

[0518] S1905, When the seventh voice includes a wake-up-free instruction, the vehicle responds to the seventh voice and determines that the seventh voice comes from the second sound zone based on the seventh voice.

[0519] S1906, The vehicle displays a second wake-up-free identifier on the second display screen based on the second sound zone. The second wake-up-free identifier is used to prompt the user that the second display screen is processing the wake-up-free instruction.

[0520] Exemplarily, the second wake-up-free identifier may be the wake-up-free identifier 601 shown above Figure 6C shown above.

[0521] In this way, voice interaction can be performed with different users through different display screens based on the corresponding relationship between the sound zone and the display screen, improving the voice interaction efficiency and providing better voice interaction services for users.

[0522] S1907, When the seventh voice includes a wake-up word, the vehicle responds to the seventh voice and stops displaying the third conversation identifier on the first display screen.

[0523] Exemplarily, the third conversation identifier may be the voice conversation identifier 1101 shown above Figure 11A shown above.

[0524] S1908, The vehicle determines that the seventh voice comes from the second sound zone based on the seventh voice.

[0525] S1909, The vehicle displays a fourth conversation identifier on the second display screen based on the second sound zone. The fourth conversation identifier is used to prompt the user that the second display screen is performing voice interaction with the second user.

[0526] Exemplarily, the fourth conversation identifier may be the voice conversation identifier 1111 shown above Figure 11D shown above.

[0527] In this way, when the wake-up word sent by the second user is received, the vehicle can end the voice interaction between the first display screen and the first user, and conduct voice interaction with the second user through the second display screen. It can be ensured that at the same moment, the vehicle only needs to process one voice interaction activated by the wake-up word.

[0528] In a possible implementation manner, the voice interaction mode between the first display screen and the first user is full-duplex interaction; when the seventh voice includes a hands-free wake-up instruction, the voice interaction mode between the second display screen and the second user is single-round interaction; when the seventh voice includes a wake-up word, the voice interaction mode between the second display screen and the second user is full-duplex interaction.

[0529] In this way, it can be determined that the voice interaction mode activated by the wake-up word is full-duplex interaction, and the voice interaction mode activated by the hands-free wake-up instruction is single-round interaction. In this way, it can be ensured to provide timely voice interaction services for more users.

[0530] Figure 20 The flowchart of another voice interaction method provided by an embodiment of the present application is shown.

[0531] As Figure 20 shown, the specific process for the vehicle to execute the voice interaction method may include the following steps:

[0532] S2001. The vehicle receives the eighth voice sent by the first user from the first sound area. The eighth voice includes a hands-free wake-up instruction. The vehicle includes a first display screen and a second display screen, and the interior of the vehicle includes a first sound area and a second sound area; the first display screen is used to respond to the voice from the first sound area, and the second display screen is used to respond to the voice from the second sound area.

[0533] The vehicle may be the intelligent vehicle 100 in the above embodiment.

[0534] Exemplarily, reference may be made to the above Figures 7A to 7D shown embodiment. In the above embodiment, the first display screen may be the central control screen 10, and the second display screen may be the right rear projection screen 30. The first user may be user A, the first sound area may be sound area a (such as the driver's sound area, the left sound area in the second row, etc.), the second user may be user B, and the second sound area may be sound area b (such as the right sound area in the second row). The eighth voice may include the hands-free wake-up instruction sent by user A.

[0535] Another exemplarily, reference may be made to the above Figures 9A to 9D shown embodiment. In the above embodiment, the first display screen may be the right rear projection screen 30, and the second display screen may be the central control screen 10. The first user may be user A, the first sound area may be sound area a (such as the right sound area in the second row, etc.), the second user may be user B, and the second sound area may be sound area b (such as the driver's sound area).

[0536] In S2002, in response to the eighth voice, the vehicle determines that the eighth voice comes from the first voice zone based on the eighth voice.

[0537] In S2003, the vehicle displays a third no-wake-up identifier on the first display screen based on the first voice zone. The third no-wake-up identifier is used to prompt the user that the first display screen is processing a no-wake-up instruction.

[0538] Exemplarily, the third no-wake-up identifier may be the Figure 7C no-wake-up identifier 701 shown above.

[0539] In S2004, the vehicle receives a ninth voice emitted by a second user from the second voice zone.

[0540] In S2005, when the ninth voice includes a no-wake-up instruction, the vehicle responds to the ninth voice and determines that the ninth voice comes from the second voice zone based on the ninth voice.

[0541] In S2006, the vehicle displays a fourth no-wake-up identifier on the second display screen based on the second voice zone. The fourth no-wake-up identifier is used to prompt the user that the second display screen is processing a no-wake-up instruction.

[0542] Exemplarily, the fourth no-wake-up identifier may be Figure 7D the no-wake-up identifier 711 shown.

[0543] In this way, based on the correspondence between the voice zone and the display screen, different display screens can process no-wake-up instructions of different users, improving the voice interaction efficiency and providing better voice interaction services for users.

[0544] In S2007, when the ninth voice includes a wake-up word, the vehicle responds to the ninth voice and determines that the ninth voice comes from the second voice zone based on the ninth voice.

[0545] In S2008, the vehicle displays a fifth conversation identifier on the second display screen based on the second voice zone. The fifth conversation identifier is used to prompt the user that the second display screen is having a voice interaction with the second user.

[0546] Exemplarily, the fifth conversation identifier may be the Figure 9C voice conversation identifier 911 shown above.

[0547] In this way, when receiving a wake-up word issued by the second user, the vehicle can, while processing the no-wake-up instruction of the first user through the first display screen, have a voice interaction with the second user through the second display screen.

[0548] In a possible implementation, the voice interaction mode between the first display screen and the first user is single-round interaction; when the ninth voice includes a wake-up-free instruction, the voice interaction mode between the second display screen and the second user is single-round interaction; when the ninth voice includes a wake-up word, the voice interaction mode between the second display screen and the second user is full-duplex interaction.

[0549] In this way, it can be determined that the voice interaction mode enabled by the wake-up word is full-duplex interaction, and the voice interaction mode enabled by the wake-up-free instruction is single-round interaction. In this way, it can be ensured that timely voice interaction services are provided to more users.

[0550] The various embodiments of the present application can be combined arbitrarily to achieve different technical effects.

[0551] In the above embodiments, they can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the present application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state disk (SSD)), etc.

[0552] Those of ordinary skill in the art can understand all or part of the processes in the methods of the above embodiments. These processes can be completed by relevant hardware instructed by a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The foregoing storage medium includes: various media such as ROM or random access memory RAM, magnetic disk, or optical disc that can store program codes.

[0553] In summary, the above description is only an embodiment of the technical solution of the present invention and is not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made according to the disclosure of the present invention shall be included within the protection scope of the present invention.

Claims

1. A voice interaction method, characterized in that, it is applied to a vehicle, and the vehicle includes a first display screen; the method includes: receiving a first voice issued by a first user, the first voice including a wake-up word; in response to the first voice, displaying a first dialogue identifier on the first display screen, the first dialogue identifier being used to prompt the user that the first display screen is performing voice interaction with the first user; receiving a second voice issued by a second user; when the second voice includes a wake-up-free instruction, in response to the second voice, displaying a voice sub-identifier on the first display screen, the voice sub-identifier being used to prompt the user that the first display screen is processing multiple voices.

2. The method according to claim 1, characterized in that, the first voice further includes a first instruction, the first instruction being used to instruct the vehicle to perform the first operation; the method further includes: in response to the first voice, performing the first operation.

3. The method according to claim 1, characterized in that, after displaying the first dialogue identifier on the first display screen, the method further includes: receiving a third voice issued by the first user, the third voice being used to instruct the vehicle to perform a first operation; in response to the third voice, performing the first operation.

4. The method according to any one of claims 1-3, characterized in that, the method further includes: after displaying the first dialogue identifier on the first display screen, outputting a first feedback, the first feedback being used to prompt the user that the vehicle has received the first voice; after displaying the voice sub-identifier on the first display screen, outputting a second feedback, the second feedback being used to prompt the user that the vehicle has received the second voice.

5. The method according to any one of claims 1-4, characterized in that, the method further includes: when the second voice includes the wake-up word, stopping displaying the first dialogue identifier and displaying a second dialogue identifier on the first display screen, the second dialogue identifier being used to prompt the user that the first display screen is performing voice interaction with the second user.

6. The method according to claim 5, characterized in that, the method further includes: when the second voice includes the wake-up word, outputting an interruption prompt, the interruption prompt being used to prompt the first user that the voice interaction with the first display screen is interrupted.

7. The method according to any one of claims 1-6, characterized in that, the interior of the vehicle includes a first sound zone and a second sound zone, the first user is located in the first sound zone, and the second user is located in the second sound zone; the first display screen is used to respond to voices from the first sound zone and the second sound zone; the displaying the first dialogue identifier on the first display screen specifically includes: determining that the first voice comes from the first sound zone based on the first voice; displaying the first dialogue identifier on the first display screen based on the first sound zone; the displaying the voice sub-identifier on the first display screen specifically includes: determining that the second voice comes from the second sound zone based on the second voice; Based on the second voice zone, display a voice sub-identifier on the first display screen.

8. The method according to claim 5 or 6, wherein, the interior of the vehicle includes a first voice zone and a second voice zone, the first user is located in the first voice zone, and the second user is located in the second voice zone; the first display screen is used to respond to voices from the first voice zone and the second voice zone; The displaying of the first conversation identifier on the first display screen specifically includes: Determine that the first voice comes from the first voice zone based on the first voice; Based on the first voice zone, display the first conversation identifier on the first display screen; The displaying of the second conversation identifier on the first display screen specifically includes: Determine that the second voice comes from the second voice zone based on the second voice; Based on the second voice zone, display the second conversation identifier on the first display screen.

9. The method according to claim 7 or 8, wherein, the method further includes: In response to the first voice, display a first indicator on the first display screen, and the first indicator is used to indicate the position of the first voice zone relative to the first display screen; When the second voice includes a hands-free wake-up instruction, in response to the second voice, replace the first indicator with a second indicator, and the second indicator is used to indicate the positions of the first voice zone and the second voice zone relative to the display screen; When the second voice includes the wake-up word, in response to the second voice, replace the first indicator with a third indicator, and the third indicator is used to indicate the position of the second voice zone relative to the display screen.

10. The method according to any one of claims 1-9, wherein, the voice interaction mode between the first display screen and the first user is full-duplex interaction; when the second voice includes a hands-free wake-up instruction, the voice interaction mode between the first display screen and the second user is single-round interaction; when the second voice includes the wake-up word, the voice interaction mode between the first display screen and the second user is full-duplex interaction.

11. A voice interaction method, wherein, applied to a vehicle, the vehicle includes a first display screen; the method includes: Receive a fourth voice emitted by a first user, and the fourth voice includes a hands-free wake-up instruction; In response to the fourth voice, and display a first hands-free wake-up identifier on the first display screen, and the first hands-free wake-up identifier is used to prompt the user that the first display screen is processing a hands-free wake-up instruction; Receive a fifth voice emitted by a second user; When the fifth voice includes a hands-free wake-up instruction, in response to the fifth voice, stop displaying the first hands-free wake-up identifier, and display a composite hands-free wake-up identifier on the first display screen, and the composite hands-free wake-up identifier is used to prompt the user that the first display screen is processing multiple hands-free wake-up instructions.

12. The method according to claim 11, wherein, the method further includes: When the fifth voice includes a wake-up word, in response to the fifth voice, stop displaying the first hands-free wake-up identifier, and display a third conversation identifier and a voice sub-identifier on the first display screen. The third conversation identifier is used to prompt the user that the first display screen is performing voice interaction with the second user, and the voice sub-identifier is used to prompt the user that the first display screen is processing multiple voices.

13. The method according to claim 11 or 12, wherein, the interior of the vehicle includes a first sound zone and a second sound zone, the first user is located in the first sound zone, and the second user is located in the second sound zone; the first display screen is used to respond to voices from the first sound zone and the second sound zone; The step of displaying the first hands-free wake-up identifier on the first display screen specifically includes: Determine that the fourth voice comes from the first sound zone based on the fourth voice; Based on the first sound zone, display the first hands-free wake-up identifier on the first display screen; The step of displaying the composite hands-free wake-up identifier on the first display screen specifically includes: Determine that the fifth voice comes from the second sound zone based on the fifth voice; Based on the second sound zone, display the composite hands-free wake-up identifier on the first display screen.

14. The method according to claim 12, wherein, the interior of the vehicle includes a first sound zone and a second sound zone, the first user is located in the first sound zone, and the second user is located in the second sound zone; the first display screen is used to respond to voices from the first sound zone and the second sound zone; The step of displaying the first hands-free wake-up identifier on the first display screen specifically includes: Determine that the fourth voice comes from the first sound zone based on the fourth voice; Based on the first sound zone, display the first hands-free wake-up identifier on the first display screen; The step of displaying the third conversation identifier and the voice sub-identifier on the first display screen specifically includes: Determine that the fifth voice comes from the second sound zone based on the fifth voice; Based on the second sound zone, display the third conversation identifier and the voice sub-identifier on the first display screen.

15. The method according to claim 13 or 14, wherein, the method further includes: When the fifth voice includes the wake-up word, in response to the fifth voice, display a fourth indicator on the first display screen, and the fourth indicator is used to indicate the positions of the first sound zone and the second sound zone relative to the display screen.

16. A voice interaction method, wherein, applied to a vehicle, the vehicle includes a first display screen and a second display screen, and the interior of the vehicle includes a first sound zone and a second sound zone; The first display screen is used to respond to voices from the first sound zone, and the second display screen is used to respond to voices from the second sound zone; the method includes: Receive a sixth voice emitted by a first user from the first sound zone, and the sixth voice includes a wake-up word; In response to the sixth voice, determine that the sixth voice comes from the first sound zone based on the sixth voice; Based on the first voice zone, display a third conversation identifier on the first display screen, where the third conversation identifier is used to prompt the user that the first display screen is performing voice interaction with the first user; Receive a seventh voice sent by a second user from the second voice zone; When the seventh voice includes a hands-free wake-up instruction, in response to the seventh voice, determine that the seventh voice comes from the second voice zone based on the seventh voice; Based on the second voice zone, display a second hands-free wake-up identifier on the second display screen, where the second hands-free wake-up identifier is used to prompt the user that the second display screen is processing a hands-free wake-up instruction.

17. The method according to claim 16, wherein, the method further includes: When the seventh voice includes the wake-up word, in response to the seventh voice, stop displaying the third conversation identifier on the first display screen; Determine that the seventh voice comes from the second voice zone based on the seventh voice; Based on the second voice zone, display a fourth conversation identifier on the second display screen, where the fourth conversation identifier is used to prompt the user that the second display screen is performing voice interaction with the second user.

18. A voice interaction method, wherein, it is applied to a vehicle, the vehicle includes a first display screen and a second display screen, and the interior of the vehicle includes a first voice zone and a second voice zone; The first display screen is used to respond to voices from the first voice zone, and the second display screen is used to respond to voices from the second voice zone; the method includes: Receive an eighth voice sent by a first user from the first voice zone, where the eighth voice includes a hands-free wake-up instruction; In response to the eighth voice, determine that the eighth voice comes from the first voice zone based on the eighth voice; Based on the first voice zone, display a third hands-free wake-up identifier on the first display screen, where the third hands-free wake-up identifier is used to prompt the user that the first display screen is processing a hands-free wake-up instruction; Receive a ninth voice sent by a second user from the second voice zone; When the ninth voice includes a hands-free wake-up instruction, in response to the ninth voice, determine that the ninth voice comes from the second voice zone based on the ninth voice; Based on the second voice zone, display a fourth hands-free wake-up identifier on the second display screen, where the fourth hands-free wake-up identifier is used to prompt the user that the second display screen is processing a hands-free wake-up instruction.

19. The method according to claim 18, wherein, the method further includes: When the ninth voice includes the wake-up word, in response to the ninth voice, determine that the ninth voice comes from the second voice zone based on the ninth voice; Based on the second voice zone, display a fifth conversation identifier on the second display screen, where the fifth conversation identifier is used to prompt the user that the second display screen is performing voice interaction with the second user.

20. A vehicle, wherein, includes: One or more processors, one or more memories; the one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer program code, the computer program code includes computer instructions, and when the one or more processors execute the computer instructions, the vehicle is caused to execute the method according to any one of claims 1-19 above.

21. A computer-readable storage medium, comprising computer instructions, characterized in that when the computer instructions run on a vehicle, the vehicle is caused to execute the method according to any one of claims 1-19 above.

Citation Information

Cited By

  • Voice interaction method and device, and storage medium

    CN119517020A