Speech interaction method and system, and related apparatus

By displaying dialogue logos and sub-logoes on the display screen of smart cars, the problem that smart cars cannot handle multiple user voice commands at the same time is solved, achieving more efficient voice interaction and better user experience.

WO2025113522A1PCT designated stage expired Publication Date: 2025-06-05HUAWEI TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/135052
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-30
Filing Date
2024-11-27
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

When multiple users in the car issue different voice commands at the same time, smart cars cannot process multiple voice commands at the same time, affecting the user's user experience.

Method used

Multiple voice interaction is realized by displaying dialogue logos and sub-identities on the display screen, allowing the vehicle to process voice commands of multiple users at the same time.

Benefits of technology

It improves the efficiency of voice interaction, provides a better user experience, and can handle voice commands of multiple users at the same time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024135052_05062025_PF_FP_ABST
    Figure CN2024135052_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application is applied to a vehicle. Disclosed are a speech interaction method and system, and a related apparatus, wherein the vehicle comprises a first display screen. The method comprises: receiving first speech that is sent by a first user, wherein the first speech comprises a wake-up word; in response to the first speech, displaying a first dialogue identifier on a first display screen, wherein the first dialogue identifier is used for prompting a user that speech interaction is being performed between the first display screen and the first user; receiving second speech that is sent by a second user; and when the second speech comprises a wake-up-free instruction, in response to the second speech, displaying a speech sub-identifier on the first display screen, wherein the speech sub-identifier is used for prompting the user that the first display screen is processing multiple paths of speech. In this way, a vehicle can simultaneously process speech instructions that are sent by a plurality of users, thereby providing a better speech interaction service for users.
Need to check novelty before this filing date? Find Prior Art

Description

A voice interaction method, system and related device

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 30, 2023, with application number 202311655712.2 and application name “A Voice Interaction Method, System and Related Devices”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of electronic technology, and in particular to a voice interaction method, system and related devices. Background Art

[0003] With the continuous development of electronic technology, the functions of smart cars are becoming more diverse. To provide better services to users, more and more smart cars are equipped with voice assistant functions. When the user is inside the smart car, the smart car can receive and respond to the user's voice commands and perform the functions indicated by the voice commands, such as playing audio, starting navigation, etc.

[0004] However, if there are multiple users in the car and multiple users issue different voice commands at the same time, the smart car cannot process multiple voice commands at the same time, which will affect the user experience. Summary of the Invention

[0005] The present application provides a voice interaction method, system and related devices, which can realize the simultaneous processing of different voice commands of multiple users and provide users with better voice interaction services.

[0006] In a first aspect, the present application provides a voice interaction method, which is applied to a vehicle, the vehicle including a first display screen; the method comprising:

[0007] Receive a first voice from a first user, where the first voice includes a wake-up word; in response to the first voice, display a first dialogue identifier on a first display screen, where the first dialogue identifier is used to prompt the user that the first display screen is conducting voice interaction with the first user; receive a second voice from a second user; when the second voice includes a wake-up-free instruction, display a voice sub-identifier on the first display screen in response to the second voice, where the voice sub-identifier is used to prompt the user that the first display screen is processing multiple voices.

[0008] In this way, the vehicle can process the second user's wake-up instructions through the first display screen while conducting voice interaction with the first user on the first display screen, thereby improving the efficiency of voice interaction and providing users with better voice interaction services.

[0009] In one possible implementation, the first voice also includes a first instruction, and the first instruction is used to instruct the vehicle to perform a first operation; the method also includes: performing the first operation in response to the first voice.

[0010] In this way, after receiving the first voice, the vehicle can execute the operation corresponding to the first instruction in the first voice, such as opening navigation, playing music, opening the window, etc.

[0011] In one possible implementation, after the first dialogue identifier is displayed on the first display screen, the method further includes: receiving a third voice issued by the first user, the third voice being used to instruct the vehicle to perform the first operation; and performing the first operation in response to the third voice.

[0012] In this way, the first user can also issue a voice command after turning on the voice assistant function of the first display screen by using the wake-up word. In this case, the first display screen has already started voice interaction with the first user in response to the first user's wake-up word. Therefore, the first display screen can continue to receive the voice command issued by the user and perform the operation corresponding to the voice command, such as opening navigation, playing music, opening the car window, etc.

[0013] In one possible implementation, the method further includes: after displaying the first dialogue identifier on the first display screen, outputting a first feedback, the first feedback being used to prompt the user that the vehicle has received the first voice; after displaying the voice sub-identifier on the first display screen, outputting a second feedback, the second feedback being used to prompt the user that the vehicle has received the second voice.

[0014] The vehicle may output feedback information (e.g., first feedback, second feedback, etc.) in one or more ways, such as display screen display, voice broadcast, indicator light flashing, vibration, etc. The first feedback and the second feedback may be output in different ways, or they may be output in the same way.

[0015] In this way, the user can be prompted by outputting feedback information that the vehicle has received the voice issued by the user.

[0016] In a possible implementation, the first feedback is further used to prompt the user whether the operation indicated by the first voice is executed, and the second feedback is further used to prompt the user whether the operation indicated by the second voice is executed.

[0017] In this way, the user can be prompted through feedback information on the execution status of the vehicle's voice command issued by the user, such as prompting the user that the execution is completed, etc.

[0018] In one possible implementation, the method further includes: when the second voice includes a wake-up word, stopping displaying the first dialogue identifier and displaying a second dialogue identifier on the first display screen, wherein the second dialogue identifier is used to prompt the user that the first display screen is performing voice interaction with the second user.

[0019] In this way, when the vehicle receives the wake-up word from the second user, it can end the voice interaction between the first display and the first user and continue the voice interaction with the second user through the first display. This ensures that at the same time, the vehicle only needs to process the voice interaction initiated by the wake-up word on one route.

[0020] In one possible implementation, the method further includes: when the second voice includes a wake-up word, outputting an interruption prompt, where the interruption prompt is used to prompt that the voice interaction between the first user and the first display screen is interrupted.

[0021] In one possible implementation, the vehicle may output the interruption prompt in one or more ways, such as display screen display, flashing indicator light, voice broadcast, vibration, etc.

[0022] In this way, the interruption prompt can be used to indicate that the voice interaction between the first user and the first display screen has ended.

[0023] In one possible implementation, the interior of a vehicle includes a first sound zone and a second sound zone, a first user is located in the first sound zone, and a second user is located in the second sound zone; a first display screen is used to respond to voices from the first sound zone and the second sound zone; a first dialogue identifier is displayed on the first display screen, specifically including: determining that the first voice comes from the first sound zone based on the first voice; displaying the first dialogue identifier on the first display screen based on the first sound zone; displaying a voice sub-identifier on the first display screen, specifically including: determining that the second voice comes from the second sound zone based on the second voice; and displaying the voice sub-identifier on the first display screen based on the second sound zone.

[0024] In this way, when the vehicle receives the user's voice, it can determine the display screen for displaying the voice identifier (such as the first dialogue identifier, voice sub-identifier, etc.) based on the correspondence between the sound zone of the voice source and the display screen.

[0025] In one possible implementation, the interior of a vehicle includes a first sound zone and a second sound zone, a first user is located in the first sound zone, and a second user is located in the second sound zone; the first display screen is used to respond to voices from the first sound zone and the second sound zone; a first dialogue identifier is displayed on the first display screen, specifically including: determining that the first voice comes from the first sound zone based on the first voice; displaying the first dialogue identifier on the first display screen based on the first sound zone; displaying the second dialogue identifier on the first display screen, specifically including: determining that the second voice comes from the second sound zone based on the second voice; displaying the second dialogue identifier on the first display screen based on the second sound zone.

[0026] In this way, when the vehicle receives the user's voice, it can determine the display screen used to display the voice identifier (such as the first dialogue identifier, the second dialogue identifier, etc.) based on the correspondence between the sound zone of the voice source and the display screen.

[0027] In one possible implementation, the method also includes: in response to the first voice, displaying a first indicator on the first display screen, the first indicator being used to indicate the position of the first sound zone relative to the first display screen; when the second voice includes a wake-up-free instruction, in response to the second voice, replacing the first indicator with a second indicator, the second indicator being used to indicate the positions of the first sound zone and the second sound zone relative to the display screen; when the second voice includes a wake-up word, in response to the second voice, replacing the first indicator with a third indicator, the third indicator being used to indicate the position of the second sound zone relative to the display screen.

[0028] In this way, the user can be prompted with the sound zone indicator (eg, the first indicator, the second indicator, the third indicator, etc.) of the source sound zone of the speech currently being processed on the display screen.

[0029] In one possible implementation, the voice interaction mode between the first display and the first user is full-duplex interaction; when the second voice includes a wake-up-free command, the voice interaction mode between the first display and the second user is single-round interaction; when the second voice includes a wake-up word, the voice interaction mode between the first display and the second user is full-duplex interaction.

[0030] This ensures that voice interaction initiated by a wake-up word is full-duplex, while voice interaction initiated without a wake-up command is single-round. This ensures that timely voice interaction services are provided to more users.

[0031] In the second aspect, the present application provides a voice interaction method, which is applied to a vehicle, and the vehicle includes a first display screen; the method includes: receiving a fourth voice issued by a first user, the fourth voice including a wake-up-free instruction; responding to the fourth voice, and displaying a first wake-up-free indicator on the first display screen, the first wake-up-free indicator is used to prompt the user that the first display screen is processing the wake-up-free instruction; receiving a fifth voice issued by a second user; when the fifth voice includes a wake-up-free instruction, responding to the fifth voice, stopping displaying the first wake-up-free indicator, and displaying a composite wake-up-free indicator on the first display screen, the composite wake-up-free indicator is used to prompt the user that the first display screen is processing multiple wake-up-free instructions.

[0032] In this way, the vehicle can process wake-up-free commands issued by multiple users through the first display screen, thereby improving the efficiency of voice interaction and providing users with better voice interaction services.

[0033] In one possible implementation, the method also includes: when the fifth voice includes a wake-up word, in response to the fifth voice, stopping displaying the first wake-up-free identifier, and displaying a third dialogue identifier and a voice sub-identifier on the first display screen, the third dialogue identifier is used to prompt the user that the first display screen is conducting voice interaction with the second user, and the voice sub-identifier is used to prompt the user that the first display screen is processing multiple voices.

[0034] In this way, the vehicle can process the wake-up command issued by the first user through the first display screen, receive the wake-up word issued by the second user, and conduct voice interaction with the second user through the first display screen.

[0035] In one possible implementation, the interior of the vehicle includes a first sound zone and a second sound zone, the first user is located in the first sound zone, and the second user is located in the second sound zone; the first display screen is used to respond to voices from the first sound zone and the second sound zone; a first wake-up-free mark is displayed on the first display screen, specifically including: determining that the fourth voice comes from the first sound zone based on the fourth voice; displaying the first wake-up-free mark on the first display screen based on the first sound zone; displaying a composite wake-up-free mark on the first display screen, specifically including: determining that the fifth voice comes from the second sound zone based on the fifth voice; and displaying the composite wake-up-free mark on the first display screen based on the second sound zone.

[0036] In this way, when the vehicle receives the user's wake-up-free command, it can determine the display screen used to display the wake-up-free logo based on the correspondence between the sound zone of the voice source and the display screen.

[0037] In one possible implementation, the interior of the vehicle includes a first sound zone and a second sound zone, the first user is located in the first sound zone, and the second user is located in the second sound zone; the first display screen is used to respond to voices from the first sound zone and the second sound zone; the first wake-up-free indicator is displayed on the first display screen, specifically including: determining that the fourth voice comes from the first sound zone based on the fourth voice; displaying the first wake-up-free indicator on the first display screen based on the first sound zone; displaying the third dialogue indicator and voice sub-identifier on the first display screen, specifically including: determining that the fifth voice comes from the second sound zone based on the fifth voice; and displaying the third dialogue indicator and voice sub-identifier on the first display screen based on the second sound zone.

[0038] In this way, when the vehicle receives the user's wake-up-free command or wake-up word, it can determine the display screen used to display the wake-up-free logo based on the correspondence between the sound zone of the voice source and the display screen.

[0039] In one possible implementation, the method further includes: when the fifth voice includes a wake-up word, in response to the fifth voice, displaying a fourth indicator on the first display screen, the fourth indicator being used to indicate positions of the first sound zone and the second sound zone relative to the display screen.

[0040] In this way, the fourth indicator can be displayed on the first display screen when the wake-up word is received.

[0041] In one possible implementation, the method also includes: in response to the fourth voice, displaying a first indicator on the first display screen, the first indicator being used to indicate the position of the first sound zone relative to the first display screen; when the fifth voice includes a wake-up-free instruction, in response to the fifth voice, replacing the first indicator with a second indicator, the second indicator being used to indicate the positions of the first sound zone and the second sound zone relative to the display screen; when the fifth voice includes a wake-up word, in response to the fifth voice, replacing the first indicator with a fourth indicator, the fourth indicator being used to indicate the positions of the first sound zone and the second sound zone relative to the display screen.

[0042] In this way, the user can be prompted with the sound zone indicator (eg, the first indicator, the second indicator, the fourth indicator, etc.) of the source sound zone of the speech currently being processed on the display screen.

[0043] In one possible implementation, the voice interaction mode between the first display and the first user is single-round interaction; when the fifth voice includes a wake-up-free command, the voice interaction mode between the first display and the second user is single-round interaction; when the fifth voice includes a wake-up word, the voice interaction mode between the first display and the second user is full-duplex interaction.

[0044] This ensures that voice interaction initiated by a wake-up word is full-duplex, while voice interaction initiated without a wake-up command is single-round. This ensures that timely voice interaction services are provided to more users.

[0045] In a third aspect, the present application provides a voice interaction method, which is applied to a vehicle, the vehicle including a first display screen and a second display screen, and the interior of the vehicle including a first sound zone and a second sound zone; the first display screen is used to respond to voice from the first sound zone, and the second display screen is used to respond to voice from the second sound zone; the method includes: receiving a sixth voice emitted by a first user from the first sound zone, the sixth voice including a wake-up word; in response to the sixth voice, determining that the sixth voice comes from the first sound zone based on the sixth voice; displaying a third dialogue identifier on the first display screen based on the first sound zone, the third dialogue identifier being used to prompt the user that the first display screen is conducting voice interaction with the first user; receiving a seventh voice emitted by a second user from the second sound zone; when the seventh voice includes a wake-up-free instruction, in response to the seventh voice, determining that the seventh voice comes from the second sound zone based on the seventh voice; displaying a second wake-up-free identifier on the second display screen based on the second sound zone, the second wake-up-free identifier being used to prompt the user that the second display screen is processing the wake-up-free instruction.

[0046] In this way, based on the correspondence between the sound zones and the display screens, voice interaction can be performed with different users through different display screens, thereby improving the efficiency of voice interaction and providing users with better voice interaction services.

[0047] In one possible implementation, the method also includes: when the seventh voice includes a wake-up word, in response to the seventh voice, stopping displaying the third dialogue identifier in the first display screen; determining based on the seventh voice that the seventh voice comes from the second sound zone; and displaying a fourth dialogue identifier on the second display screen based on the second sound zone, the fourth dialogue identifier being used to prompt the user that the second display screen is conducting voice interaction with the second user.

[0048] In this way, when the vehicle receives the wake-up word from the second user, it can end the voice interaction between the first display and the first user and continue the voice interaction with the second user through the second display. This ensures that at the same time, the vehicle only needs to process the voice interaction initiated by the wake-up word on one route.

[0049] In one possible implementation, the voice interaction mode between the first display and the first user is full-duplex interaction; when the seventh voice includes a wake-up-free command, the voice interaction mode between the second display and the second user is single-round interaction; when the seventh voice includes a wake-up word, the voice interaction mode between the second display and the second user is full-duplex interaction.

[0050] This ensures that voice interaction initiated by a wake-up word is full-duplex, while voice interaction initiated without a wake-up command is single-round. This ensures that timely voice interaction services are provided to more users.

[0051] In a fourth aspect, the present application provides a voice interaction method, which is applied to a vehicle, the vehicle including a first display screen and a second display screen, and the interior of the vehicle including a first sound zone and a second sound zone; the first display screen is used to respond to voice from the first sound zone, and the second display screen is used to respond to voice from the second sound zone; the method includes: receiving an eighth voice emitted by a first user from the first sound zone, the eighth voice including a wake-up-free instruction; in response to the eighth voice, determining based on the eighth voice that the eighth voice comes from the first sound zone; displaying a third wake-up-free indicator on the first display screen based on the first sound zone, the third wake-up-free indicator being used to prompt the user that the first display screen is processing the wake-up-free instruction; receiving a ninth voice emitted by a second user from the second sound zone; when the ninth voice includes a wake-up-free instruction, in response to the ninth voice, determining based on the ninth voice that the ninth voice comes from the second sound zone; displaying a fourth wake-up-free indicator on the second display screen based on the second sound zone, the fourth wake-up-free indicator being used to prompt the user that the second display screen is processing the wake-up-free instruction.

[0052] In this way, based on the correspondence between the sound zones and the display screens, the wake-up-free commands of different users can be processed through different display screens, thereby improving the efficiency of voice interaction and providing users with better voice interaction services.

[0053] In one possible implementation, the method also includes: when the ninth voice includes a wake-up word, in response to the ninth voice, determining that the ninth voice comes from the second sound zone based on the ninth voice; and displaying a fifth dialogue identifier on the second display screen based on the second sound zone, the fifth dialogue identifier being used to prompt the user that the second display screen is conducting voice interaction with the second user.

[0054] In this way, when the vehicle receives the wake-up word issued by the second user, it can process the first user's wake-up command through the first display screen while performing voice interaction with the second user through the second display screen.

[0055] In one possible implementation, the voice interaction mode between the first display and the first user is single-round interaction; when the ninth voice includes a wake-up-free command, the voice interaction mode between the second display and the second user is single-round interaction; when the ninth voice includes a wake-up word, the voice interaction mode between the second display and the second user is full-duplex interaction.

[0056] This ensures that voice interaction initiated by a wake-up word is full-duplex, while voice interaction initiated without a wake-up command is single-round. This ensures that timely voice interaction services are provided to more users.

[0057] In a fifth aspect, the present application provides a vehicle comprising one or more processors and one or more memories. The one or more memories are coupled to the one or more processors, and the one or more memories are configured to store computer program code, the computer program code comprising computer instructions. When the one or more processors execute the computer instructions, the vehicle executes the voice interaction method of any possible implementation of any of the above aspects.

[0058] In a sixth aspect, an embodiment of the present application provides a computer storage medium comprising computer instructions, which, when executed on a vehicle, enables the vehicle to execute the voice interaction method in any possible implementation of any of the above aspects.

[0059] In the seventh aspect, an embodiment of the present application provides a computer program product, which, when running on a vehicle, enables the vehicle to execute the voice interaction method in any possible implementation of any of the above aspects.

[0060] The beneficial effects of the fifth to seventh aspects can refer to the beneficial effects of the first to fourth aspects mentioned above. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figures 1A-1D are schematic diagrams of four voice interaction modes provided in an embodiment of the present application;

[0062] FIG2 is a schematic diagram of the system architecture of a voice interaction system 1000 provided in an embodiment of the present application;

[0063] 3A-3B are schematic diagrams of the interior configurations of two smart cars 100 provided in embodiments of the present application;

[0064] FIG3C-FIG3D are schematic diagrams of the sound zone distribution inside two smart cars 100 provided in embodiments of the present application;

[0065] 3E-3H are schematic diagrams showing the correspondence between various sound zones and display screens provided in an embodiment of the present application;

[0066] FIG4A is a schematic diagram of the hardware structure of a smart car 100 provided in an embodiment of the present application;

[0067] FIG4B is a schematic diagram of the hardware structure of a mobile phone 200 provided in an embodiment of the present application;

[0068] 5A-5F are schematic diagrams of interfaces for a group of central control screens 10 to simultaneously process multiple voice channels, provided in an embodiment of the present application;

[0069] 6A-6D are schematic diagrams of interfaces for a central control screen 10 and a passenger screen 20 to simultaneously process different voices, respectively, according to an embodiment of the present application;

[0070] 7A-7D are schematic diagrams of interfaces in which a central control screen 10 and a right rear projection screen 30 respectively process different voices simultaneously, according to an embodiment of the present application;

[0071] 8A-8C are schematic diagrams of interfaces for a group of central control screens 10 to simultaneously process multiple wake-up-free commands according to an embodiment of the present application;

[0072] 9A-9D are schematic diagrams of interfaces in which a central control screen 10 and a right rear projection screen 30 respectively process different voices simultaneously, according to an embodiment of the present application;

[0073] 10A-10C are schematic diagrams of interfaces for a central control screen 10 to simultaneously process multiple voice channels, provided in an embodiment of the present application;

[0074] Figures 11A to 11D are schematic diagrams of interfaces in which a central control screen 10 and a right rear projection screen 30 sequentially process wake-up words issued by different users, provided in an embodiment of the present application;

[0075] 12A-12B are schematic diagrams of interfaces of a group of central control screens 10 provided in an embodiment of the present application for successively processing wake-up words issued by different users;

[0076] FIG13 is a flow chart of a voice interaction method according to an embodiment of the present application;

[0077] FIG14 is a schematic diagram of a process after a smart car 100 receives a wake-up-free instruction according to an embodiment of the present application;

[0078] FIG15A is a schematic diagram of an application scenario of a voice interaction method provided in an embodiment of the present application;

[0079] 15B-15E are schematic diagrams of interfaces of different display screens in a set of application scenarios provided by an embodiment of the present application;

[0080] FIG16 is a schematic diagram of functional modules of a smart car 100 provided in an embodiment of the present application;

[0081] FIG17 is a flow chart of a voice interaction method according to an embodiment of the present application;

[0082] FIG18 is a flow chart of another voice interaction method provided in an embodiment of the present application;

[0083] FIG19 is a flow chart of another voice interaction method provided in an embodiment of the present application;

[0084] Figure 20 is a flow chart of another voice interaction method provided in an embodiment of the present application. DETAILED DESCRIPTION

[0085] The following is a clear and detailed description of the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in the text is only a description of the association relationship between related objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.

[0086] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of the technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of this application, unless otherwise specified, "plurality" means two or more.

[0087] The term "user interface (UI)" in the following embodiments of this application refers to a medium interface for interaction and information exchange between an application or operating system and a user, which realizes the conversion between the internal form of information and the form acceptable to the user. The user interface is a source code written in a specific computer language such as Java and extensible markup language (XML). The interface source code is parsed and rendered on an electronic device and finally presented as content that the user can recognize. The commonly used form of user interface is graphical user interface (GUI), which refers to a user interface related to computer operations that is displayed in a graphical manner. It can be a visual interface element such as text, icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, widgets, etc. displayed on the display screen of an electronic device.

[0088] The following introduces some terms involved in the embodiments of this application.

[0089] Wake-up word: A wake-up word is a pre-set word used to trigger the electronic device to turn on (also called wake up) the voice assistant function (hereinafter referred to as voice assistant). Electronic devices with voice assistants can be set with one or more wake-up words. When the electronic device detects that the user says the preset wake-up word, the electronic device can turn on the voice assistant and start voice interaction with the user. In this way, it is possible to better distinguish between the user's daily chat scenes and voice interaction scenes, and it can also avoid the power consumption caused by the long-term activation of the voice assistant.

[0090] Wake-up-free command: A wake-up-free command is a pre-set voice command that can trigger an electronic device to perform a preset operation. When the electronic device detects that the voice spoken by the user includes a wake-up-free command, the electronic device can turn on the voice assistant and perform the operation corresponding to the wake-up-free command (such as opening the car window, starting navigation, playing audio, etc.). It should be noted that the difference between the wake-up-free command and the wake-up word is that the wake-up word is only used to turn on the voice assistant. After turning on the voice assistant, the electronic device needs to determine the operation to be performed based on the voice command spoken by the user later; while the wake-up-free command can trigger the electronic device to turn on the voice assistant and perform the specified operation when the voice assistant is not turned on, and after the electronic device performs the operation corresponding to the wake-up-free command, the electronic device usually turns off the voice assistant.

[0091] The following introduces the working modes of various voice assistants involved in the embodiments of this application (also known as voice interaction modes).

[0092] Single-turn interaction: A single-turn interaction refers to the process of waking up the voice assistant and executing a complete voice interaction process before shutting down the voice assistant's working mode. A complete voice interaction process can include one input and one output, and the input and output cannot be performed at the same time. That is, at the same time, the electronic device with the voice assistant turned on can only perform input or output. Among them, input refers to receiving the voice from the user, and output refers to the voice feedback for the user's voice input. In the scenario of single-turn interaction, the voice assistant needs to be woken up before each voice interaction, and then voice interaction can be carried out.

[0093] Multi-turn interaction: Multi-turn interaction means that after the voice assistant is awakened, it can perform multiple complete voice interaction processes without having to wake the voice assistant again during these multiple voice interaction processes. It should be noted that in the multi-turn interaction scenario, input and output cannot be performed simultaneously.

[0094] Keep listening: After the voice assistant is activated, it continuously listens to the user's voice and can provide one or more voice feedback based on the user's voice. It should be noted that in the keep listening scenario, the electronic device with the voice assistant enabled can perform both input and output at the same time, or only perform input or output.

[0095] Full-duplex interaction (full duplex): Full-duplex interaction, also referred to as full-duplex, means that after the voice assistant is activated, it continuously listens to the user's voice and continuously outputs voice feedback based on the user's voice. Full-duplex is a term in the communications field that defines a real-time, two-way voice information interaction mode. It should be noted that in a full-duplex interaction scenario, the electronic device with the voice assistant enabled can conduct real-time, two-way voice interaction with the user.

[0096] It should be noted that, in some embodiments, the voice assistant can adopt a different working mode for each voice channel.

[0097] FIG1A shows a schematic diagram of a single-round interaction working mode provided in an embodiment of the present application.

[0098] As shown in Figure 1A, in a one-dimensional coordinate system, the coordinate axis can represent time, and the closer to the positive direction of the coordinate axis (i.e., the direction of the arrow), the larger the time value. Three color blocks of different colors are displayed in the one-dimensional coordinate system, black blocks, gray blocks, and white blocks. Among them, the black block can represent that the electronic device is in the input stage, the white block can represent that the electronic device is in the output stage, and the gray block can represent that the electronic device is in the wake-up stage. According to Figure 1A, from time t0 to time t1, the voice assistant of the electronic device is awakened; from time t1 to time t2, the electronic device receives the voice input by the user through the voice assistant; from time t2 to time t3, the electronic device outputs voice through the voice assistant, and the output voice is generated based on the voice input by the user. After the output is completed, the voice assistant of the electronic device is turned off. From time t3 to time t4, the voice assistant of the electronic device is always in the off state. From time t4 to time t5, the electronic device's voice assistant is reactivated; from time t5 to time t6, the electronic device receives user input via the voice assistant; from time t6 to time t7, the electronic device outputs voice via the voice assistant, and the output voice is generated based on the user input. After this output is completed, the electronic device's voice assistant is turned off again.

[0099] It is understandable that the embodiment shown in Figure 1A is only an example. In the embodiment of the present application, the duration of the wake-up phase, input phase and output phase of a single round of interaction may also be different from the duration of the above embodiment, and the present application does not limit it here.

[0100] FIG1B shows a schematic diagram of a multi-round interaction working mode provided in an embodiment of the present application.

[0101] As shown in FIG1B , in a one-dimensional coordinate system, the coordinate axis can represent time, and the closer to the positive direction of the coordinate axis (i.e., the direction of the arrow), the larger the time value. The one-dimensional coordinate system displays three color blocks of different colors, black blocks, gray blocks, and white blocks. The stages represented by each color block can refer to the relevant description in the embodiment shown in FIG1A above. According to FIG1B , from time t10 to time t11, the voice assistant of the electronic device is awakened; from time t11 to time t12, the electronic device receives the voice input by the user through the voice assistant; from time t12 to time t13, the electronic device outputs voice through the voice assistant, and the output voice is generated based on the voice input by the user; from time t13 to time t14, the electronic device receives the voice input by the user through the voice assistant; from time t14 to time t15, the electronic device outputs voice through the voice assistant, and the output voice is generated based on the voice input by the user between time t13 and time t14.

[0102] It can be understood that the embodiment shown in Figure 1B is only an example. In the embodiment of the present application, the duration of the wake-up phase, input phase and output phase of the multi-round interaction may also be different from the duration of the above embodiment, and the multi-round interaction may also include more or fewer voice interaction processes than the above embodiment. This application does not limit this.

[0103] FIG1C shows a schematic diagram of a continuous monitoring working mode provided in an embodiment of the present application.

[0104] As shown in Figure 1C, in a one-dimensional coordinate system, the coordinate axis can represent time, and the closer to the positive direction of the coordinate axis (i.e., the direction of the arrow), the larger the time value. The one-dimensional coordinate system displays three color blocks of different colors: black, gray, and white. The stages represented by each color block can refer to the relevant description of the embodiment shown in Figure 1A above. According to Figure 1C, from time t20 to time t21, the voice assistant of the electronic device is awakened; from time t21 to time t25, the electronic device continuously listens to the user's voice input through the voice assistant. The user's voice input can be continuous or intermittent. From time t22 to time t23, the electronic device outputs voice through the voice assistant, and the output voice is generated based on the voice input by the user between time t21 and time t23. Afterwards, from time t24 to time t26, the electronic device can also output voice through the voice assistant, and the voice output this time is generated based on the voice input by the user between time t23 and time t25, where time t25 is earlier than time t26.

[0105] It can be understood that the embodiment shown in Figure 1C is only an example. In the embodiment of the present application, the duration of the wake-up phase, input phase and output phase of continuous monitoring may also be different from the above embodiment, and continuous monitoring may also include more or fewer output phases than the above embodiment. This application does not limit this.

[0106] FIG1D shows a schematic diagram of a full-duplex interaction working mode provided in an embodiment of the present application.

[0107] As shown in Figure 1D, in a one-dimensional coordinate system, the coordinate axis can represent time, and the closer to the positive direction of the coordinate axis (i.e., the direction of the arrow), the larger the time value. Three color blocks of different colors are displayed in the one-dimensional coordinate system, black blocks, gray blocks, and white blocks. The stages represented by each color block can refer to the relevant description in the embodiment shown in Figure 1A above. According to Figure 1D, from time t30 to time t31, the voice assistant of the electronic device is awakened; from time t31 to time t32, the electronic device continuously monitors the voice input by the user through the voice assistant. The voice input by the user can be continuous or intermittent. In addition, from time t31 to time t32, the electronic device can also perform real-time voice output through the voice assistant, that is, after generating a voice useful for output based on the voice input by the user, the electronic device can output the voice in real time.

[0108] It can be understood that the embodiment shown in Figure 1D is only an example. In the embodiment of the present application, the duration of the full-duplex wake-up phase, input phase and output phase may also be different from the above embodiment, and the present application does not limit it here.

[0109] The following introduces a voice interaction system 1000 provided in an embodiment of the present application.

[0110] FIG2 shows a schematic diagram of the system architecture of a voice interaction system 1000 provided in an embodiment of the present application.

[0111] As shown in FIG2 , the voice interaction system 1000 may include a smart car 100 . The voice interaction system 1000 may also include any one or more of the following: a mobile phone 200 , a headset 300 , and a cloud server 400 .

[0112] In some embodiments, a communication connection may be established between the smart car 100 and the mobile phone 200. Based on the communication connection between the smart car 100 and the mobile phone 200, the smart car 100 may receive and respond to the user's voice commands and make / receive calls through the mobile phone 200.

[0113] In some embodiments, a communication connection may be established between the smart car 100 and the headset 300. The smart car 100 may receive and respond to a user's voice command, send specified audio data to the headset 300 via the communication connection between the smart car 100 and the headset 300, and play the audio through the headset 300.

[0114] In other embodiments, the smart car 100 and the mobile phone 200 may establish a communication connection, the smart car 100 and the headset 300 may establish a communication connection, and the mobile phone 200 and the headset 300 may also establish a communication connection. The smart car 100 can receive and respond to the user's voice commands, make / receive calls through the mobile phone 200, and play call audio through the headset 300.

[0115] In some embodiments, the smart car 100 can establish a communication connection with the cloud server 400. When the smart car 100 performs the operation corresponding to the voice command issued by the user, the smart car 100 can send a request to the cloud server 400. The request can be used to request the cloud server 400 to send specified data (such as audio data, video data, image data, web page data, etc.) to the smart car 100. In other embodiments, if the smart car 100 cannot accurately recognize the voice issued by the user, the smart car 100 can also upload the collected voice to the cloud server 400, and the cloud server 400 will analyze and recognize the voice and return the recognition result to the smart car 100. The recognition result may include the text content of the voice.

[0116] It is understandable that the voice interaction system 1000 shown in Figure 2 is only an example. In the embodiments of the present application, the voice interaction system 1000 may also include more or fewer electronic devices such as mobile phones and headphones than in the above embodiments, and may also include wearable devices such as watches and bracelets. This application does not limit this.

[0117] The following introduces the internal structure of a smart car 100 provided in an embodiment of the present application.

[0118] FIG3A shows a schematic diagram of the internal structure of a smart car 100 provided in an embodiment of the present application.

[0119] As shown in FIG3A , the interior of the smart car 100 may be provided with a plurality of seats and may also be provided with one or more display screens. For example, the interior of the smart car 100 may include three rows of seats, the first row of seats may include a main driver's seat and a front passenger seat, the second row of seats may include a second row of left seats and a second row of right seats, and the third row of seats may include three rows of left seats and three rows of right seats. The one or more display screens inside the smart car 100 may include a central control screen 10, a front passenger screen 20, and a right rear projection screen 30. The central control screen 10 may be provided between the main driver's seat and the front passenger seat, the front passenger screen 20 may be provided in front of the front passenger seat, and the right rear projection screen 30 may be provided directly in front of the second row of right seats, for example, on the backrest of the front passenger seat.

[0120] In some embodiments, each display screen may be equipped with a voice assistant for voice interaction with the user. The voice assistant is also used to display a corresponding voice identifier based on the user's voice command input. The smart car 100 may be internally provided with one or more microphones and one or more speakers to assist the voice assistant on the display screen in implementing voice interaction with the user. The microphone can be used to collect the user's voice, and the speaker can be used to output the voice and also output user-specified audio.

[0121] In some embodiments, the microphone and speaker can be disposed inside the display screen. For example, each display screen can be provided with one or more microphones and one or more speakers. The display screen can assist the voice assistant in implementing voice interaction with the user through the microphone and speaker inside the display screen.

[0122] In other embodiments, the microphone and speaker can be located outside of the display screen, for example, a microphone and speaker can be installed next to each seat. In this case, the smart car 100 can control one or more microphones in the car to collect the user's voice in the car, and after recognizing the collected voice, activate the voice assistant on the corresponding display screen, display the voice identification through the voice assistant on the display screen, and optionally call the corresponding speaker to output the voice or play the audio specified by the user.

[0123] In other embodiments, the microphone and the speaker may also be respectively arranged inside the display screen and outside the display screen. For example, one or more speakers are arranged inside each display screen, and one or more microphones are respectively arranged at different seats in the car. For another example, one or more microphones and one or more speakers are arranged inside each display screen, and one or more microphones are respectively arranged at different seats in the car, etc. In the above case, the smart car 100 can also realize voice interaction with the user by controlling the microphone and speaker in the car to assist the voice assistant of the display screen. It can be understood that the above multiple embodiments are just some examples, and this application does not limit the setting position of the microphone and speaker.

[0124] FIG3B shows a schematic diagram of the interior appearance of another smart car 100 provided in an embodiment of the present application.

[0125] As shown in FIG3A , the interior of the smart car 100 may be provided with a plurality of seats and may also be provided with one or more display screens. For example, the interior of the smart car 100 may include three rows of seats, the first row of seats may include a main driver's seat and a front passenger seat, the second row of seats may include a second row of left seats and a second row of right seats, and the third row of seats may include a third row of left seats and a third row of right seats. The one or more display screens inside the smart car 100 may include a central control screen 10, a front passenger screen 20, and a laser screen 40. The central control screen 10 may be provided between the main driver's seat and the front passenger seat, the front passenger screen 20 may be provided in front of the front passenger seat, and the laser screen 40 may be provided directly in front of the second row of seats, for example, across the backrest of the main driver's seat and the backrest of the front passenger seat.

[0126] In some embodiments, each display screen may be equipped with a voice assistant for voice interaction with the user. The voice assistant is also used to display a corresponding voice identifier based on the user's voice command input. The smart car 100 may be internally provided with one or more microphones and one or more speakers to assist the voice assistant on the display screen in implementing voice interaction with the user. The microphone can be used to collect the user's voice, and the speaker can be used to output the voice and also output user-specified audio.

[0127] The specific configuration of the microphone and speaker inside the smart car 100 can refer to the relevant description in the embodiment shown in Figure 3A above, and will not be repeated here.

[0128] It will be understood that Figures 3A to 3B are just two examples. In the embodiments of the present application, the interior of the smart car 100 may further include more, fewer, or differently arranged seats than in the above-mentioned embodiments, and may also include more, fewer, or differently distributed display screens than in the above-mentioned embodiments. The present application does not limit this.

[0129] In the embodiment of the present application, the smart car 100 can be divided into different audio zones according to the seats, and each audio zone can correspond to a display screen, which can be used to display the voice interaction status of the user in the audio zone and the smart car 100. The following describes the distribution of audio zones within the smart car 100.

[0130] FIG3C shows a sound zone distribution provided in an embodiment of the present application.

[0131] As shown in FIG3C , the interior of the smart car 100 can be divided into multiple sound zones according to the seating arrangement, for example, a driver's seat zone, a passenger driver's seat zone, a second-row left zone, a second-row right zone, a third-row left zone, a third-row right zone, etc. Each sound zone may include a seat. For example, the driver's seat zone may include the driver's seat, the passenger driver's seat zone may include the passenger driver's seat, the second-row left zone may include the second-row left seat, the second-row right zone may include the second-row right seat, the third-row left zone may include the third-row left seat, the third-row right zone may include the third-row right seat, etc.

[0132] FIG3D shows another sound zone distribution provided by an embodiment of the present application.

[0133] As shown in FIG3D , the interior of the smart car 100 can be divided into multiple sound zones based on seating arrangements, such as a driver's seat zone, a passenger seat zone, and a rear seat zone. Each sound zone may include one or more seats. For example, the driver's seat zone may include the driver's seat, the passenger seat zone may include the passenger seat, and the rear seat zone may include all seats except the driver's seat and the passenger seat.

[0134] It is understandable that Figures 3C and 3D are just two examples. In the embodiments of the present application, the smart car 100 may also include more, fewer, or different sound zones than those in the above embodiments, and the present application does not limit this.

[0135] In one application scenario, if the internal shape of the smart car 100 is the same as the embodiment shown in Figure 3A, and the internal sound zone distribution is the same as the embodiment shown in Figure 3C, then when the display screens inside the smart car 100 are all in the turned-on state, the correspondence between the sound zones and the display screens inside the smart car 100 can refer to the embodiment shown in Figure 3E below.

[0136] As shown in Figure 3E, when the central control screen 10, the passenger screen 20 and the right rear projection screen 30 inside the smart car 100 are all in the turned-on state, the central control screen 10 can correspond to the following multiple sound zones: the main driver sound zone, the second row left sound zone, the third row left sound zone and the third row right sound zone, that is, the central control screen 10 can process the voice from the main driver sound zone, the second row left sound zone, the third row left sound zone and the third row right sound zone; the passenger screen 20 can correspond to the passenger sound zone, that is, the passenger screen 20 only needs to process the voice from the passenger sound zone; the right rear projection screen 30 can correspond to the second row right sound zone, that is, the right rear projection screen 30 only needs to process the voice from the second row right sound zone.

[0137] It is understandable that the embodiment shown in Figure 3E is only an example. In the embodiment of the present application, the correspondence between the internal display screen of the smart car 100 and the sound zone may also be a different correspondence from the above embodiment, and the present application does not limit it here.

[0138] In one application scenario, if the internal shape of the smart car 100 is the same as the embodiment shown in Figure 3A, and the internal sound zone distribution is the same as the embodiment shown in Figure 3C, then when the central control screen 10 and the co-pilot screen 20 inside the smart car 100 are in the on state, and the right rear projection screen 30 is in the off state, the central control screen 10 can take over the sound zone originally corresponding to the right rear projection screen 30. At this time, the correspondence between the sound zones and display screens inside the smart car 100 can refer to the embodiment shown in Figure 3F below.

[0139] As shown in Figure 3F, when the central control screen 10 and the passenger screen 20 inside the smart car 100 are both in the on state, and the right rear projection screen 30 is in the off state, the central control screen 10 can correspond to the following multiple sound zones: main driver sound zone, second row left sound zone, second row right sound zone, third row left sound zone and third row right sound zone, that is, the central control screen 10 can process voices from the main driver sound zone, second row left sound zone, second row right sound zone, third row left sound zone and third row right sound zone; the passenger screen 20 can correspond to the passenger sound zone, that is, the passenger screen 20 only needs to process voices from the passenger sound zone.

[0140] It is understood that Figures 3E to 3F are merely illustrative. When other display screens within the smart car 100 are off, the central control screen 10 can take over the audio zone corresponding to the off display screen and process the audio in that audio zone. In the embodiment of the present application, the off display screen can also be the passenger screen 20, in which case the central control screen 10 can also take over the passenger audio zone corresponding to the passenger screen 20. Furthermore, if all display screens other than the central control screen 10 are off, or if the smart car 100 only has the central control screen 10, the central control screen 10 can process audio in all audio zones within the entire smart car 100, but this application does not limit this.

[0141] In one application scenario, if the internal shape of the smart car 100 is the same as the embodiment shown in Figure 3B, and the internal sound zone distribution is the same as the embodiment shown in Figure 3D, then when the display screens inside the smart car 100 are all in the turned-on state, the correspondence between the sound zones and the display screens inside the smart car 100 can refer to the embodiment shown in Figure 3G below.

[0142] As shown in Figure 3G, when the central control screen 10, the passenger screen 20 and the laser curtain 40 inside the smart car 100 are all in the turned-on state, the central control screen 10 can correspond to the main driver sound zone, that is, the central control screen 10 can process the voice from the main driver sound zone; the passenger screen 20 can correspond to the passenger sound zone, that is, the passenger screen 20 only needs to process the voice from the passenger sound zone; the laser curtain 40 can correspond to the rear sound zone, that is, the laser curtain 40 only needs to process the voice from the rear sound zone.

[0143] It is understandable that the embodiment shown in Figure 3G is only an example. In the embodiment of the present application, the correspondence between the internal display screen of the smart car 100 and the sound zone may also be a different correspondence from the above embodiment, and the present application does not limit it here.

[0144] In one application scenario, if the internal shape of the smart car 100 is the same as the embodiment shown in Figure 3B, and the internal sound zone distribution is the same as the embodiment shown in Figure 3D, then when the central control screen 10 and the co-pilot screen 20 inside the smart car 100 are in the open state and the laser curtain 40 is in the closed state, the central control screen 10 can take over the sound zone originally corresponding to the laser curtain 40. At this time, the correspondence between the sound zone and the display screen inside the smart car 100 can refer to the embodiment shown in Figure 3H below.

[0145] As shown in Figure 3H, when the central control screen 10 and the passenger screen 20 inside the smart car 100 are in the on state and the laser curtain 40 is in the off state, the central control screen 10 can correspond to the main driver sound zone and the rear sound zone, that is, the central control screen 10 can process the voice from the main driver sound zone and the rear sound zone; the passenger screen 20 can correspond to the passenger sound zone, that is, the passenger screen 20 only needs to process the voice from the passenger sound zone.

[0146] It is understood that Figures 3G to 3H are merely illustrative. When other display screens within the smart car 100 are off, the central control screen 10 can take over the audio zone corresponding to the off display screen and process the audio in that audio zone. In the embodiment of the present application, the off display screen can also be the passenger screen 20. In this case, the central control screen 10 can also take over the passenger audio zone corresponding to the passenger screen 20. In addition, if all display screens other than the central control screen 10 are off, or if the smart car 100 only has the central control screen 10, the central control screen 10 can process audio in all audio zones of the entire smart car 100. This application does not limit this.

[0147] The following introduces the hardware structure of a smart car 100 provided in an embodiment of the present application.

[0148] FIG4A shows a schematic diagram of the hardware structure of a smart car 100 provided in an embodiment of the present application.

[0149] As shown in FIG4A , the smart car 100 includes a controller area network (CAN) bus 11, multiple electronic control units (ECUs), an engine 13, a telematics box (T-box) 14, a transmission 15, a driving recorder 16, an anti-lock braking system (ABS) 17, a sensor system 18, a camera system 19, a microphone 20, and the like.

[0150] The CAN bus 11 is a serial communication network that supports distributed control or real-time control and is used to connect the various components of the smart car 100. Any component on the CAN bus 11 can monitor all data transmitted on the CAN bus 11. The frames transmitted by the CAN bus 11 can include data frames, remote frames, error frames, and overload frames, and different frames transmit different types of data. In an embodiment of the present application, the CAN bus 11 can be used to transmit data involved in the voice command-based control method of each component. The specific implementation of this method can be referred to the detailed description of the method embodiment below.

[0151] The CAN bus 11 is not limited to the CAN bus. In other embodiments, the various components of the smart car 100 can also be connected and communicated through other means. For example, the various components can also communicate through an in-vehicle Ethernet (Ethernet), a local interconnect network (LIN) bus, FlexRay, or a common in-vehicle media oriented system (MOST) bus, etc., and the present embodiment does not limit this. The following embodiments are described as the various components communicating through the CAN bus 11.

[0152] The ECU is the processor or brain of the smart car 100, instructing components to perform actions based on instructions received from the CAN bus 11 or user input. The ECU can consist of a security chip, a microprocessor (MCU), random access memory (RAM), read-only memory (ROM), input / output (I / O) interfaces, an analog / digital converter (A / D converter), and large-scale integrated circuits for input, output, shaping, and driving functions.

[0153] There are many types of ECUs, and different types of ECUs can be used to achieve different functions.

[0154] The multiple ECUs in the smart car 100 may include, for example, an engine ECU 121 , a telematics box (T-box) ECU 122 , a transmission ECU 123 , a driving recorder ECU 124 , an anti-lock brake system (ABS) ECU 125 , and the like.

[0155] The engine ECU 121 manages the engine and coordinates its various functions, such as starting and shutting it down. The engine is the device that powers the smart car 100. An engine is a machine that converts some form of energy into mechanical energy. The smart car 100 can be used to convert chemical energy from liquid or gas combustion, or electrical energy, into mechanical energy and output it as power. The engine consists of two major mechanisms: the crankshaft and the valve train, as well as five major systems: cooling, lubrication, ignition, energy supply, and starting. The main components of the engine include the cylinder block, cylinder head, piston, piston pin, connecting rod, crankshaft, and flywheel.

[0156] The T-box ECU 122 is used to manage the T-box 14 .

[0157] T-box 14 is mainly responsible for communicating with the Internet, providing a remote communication interface for the smart car 100, and providing services including navigation, entertainment, driving data collection, driving trajectory recording, vehicle fault monitoring, vehicle remote query and control (such as unlocking and closing, air conditioning control, window control, engine torque limit, engine start and stop, seat adjustment, query of battery power, fuel level, door status, etc.), driving behavior analysis, wireless hotspot sharing, roadside assistance, abnormal reminders, etc.

[0158] T-box 14 can be used to communicate with a telematics service provider (TSP) and user (e.g., driver) side electronic devices to display and control the vehicle status on the electronic devices. When the user sends a control command through the vehicle management application on the electronic device, the TSP will issue a request instruction to the T-box 14. After obtaining the control command, the T-box 14 sends a control message through the CAN bus and controls the smart car 100, and finally feeds back the operation results to the vehicle management application on the user side electronic device. In other words, the data read by the T-box 14 through the CAN bus 11, such as vehicle condition reports, driving reports, fuel consumption statistics, violation inquiries, location trajectories, driving behavior and other data, can be transmitted to the TSP background system via the network, and forwarded by the TSP background system to the user side electronic device for the user to view.

[0159] The T-box 14 may specifically include a communication module and a display screen.

[0160] Among them, the communication module can be used to provide wireless communication functions, supporting the smart car 100 to communicate with other devices through wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), ultra-wideband (UWB) and other wireless communication technologies. The communication module can also be used to provide mobile communication functions, supporting the smart car 100 to communicate with other devices through communication technologies such as global system for mobile communications (GSM), universal mobile telecommunications system (UMTS), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), 5G and the future 6G.

[0161] The communication module can establish connections and communicate with other devices such as servers and user-side electronic devices through cellular V2X (vehicle to everything) communication technology (C-V2X) based on cellular networks. C-V2X can include, for example, long-term evolution (LTE)-based V2X (LTE-V2X) and 5G-V2X.

[0162] The display screen is used to provide a visual interface for the driver. The smart car 100 may include one or more display screens, for example, it may include an on-board display screen arranged in front of the driver's seat, a display screen arranged above the seat for displaying the surrounding situation, and a head-up digital display (HUD) that projects information onto the windshield, etc. The display screen for displaying the user interface in the smart car 100 provided in the subsequent embodiments may be an on-board display screen arranged next to the seat, a display screen arranged above the seat, a HUD, etc., which are not limited here. The user interface displayed on the display screen in the smart car 100 can be specifically described in detail in the subsequent embodiments, and will not be repeated here.

[0163] T-box 14 may also be referred to as a vehicle system, a telematics processor, a vehicle gateway, etc., which is not limited in the embodiments of the present application.

[0164] The transmission ECU 123 is used to manage the transmission.

[0165] The transmission 15 is a mechanism used to change the engine's speed and torque. It can change the output-to-input ratio in a fixed or step-by-step manner. Transmission 15 components may include a transmission mechanism, an operating mechanism, and a power take-off mechanism. The transmission mechanism primarily changes the magnitude and direction of torque and speed; the operating mechanism primarily controls the transmission mechanism to achieve gear shifting, thereby changing the transmission ratio and achieving speed and torque variations.

[0166] The drive recorder ECU 124 is used to manage the drive recorder 16 .

[0167] The components of the driving recorder 16 may include a host computer, a vehicle speed sensor, and data analysis software. The driving recorder 16 is a device that records images and sounds of a vehicle while it is in motion, including relevant information such as driving time, speed, and location. In this embodiment of the present application, while the vehicle is in motion, the vehicle speed sensor collects wheel speed information and transmits it to the driving recorder 16 via the CAN bus.

[0168] The ABS ECU 125 is used to manage the ABS 17 .

[0169] ABS17 automatically controls the braking force during vehicle braking to prevent the wheels from locking and maintain a rolling and sliding state, thereby ensuring maximum adhesion between the wheels and the ground. During braking, if the electronic control unit determines that a wheel is approaching locking based on the wheel speed signal input by the wheel speed sensor, the ABS will enter the anti-lock brake pressure adjustment process.

[0170] The sensor system 18 may include: an accelerometer, a vehicle speed sensor, a vibration sensor, a gyroscope sensor, a radar sensor, a signal transmitter, a signal receiver, and the like. The accelerometer and vehicle speed sensor are used to detect the speed of the smart car 100. The vibration sensor can be installed under the seat, on the seat belt, on the seat back, on the operating panel, on the airbag, or in other locations to detect whether the smart car 100 has been hit and the user's location. The gyroscope sensor can be used to determine the motion posture of the smart car 100. Radar sensors may include lidar, ultrasonic radar, millimeter-wave radar, and the like. Radar sensors are used to transmit electromagnetic waves to illuminate a target and receive their echoes, thereby obtaining information such as the distance from the target to the electromagnetic wave emission point, the rate of change of distance (radial velocity), direction, and altitude, thereby identifying other vehicles, pedestrians, or roadblocks near the smart car 100. The signal transmitter and signal receiver are used to send and receive signals, which can be used to detect the user's location. The signal can be, for example, ultrasonic, millimeter-wave, or laser.

[0171] The camera system 19 may include multiple cameras for capturing still images or videos. The cameras in the camera system 19 may be located in front of, behind, on the sides of, or inside the vehicle, to facilitate functions such as assisted driving, driving recording, panoramic surround view, and in-vehicle monitoring.

[0172] The sensor system 18 and the camera system 19 can be used to detect the surrounding environment, so that the smart car 100 can make corresponding decisions to cope with environmental changes. For example, they can be used to complete the task of paying attention to the surrounding environment during the autonomous driving stage.

[0173] The microphone 20, also known as a "microphone" or "microphone," is used to convert sound signals into electrical signals. When making a call or outputting a voice command, the user can speak by placing their mouth close to the microphone 20 to input the sound signal into the microphone 20. The smart car 100 can be provided with at least one microphone 20. In other embodiments, the smart car 100 can be provided with two microphones 20, which, in addition to collecting sound signals, can also implement noise reduction functions. In other embodiments, the smart car 100 can also be provided with three, four, or more microphones 20 to form a microphone array to collect sound signals, reduce noise, identify the sound source, implement directional recording functions, and the like.

[0174] In addition, the smart car 100 may also include multiple interfaces, such as a USB interface, an RS-232 interface, an RS485 interface, etc., which can be connected to external cameras, microphones, headphones, and user-side electronic devices, such as a mobile phone 200.

[0175] In the embodiments of the present application, microphone 20 can be used to detect voice commands input by a user. Sensor system 18, camera system 19, T-box 14, etc. can be used to obtain role information of the user inputting the voice command. The manner in which the various components of smart car 100 obtain user role information can be found in the relevant descriptions in the subsequent method embodiments. T-box ECU 122 can be used to determine whether the user currently has the authority corresponding to the voice command based on this role information. Only if the user currently has the authority will T-box ECU 122 dispatch the corresponding components of smart car 100 to respond to the voice command.

[0176] In some embodiments, sensor system 18, camera system 19, T-box 14, etc. are used not only to obtain the role information of the user inputting the voice command, but also to obtain the role information of other users. T-box ECU 122 can combine the role information of the user inputting the voice command and the role information of other users to determine whether the current user has the authority corresponding to the voice command.

[0177] In some embodiments, the sensor system 18, the camera system 19, the T-box 14, etc. may be used to obtain the vehicle status of the smart car 100. The T-box ECU 122 may be used to determine whether the user currently has the authority corresponding to the voice command based on the vehicle status and the user's role information.

[0178] In some embodiments, the memory in the smart car 100 can be used to store the binding relationship between the vehicle and the user.

[0179] It should be understood that the structures illustrated in the embodiments of this application do not constitute specific limitations on the vehicle system. The smart car 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0180] For example, the smart car 100 may also include a separate memory, battery, headlights, wipers, instrument panel, audio, a transmission control unit (TCU), an auxiliary control unit (ACU), a passive entry and start system (PEPS), an on-board unit (OBU), a body control module (BCM), a charging port, and the like. The memory may be used to store permission information for different roles in the smart car 100, which indicates the permissions that the role has or does not have for using the smart car 100. In some embodiments, the memory may be used to store permission information for different roles in the smart car 100 under different vehicle states. In some embodiments, the memory may be used to store permission information for different roles in the smart car 100 when they have different other roles.

[0181] The specific functions of the various components of the smart car 100 can also be referred to the introduction of the subsequent method embodiments, which will not be repeated here.

[0182] The following describes the hardware structure of a mobile phone 200 provided in an embodiment of the present application.

[0183] FIG4B shows the hardware structure of a mobile phone 200 provided in an embodiment of the present application.

[0184] The mobile phone 200 may include a processor 210, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a sensor module 280, a button 290, a motor 291, an indicator 292, and a display screen 294. The sensor module 280 may include a touch sensor 280K. Optionally, the sensor module 280 may further include any one or more of the following: a pressure sensor, a gyroscope sensor, an air pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, an ambient light sensor, a bone conduction sensor, and the like.

[0185] It should be understood that the structure illustrated in the embodiment of the present invention does not constitute a specific limitation on the mobile phone 200. In other embodiments of the present application, the mobile phone 200 may include more or fewer components than shown, or some components may be combined or separated, or the components may be arranged differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0186] The processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors.

[0187] The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of instruction fetching and execution.

[0188] Processor 210 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 210 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 210. If processor 210 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 210 latency, and thus improves system efficiency.

[0189] In some embodiments, the processor 210 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.

[0190] USB interface 230 is an interface that complies with USB standards and specifications, and may be a Mini USB interface, a Micro USB interface, a USB Type-C interface, or the like. USB interface 230 can be used to connect a charger to charge mobile phone 200, or to transfer data between mobile phone 200 and peripheral devices. It can also be used to connect headphones to play audio. This interface can also be used to connect other electronic devices, such as AR devices.

[0191] It is understood that the interface connection relationship between the modules illustrated in the embodiment of the present invention is merely an illustrative illustration and does not constitute a structural limitation on the mobile phone 200. In other embodiments of the present application, the mobile phone 200 may also adopt a different interface connection method from the above embodiment, or a combination of multiple interface connection methods.

[0192] The charging management module 240 is configured to receive charging input from a charger. The charger can be either a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 240 can receive charging input from the wired charger via the USB interface 230. In some wireless charging embodiments, the charging management module 240 can receive wireless charging input via the wireless charging coil of the mobile phone 200. While charging the battery 242, the charging management module 240 can also power the electronic device through the power management module 241.

[0193] The power management module 241 is used to connect the battery 242, the charging management module 240, and the processor 210. The power management module 241 receives input from the battery 242 and / or the charging management module 240 and provides power to the processor 210, the internal memory 221, the display 294, the wireless communication module 260, and the like. The power management module 241 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some other embodiments, the power management module 241 can also be provided in the processor 210. In other embodiments, the power management module 241 and the charging management module 240 can also be provided in the same device.

[0194] The wireless communication function of the mobile phone 200 can be implemented through the antenna 1, the antenna 2, the mobile communication module 250, the wireless communication module 260, the modem processor and the baseband processor.

[0195] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in mobile phone 200 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.

[0196] The mobile communication module 250 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the mobile phone 200. The mobile communication module 250 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 250 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 250 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 250 can be set in the processor 210. In some embodiments, at least some of the functional modules of the mobile communication module 250 can be set in the same device as at least some of the modules of the processor 210.

[0197] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 270A, the receiver 270B, etc.) or displays an image or video through the display screen 294. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 210 and be set in the same device as the mobile communication module 250 or other functional modules.

[0198] The wireless communication module 260 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc., which are applied to the mobile phone 200. The wireless communication module 260 can be one or more devices that integrate at least one communication processing module. The wireless communication module 260 receives electromagnetic waves via the antenna 2, demodulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 210. The wireless communication module 260 can also receive the signal to be sent from the processor 210, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.

[0199] In some embodiments, the antenna 1 of the mobile phone 200 is coupled to the mobile communication module 250, and the antenna 2 is coupled to the wireless communication module 260, so that the mobile phone 200 can communicate with the network and other devices through wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS may include a global positioning system (GPS), a global navigation satellite system (GLONASS), a Beidou navigation satellite system (BDS), a quasi-zenith satellite system (QZSS) and / or a satellite based augmentation system (SBAS).

[0200] Mobile phone 200 implements display functionality through a GPU, display screen 294, and an application processor. The GPU is a microprocessor for image processing that connects display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 210 may include one or more GPUs that execute program instructions to generate or modify display information.

[0201] Display screen 294 is used to display images, videos, and the like. Display screen 294 includes a display panel. The display panel can be a liquid crystal display (LCD). Alternatively, the display panel can be made of an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, or a quantum dot light-emitting diode (QLED). In some embodiments, mobile phone 200 can include one or N display screens 294, where N is a positive integer greater than one.

[0202] The internal memory 221 may include one or more random access memories (RAM) and one or more non-volatile memories (NVM). The RAM can be directly read and written by the processor 210 and can be used to store executable programs (e.g., machine instructions) of the operating system or other running programs, as well as user and application data. The NVM can also store executable programs and user and application data, and can be pre-loaded into the RAM for direct reading and writing by the processor 210.

[0203] The audio module 270 may include a speaker 270A, a receiver 270B, and a microphone 270C. The mobile phone 200 may implement audio functions such as music playback and recording through the audio module 270 and the application processor.

[0204] The audio module 270 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 270 can also be used to encode and decode audio signals. In some embodiments, the audio module 270 can be provided in the processor 210, or some functional modules of the audio module 270 can be provided in the processor 210.

[0205] The speaker 270A, also called a "horn," is used to convert audio electrical signals into sound signals. The mobile phone 200 can listen to music or make hands-free calls through the speaker 270A.

[0206] The receiver 270B, also called the "earpiece", is used to convert audio electrical signals into sound signals. When the mobile phone 200 receives a call or a voice message, the voice can be heard by placing the receiver 270B close to the ear.

[0207] The microphone 270C, also known as a "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to the microphone 270C to input the sound signal into the microphone 270C. The mobile phone 200 can be provided with at least one microphone 270C. In other embodiments, the mobile phone 200 can be provided with two microphones 270C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the mobile phone 200 can also be provided with three, four or more microphones 270C to realize sound signal collection, noise reduction, and identification of sound sources, and realize directional recording function, etc.

[0208] The touch sensor 280K is also referred to as a "touch device." The touch sensor 280K can be disposed on the display screen 294. The touch sensor 280K and the display screen 294 form a touch screen, also referred to as a "touch screen." The touch sensor 280K is used to detect touch operations applied thereto or in the vicinity thereof. The touch sensor can transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operations can be provided via the display screen 294. In other embodiments, the touch sensor 280K can also be disposed on the surface of the mobile phone 200, at a location different from that of the display screen 294.

[0209] Keys 290 include a power button, a volume button, and the like. Keys 290 may be mechanical keys or touch-sensitive keys. Mobile phone 200 may receive key inputs and generate key signal inputs related to user settings and function control of mobile phone 200.

[0210] Motor 291 can generate vibration prompts. Motor 291 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playback, etc.) can correspond to different vibration feedback effects. For touch operations acting on different areas of the display screen 294, motor 291 can also correspond to different vibration feedback effects. Different application scenarios (for example: time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.

[0211] The indicator 292 may be an indicator light, which may be used to indicate the charging status, power level change, messages, missed calls, notifications, etc.

[0212] It should be noted that the hardware structure of the headset 300 can also refer to the hardware structure of the mobile phone 200 shown in Figure 4B above. Moreover, the headset 300 may include fewer, more or different components than the mobile phone 200, and this application does not limit this.

[0213] An embodiment of the present application provides a voice interaction method. A smart car 100 may include one or more display screens, such as display screen A and display screen B. The smart car 100 may carry multiple users (e.g., user A, user B, etc.), with user A located in audio zone a and user B located in audio zone b. During the process of user A performing voice interaction with display screen A, the smart car 100 may receive and respond to voice 1 emitted by user B. If it is determined that voice 1 is a wake-up-free command or that voice 1 includes a wake-up word, the smart car 100 may control display screen B corresponding to audio zone b to perform voice interaction with user B. Display screen A may be the same as display screen B or different from display screen B.

[0214] In this way, the smart car 100 can perform voice interaction with multiple users at the same time. Moreover, in a multi-user voice interaction scenario, the user can see in real time whether the voice command is recognized without waiting, which brings a better user experience.

[0215] In some application scenarios, a smart car 100 carries user A and user B, with user A located in audio zone a and user B located in audio zone b. When screen A of the smart car 100 displays interface 1, the smart car 100 receives a wake-up word from user A. In response to the wake-up word from user A in audio zone a, the smart car 100, after determining that the display corresponding to audio zone a is display A, may designate display A as the primary display and display a voice dialogue identifier 1 on interface 1 of display A. Voice dialogue identifier 1 serves to notify the user that the voice assistant on display A has been activated using the wake-up word. While display A displays voice dialogue identifier 1, the smart car 100 may receive and respond to a wake-up-free command from the user in audio zone b. Upon determining that the display corresponding to audio zone b is display A, the smart car 100 may split voice sub-identifiers from voice dialogue identifier 1 on interface 1 to notify the user that display A is processing multiple voice channels.

[0216] In this way, while user A is interacting with display screen A through the voice assistant, user B can also interact with display screen A through the wake-up-free command. The smart car 100 can interact with multiple users at the same time, providing users with more convenient services.

[0217] 5A , the central control screen 10 may display a main interface 500. The main interface 500 may display one or more application icons, such as a music application icon, a navigation application icon, a setting application icon, and the like.

[0218] The smart car 100 can receive and respond to the wake-up word issued by user A from sound zone a. If the smart car 100 determines that the display screen corresponding to sound zone a is the central control screen 10, the smart car 100 can determine the central control screen 10 as the main display screen and display the voice dialogue logo 501 as shown in Figure 5B on the central control screen 10.

[0219] As shown in Figure 5B, the main interface 500 includes a voice dialogue logo 501. The voice dialogue logo 501 may include a rounded rectangular frame 502 and a dialogue symbol 503. Optionally, it may also include text 504. The voice dialogue logo 1 is used to prompt the user that the voice assistant of the central control screen 10 has been awakened by the wake-up word. In the voice dialogue logo 501, the dialogue symbol 503 and the text 504 may be displayed in the rounded rectangular frame 502. The dialogue symbol 503 may be a symbol composed of multiple vertical lines of varying lengths, and the length of each of the multiple vertical lines may vary over time to simulate changes in sound volume. The dialogue symbol 503 may be used to prompt the user that the central control screen 10 is interacting with the user by voice. The text 504 may be used to prompt the user to start issuing voice commands. For example, the text 504 may be "You can speak at any time". In this embodiment of the present application, the rounded rectangular voice dialogue logo 501 shown in Figure 5B may also be referred to as a voice capsule. It is understandable that the voice dialogue identifier 501 may also adopt a shape different from the above embodiments, such as a rectangle, a circle, or a heart shape, the dialogue symbol 503 may also adopt a symbol different from the above embodiments, and the text inside the voice dialogue identifier 501 may also be text different from the above embodiments, and this application does not limit this. Further, optionally, a sound zone indicator 505 may also be displayed in the main interface 500. The sound zone indicator 505 can be used to indicate the sound zone in which the user (i.e., user A) is currently interacting with the central control screen 10 by voice. The display area of ​​the sound zone indicator 505 shown in Figure 5B includes the lower left area of ​​the voice dialogue identifier 501. The sound zone indicator 505 can be used to indicate that the sound zone a in which user A is located is the sound zone on the left side of the central control screen 10 (e.g., the main driving sound zone). It is understandable that the sound zone indicator 505 here is only an example. In the embodiment of the present application, the sound zone indicator may also adopt a shape, color, or a different symbol from the above-mentioned sound zone indicator 505, and this application does not limit this.

[0220] After the voice assistant is turned on on the central control screen 10, the smart car 100 can receive and respond to the voice command issued by user A from the sound zone a (for example, "The weather is nice today, open the car window"). As shown in Figure 5C, the central control screen 10 can replace the text 504 displayed in the voice dialogue identifier 501 with text 506.

[0221] As shown in FIG5C , the voice dialogue identifier 501 may display text 506 , and the content of the text 506 may be the same as or similar to the voice command issued by the user. For example, the text 506 may include “The weather is nice today, open the windows.” It should be noted that due to the limited size of the voice dialogue identifier 501 , if the number of characters included in the text 506 is too large and exceeds the display range of the voice dialogue identifier 501 , the central control screen 10 may use a scrolling display method to scroll the text 506 in the voice dialogue identifier 501 . The text 506 can be used to prompt the user that the central control screen 10 is receiving and processing the voice command issued by the user.

[0222] While receiving a voice command from user A (e.g., "It's a nice day today, open the windows"), the smart car 100 may also receive a wake-up-free command from user B in audio zone b. In response to the wake-up-free command from user B in audio zone b, if the display corresponding to audio zone b is determined to be the central control screen 10, the central control screen 10 of the smart car 100 may split the voice dialogue indicator 501, displaying a split indicator 510 as shown in FIG5D .

[0223] As shown in FIG5D , a split mark 510 is displayed in the main interface 500. The split mark 510 may include two voice marks that are not completely split, one of which is a voice dialogue mark 520 and the other is a voice sub-marker 530. There may also be a connection channel between the voice dialogue mark 520 and the voice sub-marker 530, indicating that the two voice marks are still in the process of splitting and have not yet been completely split. The split mark 510 can be used to indicate that the central control screen 10 is processing the voices of multiple users at the same time. Among them, the specific content of the voice dialogue mark 520 can refer to the relevant description of the voice dialogue mark 501 in the embodiment shown in FIG5C above, and will not be repeated here. The voice sub-marker 530 may include a rounded rectangular frame and a sub-symbol 531. The sub-symbol 531 can be used to prompt the user that the voice sub-marker 530 where the sub-symbol 531 is located is a split voice mark. In some embodiments, the sub-symbol 531 may be composed of multiple points of the same size (or line segments of the same length). In other embodiments, the sub-symbol 531 may also use other symbols, which is not limited in this application. Optionally, a sound zone indicator 511 can also be displayed in the main interface 500. The display area of ​​the sound zone indicator 511 is the area below the split mark 510 (including the lower left area, the area directly below, and the lower right area). The sound zone indicator 511 is used to indicate that the sound zones of the multi-channel voice being processed by the central control screen 10 are located at different positions of the central control screen 10, that is, the sound zone a where user A is located is the sound zone on the left side of the central control screen 10, and the sound zone b where user B is located is the sound zone on the right side of the central control screen 10 (for example, the right sound zone of the third row).

[0224] After the split identifier 510 is split, as shown in Figure 5E, the main interface 500 of the central control screen 10 can display the completely split voice dialogue identifier 520 and voice sub-identifier 530, that is, there is no connection channel between the voice dialogue identifier 520 and the voice sub-identifier 530.

[0225] After the smart car 100 executes the wake-up-free command issued by user B, as shown in FIG5F , the central control screen 10 may stop displaying the voice sub-identifier 530. Optionally, the main interface 500 may also display a sound zone indicator 521, which may be the same as the sound zone indicator 505 shown in FIG5B above.

[0226] It should be noted that after the central control screen 10 executes the voice command issued by user A after the wake-up word, when it is detected that the shutdown condition is met (for example, the continuous time during which user A does not make a voice sound is greater than the preset time, etc.), the central control screen 10 can turn off the voice dialogue logo 520 (and the sound zone indicator 521) displayed in the main interface 500, that is, the central control screen 10 can redisplay the main interface 500 shown in Figure 5A above.

[0227] In some embodiments, while executing a user's voice command (or after executing the voice command), the smart car 100 may also output feedback information through one or more methods, such as voice broadcast, flashing indicator lights, vibration, and display screen display. The feedback information is used to inform the user that the smart car 100 has executed the operation corresponding to the voice command. The specific content of the feedback information can be found in the relevant description of the embodiment shown in Figure 6D below and is not described in detail here.

[0228] It can be understood that the embodiment shown in Figures 5A to 5F above is only an example. In the embodiment of the present application, the user can also wake up the voice assistant of other display screens (such as the co-pilot screen 20, the right rear projection screen 30, the laser curtain 40, etc.) through the wake-up word, and the interface 1 displayed on the display screen A can also be other interfaces other than the main interface. In addition, the voice dialogue identifier and the voice sub-identifier can also adopt shapes and / or symbols different from those in the above embodiment, or include content different from that in the above embodiment. In addition, the voice command issued by user A can also be a voice command different from that in the above embodiment, and the present application does not limit this.

[0229] In some embodiments, the same display screen in the smart car 100 (e.g., the central control screen 10, the laser screen 40, etc.) can also receive and process wake-up-free commands issued by two or more other users while processing non-wake-up-free commands (e.g., a voice command issued by user A after saying the wake-up word). In this case, the display screen can also display a voice dialogue identifier and a voice sub-identifier. The voice dialogue identifier is used to prompt the user of the voice interaction between the display screen and user A, and the voice sub-identifier is used to prompt that the display screen is also processing one or more other wake-up-free commands. The specific interface can also refer to the relevant description in the embodiment shown in Figure 5E above, which will not be repeated here.

[0230] In some application scenarios, a smart car 100 carries user A and user B, with user A located in audio zone a and user B located in audio zone b. When screen A of the smart car 100 displays interface 1 and screen B displays interface 2, the smart car 100 receives a wake-up word from user A. In response to the wake-up word from user A in audio zone a, the smart car 100, after determining that the display corresponding to audio zone a is screen A, may designate screen A as the primary display and display a voice dialogue indicator 1 on screen A's interface 1. Voice dialogue indicator 1 serves to prompt the user that the voice assistant on screen A has been activated using the wake-up word. While screen A displays voice dialogue indicator 1, the smart car 100 may receive and respond to a wake-up-free command from the user in audio zone b. Upon determining that the display corresponding to audio zone b is screen B, the wake-up-free indicator may be displayed on screen B's interface 2. The wake-up-free indicator serves to prompt the user that screen B is receiving and processing the user's wake-up-free command.

[0231] In this way, while user A is interacting with display screen A through the voice assistant, user B can also interact with display screen B through the wake-up command. The smart car 100 can interact with multiple users at the same time, providing users with more convenient services.

[0232] For example, when the smart car 100 receives the wake-up word issued by user A from sound zone a and determines that the display screen corresponding to sound zone a is the central control screen 10, the central control screen 10 of the smart car 100 can display the main interface 500 as shown in Figure 6A, and the co-pilot screen 20 can display the main interface 600 as shown in Figure 6B.

[0233] As shown in FIG6A , the central control screen 10 of the smart car 100 can display a main interface 500, and the main interface 500 can display a voice dialogue logo 501. The voice dialogue logo 501 is used to prompt the user that the voice assistant of the central control screen 10 has been turned on by the wake-up word. Among them, the specific contents of the main interface 500 and the voice dialogue logo 501 can refer to the relevant description in the embodiment shown in FIG5C above. In addition, the specific process of the central control screen 10 displaying the voice dialogue logo 501 in response to the wake-up word issued by user A can also refer to the relevant contents in the embodiments shown in FIG5A to FIG5C above, which will not be repeated here. Optionally, a sound zone indicator 505 can also be displayed in the main interface 500. The relevant contents of the sound zone indicator 505 can refer to the relevant contents in the embodiment shown in FIG5B above, which will not be repeated here.

[0234] As shown in FIG6B , the co-pilot screen 20 of the smart car 100 displays a main interface 600 , and the main interface 600 may display wallpaper, such as a landscape picture.

[0235] While the central control screen 10 is interacting with user A via voice, the smart car 100 may also receive a wake-up-free command from user B in audio zone b. In response to the wake-up-free command from user B in audio zone b, the smart car 100 may display a wake-up-free indicator 601 on the main interface 600 of the passenger screen 20 of the smart car 100, as shown in FIG6C , if the display screen corresponding to audio zone b is the passenger screen 20.

[0236] As shown in Figure 6C, the wake-up-free mark 601 may include a rounded rectangular frame 602 and a wake-up-free symbol 603. Optionally, text 604 may also be displayed in the wake-up-free mark 601. The wake-up-free mark 601 can be used to prompt the user that the display screen B is receiving and processing the user's wake-up-free instruction. In the wake-up-free mark 601, the wake-up-free symbol 603 and the text 604 can be located inside the rounded rectangular frame 602. The wake-up-free symbol 603 can be a spherical symbol as shown in Figure 6C, or a symbol of other shapes and colors, which is not limited in this application. Moreover, the wake-up-free symbol 603 is different from the dialogue symbol 503 shown in Figure 5B above. The text 604 can be the text content of the wake-up-free instruction issued by user B, such as "turn on the air conditioner". In an embodiment of the present application, the wake-up-free mark 601 shown in Figure 6C can also be referred to as a wake-up-free capsule.

[0237] It should be noted that, since the audio zone corresponding to the co-pilot screen 20 is only the co-pilot audio zone, that is, the co-pilot screen 20 only needs to process the voice from the co-pilot seat, therefore, in some embodiments, the co-pilot screen 20 may not display the audio zone indicator.

[0238] In some embodiments, after the smart car 100 recognizes the wake-up-free instruction issued by user B (for example, "turn on the air conditioner"), the smart car 100 can change the text 604 in the wake-up-free indicator 601 to text 605 while executing the operation corresponding to the wake-up-free instruction (or after executing the operation), as shown in Figure 6D.

[0239] As shown in Figure 6D, a wake-up-free mark 601 is displayed on the main interface 600, and text 605 is displayed in the wake-up-free mark 601. The text 605 can be used to prompt the user that the wake-up-free command responded by the co-pilot screen 20 has been executed. For example, the text 605 can be "The air conditioner has been turned on for you."

[0240] In some embodiments, while the passenger screen 20 displays text 605 in the wake-up-free indicator 601, the passenger screen 20 may also play a voice message to prompt the user that the wake-up-free instruction to which the passenger screen 20 responded has been executed. For example, the content of the played voice message may be the content of text 605 shown in FIG. 6D .

[0241] It is understood that the embodiment shown in FIG6D is merely an example of outputting feedback information via a display screen. In the embodiment of the present application, the smart car 100 may also output feedback information via one or more methods, such as voice announcement, flashing indicator lights (e.g., flashing indicator lights on the display screen, flashing door panel ambient lights, etc.), display screen display, vibration, etc. In another possible implementation, executing the operation corresponding to the user's voice command may also be considered a method of outputting feedback information, which is not limited in the present application.

[0242] After the smart car 100 executes the wake-up-free command of user B, or after the smart car 100 outputs feedback information (such as playing audio, outputting the above text 605), the smart car 100 can control the co-pilot screen 20 to stop displaying the wake-up-free mark 601 and redisplay the main interface 600 shown in Figure 6B above.

[0243] It should be noted that, while the co-pilot screen 20 is receiving and processing the wake-up-free command from user B, the central control screen 10 can conduct voice interaction with user A through the voice assistant. Before the central control screen 10 meets the shutdown conditions, the central control screen 10 can always display the voice dialogue logo 501 in the main interface 500. Whether the co-pilot screen 20 responds to the wake-up-free command will not affect the display content of the central control screen 10.

[0244] It is understandable that the embodiment shown in Figures 6A to 6D above is only an example. In the embodiment of the present application, the user can also wake up the voice assistant of other display screens (such as the co-pilot screen 20, the right rear projection screen 30, the laser screen 40, etc.) through the wake-up word, and the display screen B can also be other display screens other than the co-pilot screen 20 (such as the right rear projection screen 30, the laser screen 40, etc.). In addition, the interface 1 displayed on the display screen A can also be an interface other than the main interface 500, and the display screen 2 displayed on the display screen B can also be an interface other than the main interface 600. The voice dialogue mark and the wake-up-free mark can also adopt different shapes, symbols, etc. from the above embodiments. The voice dialogue mark and the wake-up-free mark can also include different content from the above embodiments, and this application does not limit them here. In addition, the voice command issued by user A and the wake-up-free command issued by user B can also be voice commands different from the above embodiments, and this application does not limit them here.

[0245] In some application scenarios, the smart car 100 carries user A and user B, user A is in sound zone a, and user B is in sound zone b. When the display screen A of the smart car 100 displays interface 1 and the display screen B displays interface 2, the smart car 100 receives a wake-up-free command issued by user A from sound zone a. In response to the wake-up-free command issued by user A from sound zone a, the smart car 100 can display a wake-up-free mark 1 on interface 1 of display screen A. The wake-up-free mark 1 is used to prompt the user that display screen A is receiving and processing the user's wake-up-free command. When display screen A displays the wake-up-free mark 1, the smart car 100 can receive and respond to the wake-up-free command issued by the user from sound zone b. When it is determined that the display screen corresponding to sound zone b is display screen B, the smart car 100 can display a wake-up-free mark 2 on interface 2 of display screen B. The wake-up-free mark 2 is used to prompt the user that display screen B is receiving and processing the user's wake-up-free command.

[0246] In this way, the smart car 100 can simultaneously receive and process wake-up commands from different users, conduct voice interaction with multiple users through different display screens, and provide users with more convenient services.

[0247] For example, the central control screen 10 of the smart car 100 may display a main interface 700 as shown in FIG. 7A , and the right rear projection screen 30 may display a main interface 710 as shown in FIG. 7B .

[0248] As shown in FIG7A , the central control screen 10 may display a main interface 700. The main interface 700 may display one or more application icons, such as a music application icon, a navigation application icon, a setting application icon, and the like.

[0249] As shown in FIG7B , the right rear projection screen 30 of the smart car 100 displays a main interface 710 , and the main interface 710 may display wallpaper, such as a landscape picture.

[0250] The smart car 100 can receive and respond to the wake-up-free command issued by user A from sound zone a, such as "open the car window". If the smart car 100 determines that the display screen corresponding to sound zone a is the central control screen 10, the smart car 100 can display the wake-up-free logo 701 as shown in Figure 7C in the main interface 700 of the central control screen 10.

[0251] As shown in Figure 7C, the main interface 700 includes a wake-up-free mark 701. The wake-up-free mark 701 may include a rounded rectangular frame 702 and a wake-up-free symbol 703. Optionally, it may also include text 704. The wake-up-free mark 701 is used to prompt the user that the central control screen 10 is receiving and processing the user's wake-up-free instruction. In the wake-up-free mark 701, the wake-up-free symbol 703 and text 704 may be displayed in a rounded rectangular frame 702. The specific content of the wake-up-free symbol 703 can refer to the relevant description of the wake-up-free mark 601 shown in Figure 6C above, which will not be repeated here. Text 704 may include the text content of the wake-up-free instruction issued by user A, such as "open the car window". Further optionally, a sound zone indicator 705 may also be displayed in the main interface 700. The sound zone indicator 705 can be used to indicate the sound zone in which the user (ie, user A) who is interacting with the central control screen 10 by voice is located. The display area of ​​the sound zone indicator 705 shown in Figure 7C includes the lower left area of ​​the wake-up-free mark 701. The sound zone indicator 705 can be used to indicate that the sound zone a where user A is located is the sound zone on the left side of the central control screen 10 (for example, the main driving sound zone).

[0252] While the central control screen 10 is processing user A's wake-up-free command, the smart car 100 may also receive a wake-up-free command from user B in audio zone b, such as "turn on the air conditioner." In response to the wake-up-free command from user B in audio zone b, if the display corresponding to audio zone b is determined to be the right rear projection screen 30, the smart car 100 may display a wake-up-free indicator 711 on the main interface 710 of the right rear projection screen 30, as shown in FIG7D .

[0253] As shown in FIG7D , the main interface 710 displays a wake-up-free indicator 711. The wake-up-free indicator 711 may include a rounded rectangular frame 712 and a wake-up-free symbol 713, and optionally, text 714. The specific content of the wake-up-free indicator 711 can be referred to the relevant description of the wake-up-free indicator 601 shown in FIG6C above, and will not be repeated here. The wake-up-free indicator 711 can be used to indicate that the right rear projection screen 30 is receiving and processing the user's wake-up-free command.

[0254] It should be noted that since the sound zone corresponding to the right rear projection screen 30 is only the second row right sound zone, that is, the right rear projection screen 30 only needs to process the voice from the second row right seats, therefore, in some embodiments, the right rear projection screen 30 may not display the sound zone indicator.

[0255] In some embodiments, after executing the operation corresponding to the wake-up-free instruction, the smart car 100 can control the central control screen 10 and the right rear projection screen 30 to output feedback information. The specific content of the feedback information can refer to the relevant description in the embodiment shown in Figure 6D above, and will not be repeated here.

[0256] In addition, after executing the operation corresponding to the wake-up-free instruction, the smart car 100 can also control the display screen that responds to the wake-up-free instruction (such as the central control screen 10, the right rear projection screen 30, etc.) to stop displaying the wake-up-free logo. This application does not limit this.

[0257] It is understandable that the embodiment shown in Figures 7A to 7D above is only an example. In the embodiment of the present application, the user can also perform voice interaction with other display screens (such as the co-pilot screen 20, the laser screen 40, etc.) through the wake-up-free instruction, and the interface 1 displayed on the display screen A can also be an interface other than the main interface 700, and the interface 2 displayed on the display screen B can also be an interface other than the main interface 710. This application does not limit this. In addition, the wake-up-free logo can also adopt a shape and / or symbol different from that of the above embodiment, or include content different from that of the above embodiment. Moreover, the wake-up-free instruction issued by user A and user B can also be a wake-up-free instruction different from that of the above embodiment. This application does not limit this.

[0258] In some application scenarios, the smart car 100 carries user A and user B, user A is in audio zone a, and user B is in audio zone b. When the display screen A of the smart car 100 displays interface 1, the smart car 100 receives a wake-up-free command issued by user A from audio zone a. In response to the wake-up-free command issued by user A from audio zone a, the smart car 100 may display a wake-up-free mark 1 on interface 1 of display screen A. The wake-up-free mark 1 is used to prompt the user that display screen A is receiving and processing the user's wake-up-free command. When display screen A displays the wake-up-free mark 1, the smart car 100 may receive and respond to the wake-up-free command issued by the user from audio zone b. When it is determined that the display screen corresponding to audio zone b is display screen A, the smart car 100 may display a composite wake-up-free mark on interface 1 of display screen A. The composite wake-up-free mark is used to prompt the user that display screen A is receiving and processing wake-up-free commands issued by multiple users.

[0259] In this way, the smart car 100 can simultaneously receive and process wake-up commands from different users, conduct voice interaction with multiple users through the same display screen, and provide users with more convenient services.

[0260] Exemplarily, the smart car 100 can receive and respond to the wake-up-free command issued by user A from sound zone a. When it is determined that the display screen corresponding to sound zone a is the central control screen 10, as shown in Figure 8A, a wake-up-free mark 801 can be displayed in the main interface 800 of the central control screen 10. The wake-up-free mark 801 is used to prompt the user that the central control screen 10 is receiving and processing the wake-up-free command issued by the user. The specific content of the main interface 800 and the wake-up-free mark 801 can refer to the relevant description in the embodiment shown in Figure 7C above. In addition, the specific process of the central control screen 10 receiving and displaying the wake-up-free mark 801 in response to the wake-up-free command issued by user A can also refer to the relevant description in the embodiments shown in Figures 7A and 7C above, which will not be repeated here. Optionally, a sound zone indicator 802 can also be displayed in the main interface 800. The sound zone indicator 802 can be used to indicate the sound zone in which the user (i.e., user A) who is interacting with the central control screen 10 by voice is located. The display area of ​​the sound zone indicator 802 shown in Figure 8A includes the lower left area of ​​the composite wake-up-free logo 810. The sound zone indicator 802 can be used to indicate that the sound zone a where user A is located is the sound zone on the left side of the central control screen 10 (for example, the second row left sound zone, etc.).

[0261] While the central control screen 10 is processing user A's wake-up-free command, the smart car 100 may also receive a wake-up-free command from user B in audio zone b, such as "turn on the air conditioner." In response to the wake-up-free command from user B in audio zone b, if the display corresponding to audio zone b is determined to be the central control screen 10, the smart car 100 may display a composite wake-up-free indicator 810 on the main interface 800 of the central control screen 10, as shown in FIG8B .

[0262] As shown in Figure 8B, a composite wake-up-free indicator 810 is displayed in the main interface 800. The composite wake-up-free indicator 810 is used to prompt the user that the central control screen 10 is processing the wake-up-free instructions of multiple users. The composite wake-up-free indicator 810 may include a rounded rectangular frame 811, a composite wake-up-free symbol 812, and optionally, may also include text 813 and / or a numeric indicator 814. The composite wake-up-free symbol 812, text 813, and numeric indicator 814 are all displayed inside the rounded rectangular frame 811. The composite wake-up-free symbol 812 can be used to prompt the user that the central control screen 10 is processing the wake-up-free instructions of multiple users. The composite wake-up-free symbol 812 can be a composite of two overlapping wake-up-free symbols, or a symbol of other shapes and colors, which is not limited in this application. The text 813 can be used to prompt the user that the central control screen 10 is processing multiple wake-up-free instructions. For example, the text 813 can be "Multi-path execution". In other embodiments, the text 813 can also display multiple wake-up-free instructions being processed in a rotating manner. The digital indicator 814 can be used to prompt the user the number of wake-up-free instructions being processed by the central control screen 10. For example, the digital indicator 814 can be "2", which is used to prompt the user that the current central control screen 10 is processing two wake-up-free instructions at the same time. Further optionally, a sound zone indicator 815 can also be displayed in the main interface 800. The sound zone indicator 815 can be used to indicate the sound zone in which the users (i.e., user A and user B) are interacting with the central control screen 10 by voice. The display area of ​​the sound zone indicator 815 shown in Figure 8B includes the lower left area of ​​the composite wake-up-free logo 810. The sound zone indicator 815 can be used to indicate that the sound zone a where user A is located and the sound zone b where user B is located are both sound zones on the left side of the central control screen 10 (for example, the second row left sound zone, the third row left sound zone, etc.).

[0263] In some embodiments, after the smart car 100 executes the wake-up-free command issued by user A and the wake-up-free command issued by user B (for example, the smart car 100 turns on the air conditioner and opens the windows), the smart car 100 can control the central control screen 10 to stop displaying the composite wake-up-free indicator 810 and redisplay the main interface 800 shown in Figure 8C. The main interface 800 does not include the wake-up-free indicator, the composite wake-up-free indicator, or the audio zone indicator.

[0264] It is understood that the embodiment shown in Figures 8A to 8C above is only an example. In the embodiment of the present application, the user can also use the wake-up-free instruction to perform voice interaction with other display screens (such as the laser screen 40, etc.), and the interface 1 displayed on the display screen A can also be an interface other than the main interface 800. This application does not limit this. In addition, the wake-up-free mark can also use a shape and / or symbol different from that of the above embodiment, or include content different from that of the above embodiment. Moreover, the wake-up-free instruction issued by user A and user B can also be a wake-up-free instruction different from that of the above embodiment. This application does not limit this.

[0265] In other embodiments, one or more display screens of the smart car 100 (such as the central control screen 10, the laser curtain 40, etc.) can also simultaneously receive and process three or more wake-up-free instructions. In this case, the display screen can indicate the number of wake-up-free instructions processed simultaneously by the display screen through a digital indicator displayed in the composite wake-up-free identifier. This application does not limit this.

[0266] In some application scenarios, the smart car 100 carries user A and user B, user A is in audio zone a, and user B is in audio zone b. When the display screen A of the smart car 100 displays interface 1 and the display screen B displays interface 2, the smart car 100 receives a wake-up-free command issued by user A from audio zone a. In response to the wake-up-free command issued by user A from audio zone a, the smart car 100 can display a wake-up-free mark 1 on interface 1 of display screen A. The wake-up-free mark 1 is used to prompt the user that display screen A is receiving and processing the user's wake-up-free command. When display screen A displays the wake-up-free mark 1, the smart car 100 can receive and respond to the wake-up word issued by the user from audio zone b. When it is determined that the display screen corresponding to audio zone b is display screen B, the smart car 100 can determine display screen B as the main display screen and display a voice dialogue mark 1 on interface 2 of display screen B. The voice dialogue mark 1 is used to prompt the user that the voice assistant of display screen B has been turned on by the wake-up word.

[0267] In this way, in the process of receiving and processing the user's wake-up-free command, another user can wake up the voice assistant of another display screen through the wake-up word. The smart car 100 can conduct voice interaction with multiple users through multiple display screens, providing users with more convenient services.

[0268] For example, when the smart car 100 receives a wake-up-free command from user A from sound zone a and determines that the display screen corresponding to sound zone a is the right rear projection screen 30, the right rear projection screen 30 can display the main interface 900 shown in Figure 9A below, and the central control screen 10 can display the main interface 910 shown in Figure 9B below.

[0269] As shown in FIG9A , the right rear projection screen 30 of the smart car 100 displays a main interface 900, which may include a wallpaper, such as a landscape image. A wake-up-free indicator 901 may be displayed in the main interface 900. The specific content of the wake-up-free indicator 901 can refer to the relevant description in the embodiment shown in FIG7C above. Furthermore, the specific process of the smart car 100 receiving and responding to the wake-up-free instruction issued by user A and displaying the wake-up-free indicator 901 can also refer to the relevant content in the embodiments shown in FIG7A and FIG7C above, and will not be repeated here.

[0270] As shown in FIG9B , the central control screen 10 may display a main interface 910. The main interface 910 may display one or more application icons, such as a music application icon, a navigation application icon, a setting application icon, and the like.

[0271] While the right rear projection screen 30 is processing user A's wake-up command, the smart car 100 can also receive a wake-up word from user B in audio zone b. In response to the wake-up word from user B in audio zone b, if the display corresponding to audio zone b is determined to be the central control screen 10, the smart car 100 can determine the central control screen 10 as the primary display and display a voice dialogue indicator 911 on the main interface 910 of the central control screen 10, as shown in FIG9C .

[0272] As shown in FIG9C , a voice dialogue indicator 911 is displayed on the main interface 910. Voice dialogue indicator 911 is used to prompt the user that the voice assistant of the central control screen 10 has been activated by the wake-up word. The specific content of voice dialogue indicator 911 can refer to the relevant description of voice dialogue indicator 501 in the embodiment shown in FIG5B above, and will not be repeated here. Optionally, a sound zone indicator 912 can also be displayed on the main interface 910. The specific content of sound zone indicator 912 can refer to the relevant content in the embodiment shown in FIG5B above, and will not be repeated here.

[0273] After the voice assistant is activated on the central control screen 10, the smart car 100 can receive and respond to a voice command (e.g., "Play a song") issued by user B in audio zone b. As shown in FIG9D , the central control screen 10 can display text 913 in the voice dialogue indicator 911. Text 913 can be the text content of the user's voice command, e.g., "Play a song." Text 913 can be used to notify the user that the central control screen 10 is receiving and processing the user's voice command.

[0274] It should be noted that after the smart car 100 executes the wake-up-free command issued by user A, for example, after opening the car window, the smart car 100 can control the right rear projection screen 30 to stop displaying the wake-up-free mark 901. Optionally, the smart car 100 can also control the right rear projection screen 30 to output feedback information. The specific content of the feedback information can be referred to the relevant description of the embodiment shown in Figure 6D above, and will not be repeated here. In addition, when the smart car 100 detects that the central control screen 10 meets the shutdown conditions, it can also turn off the voice assistant of the central control screen 10 and stop displaying the voice dialogue mark.

[0275] It is understandable that the embodiment shown in Figures 9A to 9D above is only an example. In the embodiment of the present application, the user can also use the wake-up-free instruction to perform voice interaction with other display screens (such as the central control screen 10, etc.), and the interface 1 displayed on the display screen A can also be an interface other than the main interface 900. This application does not limit this. In addition, the wake-up-free logo and voice dialogue logo can also use different shapes and / or symbols from the above embodiment, or include different content from the above embodiment. Moreover, the wake-up-free instructions issued by user A and user B can also be different from the wake-up-free instructions in the above embodiment. This application does not limit this.

[0276] In some application scenarios, a smart car 100 carries user A and user B, with user A located in audio zone a and user B located in audio zone b. When display screen A of the smart car 100 displays interface 1, the smart car 100 receives a wake-up-free command issued by user A from audio zone a. In response to the wake-up-free command issued by user A from audio zone a, the smart car 100 may display a wake-up-free indicator 1 on interface 1 of display screen A. The wake-up-free indicator 1 is used to prompt the user that display screen A is receiving and processing the user's wake-up-free command. When display screen A displays the wake-up-free indicator 1, the smart car 100 may receive and respond to a wake-up word issued by the user from audio zone b. Upon determining that the display screen corresponding to audio zone b is display screen A, the smart car 100 may split the wake-up-free indicator 1 on interface 1 into a voice dialogue indicator and a voice sub-identifier. The voice dialogue indicator and the voice sub-identifier may be used to prompt the user that display screen A is processing the voices of multiple users simultaneously and that one of the users has activated the voice assistant on display screen A using a wake-up word.

[0277] In this way, the smart car 100 can simultaneously receive and process voice commands from different users, conduct voice interaction with multiple users through the same display screen, and provide users with more convenient services.

[0278] For example, when the smart car 100 receives a wake-up-free command from user A from sound zone a and determines that the display screen corresponding to sound zone a is the central control screen 10, the smart car 100 controls the central control screen 10 to display the main interface 1010 as shown in Figure 10A.

[0279] As shown in Figure 10A, the central control screen 10 can display a main interface 1010, and a wake-up-free mark 1011 can be displayed in the main interface 1010. The wake-up-free mark 1011 is used to prompt the user that the central control screen 10 is receiving and processing the user's wake-up-free command. Among them, the relevant contents of the main interface 1010 and the wake-up-free mark 1011 can refer to the relevant contents of the embodiment shown in Figure 5A and the embodiment shown in Figure 7C above, respectively, and will not be repeated here. In addition, the specific process of the central control screen 10 displaying the wake-up-free mark 1011 in response to the wake-up-free command issued by user A can also refer to the relevant contents of the embodiments shown in Figures 7A to 7C above, and will not be repeated here. Optionally, the main interface 1010 can also display a sound zone indicator 1012, and the sound zone indicator 1012 is used to indicate the positional relationship between the sound zone a where user A is located and the central control screen 10. For example, the display area of ​​the audio zone indicator 1012 shown in FIG10A includes the lower left portion of the wake-up-free logo 1011 , indicating that audio zone a is located on the left side of the central control screen 10 (eg, the second row left audio zone, etc.).

[0280] While the central control screen 10 is processing user A's wake-up command, the smart car 100 may also receive a wake-up word from user B in audio zone b. In response to the wake-up word from user B in audio zone b, if the display corresponding to audio zone b is determined to be the central control screen 10, the smart car 100 may determine the central control screen 10 as the primary display and display a split indicator 1020 as shown in FIG10B on the main interface 1010 of the central control screen 10.

[0281] As shown in Figure 10B, a split identifier 1020 is displayed in the main interface 1010. The split identifier 1020 may include two voice identifiers that are not completely split, one of which is a voice dialogue identifier 1021, and the other is a voice sub-identifier 1022. There may also be a connection channel between the voice dialogue identifier 1021 and the voice sub-identifier 1022, indicating that the two voice identifiers are still in the process of splitting and have not yet been completely split. The split identifier 1020 can be used to indicate that the central control screen 10 is processing the voices of multiple users at the same time. Among them, the specific content of the voice dialogue identifier 1021 can refer to the relevant description of the voice dialogue identifier 501 in the embodiment shown in Figure 5B above, and will not be repeated here. The voice sub-identifier 1022 may include a rounded rectangular frame and a sub-symbol 1024. The sub-symbol 1024 may be the same as the sub-symbol 531 shown in Figure 5D above. Optionally, a sound zone indicator 1023 can also be displayed in the main interface 1010. The display area of ​​the sound zone indicator 1023 is the area below the split mark 1020 (including the lower left area, the area directly below, and the lower right area). The sound zone indicator 1023 is used to indicate that the sound zones of the multi-channel voice being processed by the central control screen 10 are located at different positions of the central control screen 10, that is, the sound zone a where user A is located is the sound zone on the left side of the central control screen 10, and the sound zone b where user B is located is the sound zone on the right side of the central control screen 10 (for example, the right sound zone of the third row).

[0282] After the split identifier 1020 is split, as shown in Figure 10C, the main interface 1010 of the central control screen 10 can display the completely split voice dialogue identifier 1021 and voice sub-identifier 1022, that is, there is no connection channel between the voice dialogue identifier 1021 and the voice sub-identifier 1022.

[0283] It should be noted that after the smart car 100 executes the wake-up-free command issued by user A, for example, after opening the window, the smart car 100 can control the central control screen 10 to stop displaying the wake-up-free indicator 901. Optionally, the smart car 100 can also control the central control screen 10 to output feedback information. The specific content of the feedback information can be referred to the relevant description of the embodiment shown in Figure 6D above and will not be repeated here. In addition, when the smart car 100 detects that the central control screen 10 meets the shutdown conditions, it can also turn off the voice assistant on the central control screen 10 and stop displaying the voice dialogue indicator.

[0284] It can be understood that the embodiment shown in Figures 10A to 10C above is only an example. In the embodiment of the present application, the user can also wake up the voice assistant of other display screens (such as the co-pilot screen 20, the right rear projection screen 30, the laser curtain 40, etc.) through the wake-up word, and the interface 1 displayed on the display screen A can also be other interfaces other than the main interface. In addition, the wake-up-free mark, voice dialogue mark and voice sub-marker can also adopt shapes and / or symbols different from the above embodiment, or include content different from the above embodiment. In addition, the voice commands issued by user A and user B can also be voice commands different from the above embodiment, and the present application does not limit them here.

[0285] In some application scenarios, a smart car 100 carries user A and user B, with user A located in audio zone a and user B located in audio zone b. When screen A of the smart car 100 displays interface 1 and screen B displays interface 2, the smart car 100 receives a wake-up word from user A in audio zone a. In response to the wake-up word from user A in audio zone a, the smart car 100 may display a voice dialogue indicator 1 on interface 1 of screen A. Voice dialogue indicator 1 is used to prompt the user that the voice assistant on screen A has been activated using the wake-up word. While screen A displays voice dialogue indicator 1, the smart car 100 may receive and respond to the wake-up word from the user in audio zone b. Upon determining that the display corresponding to audio zone b is screen B, the smart car 100 may turn off voice dialogue indicator 1 displayed on screen A, identify screen B as the new primary display, and display voice dialogue indicator 2 on screen B. Voice dialogue indicator 2 is used to prompt the user that the voice assistant on screen B has been activated using the wake-up word.

[0286] In this way, when the smart car 100 receives the wake-up word from another user, it can interrupt the voice interaction between the user who previously used the wake-up word and the display screen, ensuring that at the same time, among the voices that need to be processed in the smart car 100, only one voice is not a wake-up-free command.

[0287] For example, when the smart car 100 receives the wake-up word from user A in audio zone a and determines that the display screen corresponding to audio zone a is the central control screen 10, the smart car 100 controls the central control screen 10 to display the main interface 1010 shown in FIG10A. Simultaneously, the right rear projection screen 30 may display the main interface 1110 shown in FIG11B below.

[0288] As shown in FIG11A , the central control screen 10 may display a main interface 1100, and a voice dialogue logo 1101 may be displayed in the main interface 1100. The voice dialogue logo 1101 is used to prompt the user that the voice assistant of the central control screen 10 has been turned on by the wake-up word, and that voice interaction with the central control screen 10 is being performed through the voice assistant. Text 1103 may also be displayed in the voice dialogue logo 1101. Text 1103 may be part or all of the text content of the voice command issued by user A after issuing the wake-up word, such as “The weather is good today, open the car window”. Optionally, a sound zone indicator 1102 may also be displayed in the main interface 1100. The specific content of the sound zone indicator 1102 may refer to the relevant content in the embodiment shown in FIG5B above, and will not be repeated here.

[0289] As shown in FIG11B , the right rear projection screen 30 of the smart car 100 displays a main interface 1110 , and the main interface 1110 may display wallpaper, such as a landscape picture.

[0290] While the central control screen 10 is processing the voice command issued by user A after the wake-up word, the smart car 100 may also receive the wake-up word issued by user B in audio zone b. In response to the wake-up word issued by user B in audio zone b, if the display corresponding to audio zone b is determined to be the right rear projection screen 30, the smart car 100 may determine the right rear projection screen 30 as the new primary display, as shown in FIG11C , control the central control screen 10 to stop displaying the voice dialogue icon 1101, and control the right rear projection screen 30 to display the voice dialogue icon 1111 in the main interface 1110, as shown in FIG11D .

[0291] As shown in Figure 11C, the central control screen 10 displays a main interface 1100, and the main interface 1100 does not include a voice dialogue logo 1101. Optionally, while stopping displaying the voice dialogue logo 1101, the central control screen 10 may also display an interruption prompt 1103 on the main interface 1100. The interruption prompt 1103 can be used to prompt the user that the current voice interaction has been interrupted. For example, the text "Voice interaction interrupted" may be displayed in the interruption prompt 1103. In other embodiments, the central control screen 10 may also output the interruption prompt in one or more ways such as voice broadcast, vibration, and flashing indicator light, which is not limited in this application.

[0292] As shown in Figure 11D, the right rear projection screen 30 displays a main interface 1110, and a voice dialogue logo 1111 is displayed in the main interface 1110. The voice dialogue logo 1111 can be used to prompt the user that the voice assistant of the right rear projection screen 30 has been turned on through the wake-up word. The specific content of the voice dialogue logo 1111 can refer to the relevant description in the embodiment shown in Figure 5B above, and will not be repeated here.

[0293] It should be noted that, upon detecting that the display duration of the interruption prompt 1103 exceeds a preset duration (e.g., 5 seconds), the smart car 100 may stop displaying the interruption prompt 1103. Furthermore, upon detecting that the right rear projection screen 30 meets the shutdown condition, the smart car 100 may also shut down the voice assistant on the right rear projection screen 30 and stop displaying the voice dialogue indicator.

[0294] It can be understood that the embodiment shown in Figures 11A to 11D above is only an example. In the embodiment of the present application, the user can also use the wake-up word to turn on the voice assistant of other display screens (such as the co-pilot screen 20, the laser curtain 40, etc.), and the interface 1 displayed on the display screen A and the interface 2 displayed on the display screen B can also be other interfaces other than the main interface. In addition, the voice dialogue identifier and the interruption prompt can also adopt different shapes and / or symbols from the above embodiment, or include different content from the above embodiment. In addition, the voice commands issued by user A and user B can also be different from the voice commands of the above embodiment, and the present application does not limit them here.

[0295] In some application scenarios, the smart car 100 carries user A and user B, user A is located in audio zone a, and user B is located in audio zone b. When screen A of the smart car 100 displays interface 1 and screen B displays interface 2, the smart car 100 receives a wake-up word from user A in audio zone a. In response to the wake-up word from user A in audio zone a, the smart car 100 may display a voice dialogue identifier 1 on interface 1 of screen A. Voice dialogue identifier 1 is used to prompt the user that the voice assistant on screen A has been activated using the wake-up word. When screen A displays voice dialogue identifier 1, the smart car 100 may receive and respond to the wake-up word from the user in audio zone b. If it is determined that the display corresponding to audio zone b is screen A, the smart car 100 may turn off voice dialogue identifier 1 displayed on screen A and display voice dialogue identifier 2 on screen A. Voice dialogue identifier 2 is used to prompt the user that the voice assistant on screen A has been activated using the wake-up word.

[0296] In this way, when the smart car 100 receives the wake-up word from another user, it can interrupt the voice interaction between the user who previously used the wake-up word and the display screen, ensuring that at the same time, among the voices that need to be processed in the smart car 100, only one voice is not a wake-up-free command.

[0297] For example, when the smart car 100 receives the wake-up word issued by user A from sound zone a and determines that the display screen corresponding to sound zone a is the central control screen 10, the smart car 100 controls the central control screen 10 to display the main interface 1010 as shown in Figure 10A.

[0298] As shown in FIG12A , the central control screen 10 may display a main interface 1200, in which a voice dialogue logo 1201 may be displayed. The voice dialogue logo 1201 is used to prompt the user that the voice assistant of the central control screen 10 has been activated through the wake-up word and that voice interaction with the central control screen 10 is being performed through the voice assistant. The voice dialogue logo 1201 may also display text 1203. The text 1203 may be part or all of the text content of the voice command issued by user A after issuing the wake-up word, such as “The weather is nice today, open the windows”. Optionally, a sound zone indicator 1202 may also be displayed in the main interface 1200. The display area of ​​the sound zone indicator 1202 may be the lower left corner of the voice dialogue logo 1201, to prompt the user that the sound zone a in which user A is located is on the left side of the central control screen 10.

[0299] While the central control screen 10 is processing the voice command issued by user A after the wake-up word, the smart car 100 can also receive the wake-up word issued by user B in audio zone b. In response to the wake-up word issued by user B in audio zone b, when it is determined that the display screen corresponding to audio zone b is the central control screen 10, as shown in Figure 12B, the smart car 100 can control the central control screen 10 to stop displaying the voice dialogue identifier 1201 and display the voice dialogue identifier 1211 on the main interface 1200. Alternatively, the smart car 100 can also control the central control screen 10 to update the content (e.g., text, etc.) in the voice dialogue identifier 1201, changing the voice dialogue identifier 1201 to the voice dialogue identifier 1211 shown in Figure 12B.

[0300] As shown in FIG12B , the central control screen 10 displays a main interface 1200, and the main interface 1200 displays a voice dialogue logo 1211. The voice dialogue logo 1211 may display text 1213, and the text 1213 is different from the text in the voice dialogue logo 1201 shown in FIG12A . For example, the text 1213 may be "You can speak at any time." Optionally, a sound zone indicator 1212 may also be displayed in the main interface 1200. The display area of ​​the sound zone indicator 1212 may be the lower right corner of the voice dialogue logo 1211, which is used to prompt the user B that the sound zone b is located on the right side of the central control screen 10. Further optionally, the central control screen 10 may also output an interruption prompt in one or more ways such as display screen display, voice broadcast, vibration, and flashing indicator light. The specific content of the interruption prompt can refer to the relevant description in the embodiment shown in FIG11C above, and this application will not repeat it here.

[0301] It should be noted that when the smart car 100 detects that the central control screen 10 meets the shutdown conditions, it can also turn off the voice assistant of the central control screen 10 and stop displaying the voice dialogue logo.

[0302] It can be understood that the embodiment shown in Figures 12A to 12B above is only an example. In the embodiment of the present application, the user can also use the wake-up word to turn on the voice assistant of other display screens (such as the co-pilot screen 20, the right rear projection screen 30, the laser curtain 40, etc.), and the interface 1 displayed on the display screen A can also be other interfaces other than the main interface. In addition, the voice dialogue identifier and the interruption prompt can also adopt different shapes and / or symbols from the above embodiment, or include different content from the above embodiment. In addition, the voice commands issued by user A and user B can also be different from the voice commands of the above embodiment, and the present application does not limit them here.

[0303] The following describes a process of a voice interaction method provided in an embodiment of the present application.

[0304] As shown in FIG13 , the specific process of the smart car 100 executing the voice interaction method may include the following steps:

[0305] S1301, the smart car 100 receives voice 1 sent by the user.

[0306] The smart car 100 can receive the user's voice through one or more microphones provided inside.

[0307] S1302 , the smart car 100 determines the sound zone 1 corresponding to the voice 1 .

[0308] In some embodiments, the smart car 100 may be equipped with microphones in different sound zones within the car, and the smart car 100 may store one or more sound source localization algorithms. When a user in the car (e.g., user A) speaks, because user A's position and distance relative to each microphone are different, the voice arrives at each microphone at different times and from different directions. The smart car 100 can use the sound source localization algorithm to determine the sound zone where user A is located based on the voice collected by multiple microphones. In this way, after collecting Voice 1, the smart car 100 can determine the sound zone where the user who issued Voice 1 is located based on the sound source localization algorithm. This sound zone is Voice Zone 1 corresponding to Voice 1.

[0309] In other embodiments, the smart car 100 may be provided with microphones in different sound zones within the car, and each microphone may have its own device identification, such as an identity document (ID) of the microphone. When a user in the car (e.g., user A) speaks, the microphone in the same sound zone as user A may collect the voice of user A and report the voice carrying the device identification of the microphone to the smart car 100. The smart car 100 may determine the microphone that reported the voice based on the device identification of the microphone, and thereby determine the sound zone corresponding to the voice based on the sound zone where the microphone is located. In this way, after collecting voice 1, the smart car 100 may determine the microphone that reported voice 1 based on voice 1, and further determine the sound zone 1 corresponding to voice 1.

[0310] It should be noted that in some embodiments, when multiple users in the vehicle are speaking, the microphone can determine whether the collected speech is from its local sound zone (i.e., the sound zone where the microphone is located) based on any one or more of the speech parameters such as the angle of arrival, volume, latency, voiceprint, and frequency. The microphone can then filter the collected signal to retain only the speech within its local sound zone and filter out speech from other sound zones.

[0311] For example, in smart car 100, microphone 1 is located in audio zone 1, and microphone 2 is located in audio zone 2. User A frequently sits in audio zone 1, and user B frequently sits in audio zone 2. After smart car 100 starts, user A speaks voice a, and user B also speaks voice b. Microphone 1 can capture both user A's voice a and user B's voice b, and microphone 2 can capture both user A's voice a and user B's voice b. In this case, microphone 1 can identify user A's voiceprint based on previously collected voice data and, based on user A's voiceprint, determine to retain voice a and filter voice b. Microphone 2 can similarly determine to retain voice b and filter voice a based on user B's voiceprint. Microphone 1 can then report voice a, carrying microphone 1's device identifier, to smart car 100, and microphone 2 can report voice b, carrying microphone 2's device identifier, to smart car 100. The smart car 100 can determine that the sound zone corresponding to voice a is sound zone 1 and the sound zone corresponding to voice b is sound zone 2 based on the device identifiers carried by voice a and voice b.

[0312] It can be understood that the above embodiments are merely illustrative of two ways of determining the sound zone corresponding to the speech. In the embodiments of the present application, the smart car 100 may also use more or different methods from the above embodiments to determine the sound zone corresponding to the speech, and the present application does not limit this.

[0313] S1303, the smart car 100 determines whether the voice 1 contains a wake-up word.

[0314] The smart car 100 may store one or more wake-up words, such as "Xiao A".

[0315] If the voice 1 includes a wake-up word, the smart car 100 may execute the following step S1304.

[0316] If the voice 1 does not include a wake-up word, the smart car 100 may execute the following step S1307.

[0317] S1304, the smart car 100 determines whether there is a main display screen, which is a display screen that is processing a non-wake-up-free command through a voice assistant.

[0318] The main display screen refers to the display screen that is processing a non-wake-up-free command through a voice assistant. A non-wake-up-free command refers to a voice command that is not a wake-up-free command. The non-wake-up-free command may include a wake-up word. According to the explanation of terms in the embodiment of the present application, only when the user turns on the voice assistant of the display screen through the wake-up word, the display screen can process the user's non-wake-up-free command through the voice assistant. Therefore, in the embodiment of the present application, the main display screen may also be the display screen on which the voice assistant is turned on by the user through the wake-up word, or the display screen that is triggered by the user to display a voice dialogue identifier through the wake-up word.

[0319] At the same time, there is at most one main display screen among the one or more display screens of the smart car 100. It should be noted that compared with non-wake-up-free commands, the execution process of the wake-up-free command is shorter and consumes fewer resources. The smart car 100 can be set to process multiple voices at the same time, and at most one voice is a non-wake-up-free command. In this way, the smart car 100 can ensure that each user's voice command can receive timely feedback. Therefore, at the same time, there is at most one main display screen in the smart car 100, and among the one or more voices processed by the main display screen, at most one voice is a non-wake-up-free command.

[0320] In some embodiments, upon receiving the wake-up word, the smart car 100 may determine the display screen corresponding to the audio zone corresponding to the wake-up word speech, determine the display screen corresponding to the audio zone, determine the display screen corresponding to the audio zone, and store the device identifier of the display screen as the device identifier of the main display screen. In this case, when executing step S1304, the smart car 100 may determine whether a main display screen currently exists based on whether the device identifier of the main display screen is stored.

[0321] In some embodiments, the smart car 100 can determine whether each display screen is processing a non-wake-up-free instruction. If there is a display screen that is processing a non-wake-up-free instruction, then the display screen is the main display screen, and the smart car 100 can determine that there is currently a main display screen; if there is no display screen that is processing a non-wake-up-free instruction, then the smart car 100 can determine that there is currently no main display screen.

[0322] In some embodiments, the smart car 100 can determine whether a voice dialogue logo is displayed on each display screen. If there is a display screen displaying a voice dialogue logo, the smart car 100 can determine that the display screen is the main display screen, that is, the main display screen currently exists; if there is no display screen displaying a voice dialogue logo, the smart car 100 can determine that there is no main display screen currently.

[0323] If the smart car 100 determines that a main display screen exists, the smart car 100 may execute the following step S1305 .

[0324] If the smart car 100 determines that the main display screen does not exist, the smart car 100 may execute the following step S1306.

[0325] S1305, the smart car 100 turns off the voice dialogue icon on the main display screen.

[0326] When the smart car 100 receives the wake-up word and a main display screen already exists, the smart car 100 can close the voice dialogue started by the previous user (for example, user C) through the wake-up word, that is, turn off the voice dialogue indicator of the main display screen and stop the voice interaction between the original main display screen and user C.

[0327] For example, referring to the embodiments shown in Figures 11A and 11C above, Voice 1 may be the wake-up word uttered by user B in the above embodiment. While the central control screen 10 displays the voice dialogue icon 1101 and the voice assistant receives and processes a user's non-wake-up-free command, the smart car 100 may deactivate the voice dialogue icon 1101 displayed on the central control screen 10 upon receiving a wake-up word uttered by another user.

[0328] As another example, referring to the embodiment shown in Figures 12A and 12B above, voice 1 may be the wake-up word issued by user B in the above embodiment. When the voice dialogue logo 1201 is already displayed on the central control screen 10 and the voice assistant receives and processes the user's non-wake-up command, the smart car 100 may turn off the voice dialogue logo 1201 displayed on the central control screen 10 when receiving the wake-up word issued by another user.

[0329] S1306: The smart car 100 displays a voice dialogue identifier on the display screen A corresponding to the audio zone 1.

[0330] The smart car 100 may store the correspondence between the audio zones and the display screens. For example, the correspondence between the audio zones and the display screens may refer to the relevant descriptions in the embodiments shown in FIG. 3E to FIG. 3F .

[0331] In some embodiments, the state of the display (e.g., on or off) may also affect the correspondence between the audio zones and the display. The smart car 100 may also store the correspondence between the display state, the audio zones, and the display. For example, the correspondence between the audio zones and the display in different states can be described in the embodiments shown in Figures 3E to 3F above, and will not be repeated here.

[0332] It can be understood that the above Figures 3E to 3H are just some examples. In the embodiments of the present application, the division method of the sound zones, the number of display screens, the setting method of the display screens, and the correspondence between the sound zones and the display screens may be different from the above embodiments, and the present application does not limit them here.

[0333] The smart car 100 can determine the display screen A corresponding to the sound zone 1 based on the correspondence between the sound zones and the display screens. In some embodiments, the smart car 100 can also determine the display screen A corresponding to the sound zone 1 based on the current display screen state and the correspondence between the sound zones, the display screen state, and the display screen.

[0334] When the smart car 100 receives the wake-up word and a main display screen currently exists, the smart car 100 can display the voice dialogue logo on the display screen A corresponding to the audio zone 1 while turning off the voice dialogue logo on the main display screen.

[0335] For example, referring to the embodiment shown in Figures 11A to 11D above, in the above embodiment, display screen A may be the right rear projection screen 30, and the voice dialogue identifier may be the voice dialogue identifier 1111 shown in Figure 11D . For another example, referring to the embodiment shown in Figures 12A to 12B above, in the above embodiment, display screen A may be the central control screen 10, and the newly displayed voice dialogue identifier may be the voice dialogue identifier 1211 shown in Figure 12B .

[0336] When the smart car 100 receives the wake-up word and no main display screen is currently present, the smart car 100 may display a voice dialogue identifier on display screen A corresponding to audio zone 1. For example, referring to the embodiment shown in Figures 5A and 5B above, display screen A may be the central control screen 10, and the voice dialogue identifier may be the voice dialogue identifier 501 shown in Figure 5B above.

[0337] In some embodiments, if display screen A is the same as the previous main display screen, the main display screen can also change the text displayed in the voice dialogue identifier. The newly displayed text can be used to prompt the user that the original voice interaction has been interrupted and a new round of voice interaction has started.

[0338] In some embodiments, the smart car 100 may display a voice zone indicator on display screen A while displaying a voice dialogue identifier on display screen A. The voice zone indicator is used to indicate the relative position of the source voice zone of the voice being processed by display screen A and display screen A.

[0339] It should be noted that in one possible implementation, the smart car 100 may display a voice zone indicator only on the main display screen, where the voice zone indicator indicates the relative position of the source voice zone of one or more voice channels currently being processed by the main display screen relative to the main display screen. In this case, the voice zone indicator can also be used to prompt the user to continue issuing voice commands and interacting with the main display screen. Furthermore, the voice zone indicator need not be displayed on the secondary display screen.

[0340] In another possible implementation, the smart car 100 may also determine whether to display a sound zone indicator on the display screen based on whether the display screen has the ability to process multiple voice channels. For example, when display screen A is a display screen that can process multiple voice channels, such as the central control screen 10 or the laser screen 40, display screen A may display a sound zone indicator on display screen A in response to a wake-up word or a wake-up-free command issued from sound zone 1 (or other corresponding sound zones); when display screen A is a display screen that only needs to process one voice channel, such as the passenger screen 20 or the right rear projection screen 30, display screen A may not display a sound zone indicator. For example, the sound zone indicator may be the sound zone indicator 505 shown in FIG. 5B above, or the sound zone indicator 912 shown in FIG. 9C, etc.

[0341] In another possible implementation, the smart car 100 may also display a sound zone indicator on the corresponding display screen when receiving a wake-up word or a wake-up-free command. In another possible implementation, the smart car 100 may not display a sound zone indicator on any display screen.

[0342] After executing step S1306, the smart car 100 can execute the following step S1309.

[0343] S1307, the smart car 100 determines whether the voice 1 is a wake-up-free command.

[0344] The smart car 100 may store one or more wake-up-free commands. The wake-up-free commands may be pre-set or determined during user use based on the frequency of use of different voice commands. For example, the smart car 100 may record voice commands used by the user and set voice commands used more than a preset number (e.g., five times) within a period of time (e.g., a week) as wake-up-free commands.

[0345] For example, Table 1 shows a plurality of wake-up-free instructions stored in a smart car 100 provided in an embodiment of the present application.

[0346] Table 1

[0347] As shown in Table 1, the smart car 100 may store one or more wake-up-free commands, such as "turn on the air conditioner", "open the window", "turn on the navigation and go home", and "play music".

[0348] It can be understood that the embodiment shown in Table 1 is only an example of how the smart car 100 can store one or more wake-up-free instructions. In the embodiment of the present application, the smart car 100 can store more or fewer wake-up-free instructions than those in Table 1 above, and this application does not limit this.

[0349] After recognizing the text content of voice 1, the smart car 100 can determine whether voice 1 is a wake-up-free command based on the text content of voice 1 and the stored wake-up-free commands. If the text content of voice 1 includes any wake-up-free command, the smart car 100 can determine that voice 1 is a wake-up-free command; if the text content of voice 1 does not include any wake-up-free command, the smart car 100 can determine that voice 1 is not a wake-up-free command.

[0350] If the voice 1 is a wake-up-free instruction, the smart car 100 may execute the following step S1308. Optionally, if the smart car 100 stores feedback information corresponding to the wake-up-free instruction, the smart car 100 may also output the feedback information.

[0351] If voice 1 does not belong to a wake-up-free command, the smart car 100 can execute the following step S1309.

[0352] S1308 , the smart car 100 displays a wake-up-free logo, a voice sub-logo, or a composite wake-up-free logo on the display screen A corresponding to the audio zone 1 .

[0353] The smart car 100 can determine the display screen A corresponding to the sound zone 1 based on the correspondence between the sound zone and the display screen. The specific determination method can refer to the relevant description in the above step S1306.

[0354] If Voice 1 is a wake-up-free command, the smart car 100 can control display screen A to display any of the following voice indicators: a wake-up-free indicator, a voice sub-indicator, or a combined wake-up-free indicator. The voice indicators displayed on display screen A may vary in different scenarios. The specific relationship between scenarios and voice indicators, as well as the method for determining scenarios, can be found in the embodiment shown in FIG. 14 below and will not be described in detail here.

[0355] S1309, the smart car 100 determines whether there is a main display screen and whether the voice 1 is issued by user A, and user A is the user who turns on the voice assistant of the main display screen through the wake-up word.

[0356] The specific method for the smart car 100 to determine whether there is a main display screen can refer to the relevant content in the above step S1304 and will not be repeated here.

[0357] In the absence of the main display screen, the smart car 100 may no longer process the voice 1.

[0358] When there is a main display screen, the smart car 100 can further determine whether voice 1 is issued by user A, who is the user who turns on the voice assistant of the main display screen through the wake-up word, that is, user A is the user who is interacting with the original main display screen by voice, and the voice command issued is not a wake-up-free command.

[0359] Smart car 100 can determine whether Voice 1 was sent by user A based on whether the voiceprint information of Voice 1 is consistent with the voiceprint information of user A. User A's voiceprint information can be obtained based on the voice commands (including wake-up words, etc.) collected during the previous voice interaction process. If the voiceprint information of Voice 1 and User A is consistent, smart car 100 can determine that Voice 1 was sent by user A; if the voiceprint information of Voice 1 is inconsistent with User A's voiceprint information, smart car 100 can determine that Voice 1 was not sent by user A.

[0360] If the smart car 100 determines that the main display screen does not exist, the smart car 100 may not process the voice 1.

[0361] If the smart car 100 determines that there is a main display screen and the voice 1 is not sent by user A, the smart car 100 may no longer process the voice 1.

[0362] If the smart car 100 determines that there is a main display screen and the voice 1 is issued by user A, the smart car 100 can execute the following step S1310.

[0363] In some embodiments, after step S1301, the smart car 100 may execute step S1309. In this case, if the smart car 100 determines that a main display screen exists and that Voice 1 is from user A, the smart car 100 may execute step S1310. If the smart car 100 determines that a main display screen does not exist or that Voice 1 is not from user A, the smart car 100 may execute step S1302.

[0364] S1310, the smart car 100 executes the operation corresponding to voice 1.

[0365] When voice 1 is a wake-up-free command, smart car 100 can perform the operation corresponding to the wake-up-free command. For example, if voice 1 includes the wake-up-free command "turn on the air conditioner," smart car 100 can turn on the air conditioner; another example, if voice 1 includes the wake-up-free command "play music," smart car 100 can play music, etc.

[0366] When the voice 1 includes a wake-up word and other voice commands, the smart car 100 can execute the voice command. For example, if the voice command is "play a song", the smart car 100 can play a song, etc.

[0367] When Voice 1 includes a wake-up word and no other voice commands, the smart car 100 can prompt the user to output a voice command through display A, or ask the user for the next voice command through voice. The operation of the smart car 100 prompting the user to output a voice command can be the operation corresponding to Voice 1 in this scenario.

[0368] When voice 1 includes a non-wake-up command and display A is the main display, the smart car 100 can execute the voice command. For example, if the voice command is "go home", the smart car 100 can start navigation and set the destination to "home".

[0369] In some embodiments, when executing the operation corresponding to voice 1, the smart car 100 can also interact with other electronic devices (such as a cloud server 400, a mobile phone 200, a headset 300, etc.).

[0370] For example, if the voice 1 is "play movie X", the smart car 100 can obtain the video data of movie X through the communication connection with the cloud server 400, and play movie X through the display screen A. For another example, if the voice 1 is "answer a call", the smart car 100 can send an answer instruction to the mobile phone 200 through the communication connection with the mobile phone 200, instructing the mobile phone 200 to answer the call. In this case, the smart car 100 can also obtain the call data of the mobile phone 200 through the communication connection between the smart car 100 and the mobile phone 200, and play the audio of the call through the speaker of the smart car 100 based on the call data. For another example, if the voice 1 is "play music through headphones", the smart car 100 can establish a communication connection with the headphones 300, send audio data to the headphones 300 through the communication connection with the headphones 300, and instruct the headphones 300 to play audio based on the audio data.

[0371] S1311, the smart car 100 outputs feedback information.

[0372] Step S1311 is an optional step.

[0373] While executing the operation corresponding to voice 1 (or after executing the operation corresponding to voice 1), the smart car 100 may also output feedback information by displaying on a screen, announcing by voice, flashing an indicator light, vibrating, or any one or more other methods. In another possible implementation, executing the operation corresponding to voice 1 may also be considered a method of outputting feedback information, which is not limited in this application.

[0374] In some embodiments, when the smart car 100 receives and recognizes the voice instructions contained in the voice 1 (such as wake-up-free instructions, wake-up words, non-wake-up-free instructions, etc.), it can output specified feedback information, such as "Received", "OK", "Execute for you immediately".

[0375] In other embodiments, the smart car 100 may also determine to output different feedback information based on the voice command in voice 1 during or after executing the operation corresponding to voice 1. The smart car 100 may store feedback information corresponding to the voice command. The voice command may include a wake-up-free command, a wake-up word, and a non-wake-up-free command. When receiving a voice command, the smart car 100 may not only activate the voice assistant, but also output the feedback information corresponding to the voice command by displaying on a screen and / or playing audio.

[0376] For example, Table 2 shows the correspondence between voice commands and feedback information stored in a smart car 100 provided in an embodiment of the present application.

[0377] Table 2

[0378] As shown in Table 2, the smart car 100 may store one or more voice commands, and may also store feedback information corresponding to the voice commands. For example, the feedback information corresponding to the voice command "Start navigation, go home" may be "Navigation started, destination: home"; the feedback information corresponding to the voice command "Play music" may be "Music will be played for you soon"; the feedback information corresponding to the voice command "Open the window" may be "Received, I will open the window for you immediately"; the feedback information corresponding to the wake-up word "Xiao A" may be "You can speak at any time", etc.

[0379] It can be understood that the embodiment shown in Table 2 is only an example of how the smart car 100 can store the correspondence between voice commands and feedback information. In the embodiment of the present application, the smart car 100 can store more or less voice commands and feedback information than those in Table 2 above, and can also store correspondences different from those in the embodiment shown in Table 2. This application does not limit this.

[0380] By using the voice interaction method provided in the embodiments of the present application, the smart car 100 can simultaneously process voice commands from multiple users and conduct voice interaction with multiple users through one or more display screens. Furthermore, when the smart car 100 processes multiple voice channels, since at most one voice channel is a non-wake-up-free command and the remaining one or more voice channels are wake-up-free commands, the smart car 100 can ensure that each user receives timely feedback, thereby improving the user experience.

[0381] It should be noted that the embodiment shown in Figure 13 above is only an example of how the smart car 100 can perform different operations in different scenarios. In the embodiment of the present application, when the smart car 100 receives the user's voice 1, it can also use an execution order different from the embodiment shown in Figure 13 above to execute the voice interaction method. For example, first determine whether there is a main display screen, then determine whether voice 1 is a wake-up word-free voice, then determine whether voice 1 contains a wake-up word, and then determine the sound zone a corresponding to voice 1, etc., or simultaneously determine whether there is a main display screen, whether voice 1 is a wake-up word-free voice, whether voice 1 contains a wake-up word, and the sound zone a corresponding to voice 1, etc. This application does not limit the specific execution order of each judgment step.

[0382] In some embodiments, when a voice conversation indicator is displayed on display screen A, the smart car 100 may stop displaying the voice conversation indicator upon detecting that a shutdown condition has been met. The shutdown conditions may include, but are not limited to, any one or more of the following: receiving a wake-up word, the user currently interacting with display screen A via a non-wake-up-free voice command not speaking for a period exceeding a preset time, receiving a shutdown instruction, or receiving a user shutdown operation. The shutdown instruction may be a preset voice command, such as "end conversation." The user shutdown operation may be a user pressing a designated button or a control on display screen A, etc.

[0383] In some embodiments, when the display screen A displays a wake-up-free logo, the smart car 100 can stop displaying the wake-up-free logo after the wake-up-free instruction that triggers the display of the wake-up-free logo is executed.

[0384] In some embodiments, when a composite wake-up-free indicator is displayed on display screen A, the smart car 100 may change the digital indicator in the composite wake-up-free indicator after detecting that one or more wake-up-free instructions among the multiple wake-up-free instructions being processed have been executed.

[0385] In some embodiments, when the display screen A displays a voice dialogue identifier and a voice sub-identifier, the smart car 100 may stop displaying the voice sub-identifier after detecting that all currently executed wake-up-free indicators have been completed.

[0386] The following describes a specific process for a smart car 100 to execute the above step S1308 provided in an embodiment of the present application.

[0387] As shown in FIG14 , the specific process of the smart car 100 executing step S1308 may include the following steps:

[0388] S1401, the smart car 100 determines whether there is a main display screen, which is a display screen that is processing a non-wake-up-free command through a voice assistant.

[0389] The specific content of step S1401 can refer to the relevant content of step S1304 shown in Figure 13 above, and will not be repeated here.

[0390] If the smart car 100 determines that a main display screen currently exists, the smart car 100 may execute the following step S1402 .

[0391] If the smart car 100 determines that the main display screen does not currently exist, the smart car 100 may execute the following step S1404.

[0392] S1402: The smart car 100 determines whether the display screen A corresponding to the audio zone 1 is the main display screen.

[0393] If a main display screen exists, the smart car 100 can determine display screen A corresponding to audio zone 1 based on the correspondence between audio zones and display screens, and determine whether display screen A is the main display screen. The correspondence between audio zones and display screens can be found in the above-mentioned step S1306 and will not be further described here.

[0394] If the smart car 100 determines that the display screen A corresponding to the audio zone 1 is the main display screen, the smart car 100 may execute the following step S1403 .

[0395] If the smart car 100 determines that the display screen A corresponding to the audio zone 1 is not the main display screen, the smart car 100 may execute the following step S1404.

[0396] S1403, the smart car 100 displays the voice dialogue logo and the voice sub-logo on the main display screen.

[0397] When there is a main display screen and display screen A is the main display screen, the main display screen can process multiple voice channels simultaneously, and one of the voice channels is a non-wake-up-free command, and the other one or more voice channels are wake-up-free commands.

[0398] In some embodiments, before receiving Voice 1, the main display screen only processes one voice channel, and this voice channel is a non-wake-up-free command. In this case, a voice dialogue identifier can be displayed on the main display screen, which is used to remind the user that the current display screen A is the main display screen and is processing a non-wake-up-free command. When Voice 1 is received, the smart car 100 can split the voice dialogue identifier into a voice sub-identifier, which is used to remind the user that the current display screen A (i.e., the main display screen) is processing both a non-wake-up-free command and a wake-up-free command.

[0399] For example, the speech conversation identifier may be the speech conversation identifier 501 shown in FIG5C , and the speech sub-identifier may be the speech sub-identifier 530 shown in FIG5E . The specific process of splitting the speech sub-identifier from the speech conversation identifier may refer to the relevant descriptions in the embodiments shown in FIG5C to FIG5E .

[0400] In other embodiments, before receiving Voice 1, the main display screen may process multiple voice messages, one of which is a non-wake-up-free command, while the other one or more voice messages are wake-up-free commands. In this case, the main display screen may display a voice conversation identifier and a voice sub-identifier. When Voice 1 is received, the smart car 100 may continue to display the voice conversation identifier and the voice sub-identifier. Optionally, in this case, a number may be displayed in the voice sub-identifier to indicate to the user the number of wake-up-free commands currently processed by display screen A.

[0401] S1404, the smart car 100 determines whether display screen A is processing other wake-up-free instructions.

[0402] In some embodiments, the smart car 100 can determine whether display screen A is processing other wake-up-free instructions based on whether display screen A displays a wake-up-free mark or meets the wake-up-free mark.

[0403] In other embodiments, the smart car 100 can also determine whether display screen A is processing other wake-up-free instructions based on whether display screen A turns on the voice assistant.

[0404] It is understandable that the two embodiments here are just two examples. In the embodiment of the present application, the smart car 100 can also determine whether the display screen A is processing other wake-up-free instructions based on other methods.

[0405] If the smart car 100 determines that the display screen A is processing other wake-up-free instructions, the smart car 100 can execute the following step S1405.

[0406] If the smart car 100 determines that the display screen A is not currently processing other wake-up-free instructions, the smart car 100 can execute the following step S1406.

[0407] S1405, the smart car 100 displays a composite wake-up-free logo on the display screen A corresponding to the audio zone 1. The composite wake-up-free logo is used to remind the user that the display screen A is processing multiple wake-up-free instructions.

[0408] In some embodiments, before receiving voice 1, screen A is processing only one wake-up-free instruction. In this case, a wake-up-free indicator may be displayed on screen A. Upon receiving voice 1, screen A may display a composite wake-up-free indicator on screen A in response to the wake-up-free instruction in voice 1. The composite wake-up-free indicator is used to prompt the user that screen A is processing multiple wake-up-free instructions.

[0409] For example, referring to the embodiment shown in Figures 8A to 8B above, in the above embodiment, voice 1 can be a wake-up-free command issued by user B, display screen A can be the central control screen 10, and the wake-up-free logo can be the composite wake-up-free logo 810 shown in Figure 8B above.

[0410] In other embodiments, before receiving voice 1, display screen A is processing multiple (two or more) wake-up-free instructions. In this case, a composite wake-up-free indicator may be displayed on display screen A. When voice 1 is received, display screen A may update the digital indicator in the composite wake-up-free indicator in response to the wake-up-free instruction in voice 1. The digital indicator can be used to prompt the user the number of wake-up-free instructions currently being executed by display screen A. For example, the digital indicator can refer to the digital indicator 814 in the embodiment shown in Figure 8B above.

[0411] S1406, the smart car 100 displays a wake-up-free logo on the display screen A corresponding to the audio zone 1.

[0412] Exemplarily, the wake-up-free flag may be the wake-up-free flag 601 shown in FIG. 6C , the wake-up-free flag 701 shown in FIG. 7C , the wake-up-free flag 711 shown in FIG. 7D , and so on.

[0413] By adopting the voice interaction method provided in the embodiment of the present application, the smart car 100 can conduct voice interaction with multiple users through one or more display screens, and support simultaneous processing of multiple wake-up-free commands, thereby improving the efficiency of voice interaction and providing users with better voice interaction services.

[0414] In the embodiment of the present application, the voice assistants in different display screens can adopt the same working mode or different working modes. Moreover, the voice assistant on the same display screen can also adopt different working modes for the voices initiated by different users.

[0415] In some embodiments, the voice assistant in the display screen can adopt different working modes for the voices of different users based on the wake-up method of the voice assistant. For example, when the voice assistant is awakened by the user through a wake-up word, the voice assistant can adopt a full-duplex interaction working mode to perform voice interaction with the user who issued the wake-up word. When the voice assistant is awakened by a wake-up-free instruction, the voice assistant can adopt a single-round interaction working mode to perform voice interaction with the user who issued the wake-up-free instruction. Using the above method, at the same time, the smart car 100 can ensure that at most one voice channel adopts a full-duplex interaction working mode, and the other one or more voice channels adopt a single-round interaction working mode. Since the power consumption required for single-round interaction is lower than that of other working modes, and the requirements for voice interaction are low, therefore, in the above case, the smart car 100 can ensure that the voice commands of multiple users can receive timely feedback, providing users with better voice interaction services.

[0416] It is understandable that the above embodiment is only an example. In the embodiment of the present application, when the voice assistant is awakened by the user through the wake-up word, the voice assistant can also adopt other working modes (such as continuous monitoring, multi-round interaction, etc.) to interact with the user by voice, and this application does not limit this.

[0417] The following introduces a specific application scenario of a voice interaction method provided by an embodiment of the present application.

[0418] For example, FIG15A shows a schematic diagram of an application scenario provided in an embodiment of the present application.

[0419] As shown in FIG15A , inside smart car 100, the father sits in the driver's seat, the mother sits in the passenger seat, the daughter sits in the second row on the left, the son sits in the second row on the right, the grandfather sits in the third row on the left, and the grandmother sits in the third row on the right. The father is in the driver's seat, the mother is in the passenger seat, the daughter is in the second row on the left, the son is in the second row on the right, the grandfather is in the third row on the left, and the grandmother is in the third row on the right.

[0420] The correspondence between the display screens and audio zones inside the smart car 100 is as follows: the central control screen 10 corresponds to the following multiple audio zones: main driver audio zone, second row left audio zone, third row left audio zone, third row right audio zone; the co-driver screen 20 corresponds to the co-driver audio zone; the right rear projection screen 30 corresponds to the second row right audio zone.

[0421] The smart car 100 can receive the voice containing the wake-up word issued by the father from the main driving audio zone, such as "Xiao A, navigate home", turn on the voice assistant of the central control screen 10, determine the central control screen 10 as the main display screen, and display the voice dialogue logo on the central control screen 10. The specific interface can refer to the main interface 500 and the voice dialogue logo 501 shown in Figure 5B above.

[0422] While the smart car 100 is receiving the father's voice commands, the smart car 100 can also receive voice commands from other users, such as the mother's wake-up command "Turn on the air conditioner", the grandfather's wake-up command "Close the car windows", and the son's wake-up command "I want to watch cartoon XX".

[0423] Based on the correspondence between audio zones and display screens, the central control screen 10 can be responsible for processing the following voice commands: "Little A, navigate home" issued by the father in the main driver's audio zone and "Close the windows" issued by the grandfather in the third-row left audio zone. In this case, the central control screen 10 must simultaneously process one non-wake-up command and one wake-up command. At this time, the central control screen 10 may display the main interface 1500 shown in Figure 15B. At the same time, the smart car 100 can also execute the grandfather's wake-up command "Close the windows."

[0424] As shown in Figure 15B, the main interface 1500 of the central control screen 10 can display a voice dialogue logo 1501 and a voice sub-logo 1502 split from the voice dialogue logo 1501. Optionally, a sound zone indicator 1503 can also be displayed. The voice dialogue logo 1501 and the voice sub-logo 1502 can be used to prompt the user that the central control screen 10 is processing multiple voices at the same time, and one of the voices is a non-wake-up-free command. The voice dialogue logo 1501 can also be used to prompt the user that the current main display screen is the central control screen 10. The specific content and display process of the main interface 1500, the voice dialogue logo 1501 and the voice sub-logo 1502 can refer to the relevant description in the embodiments shown in Figures 5A to 5F above, and will not be repeated here. Moreover, the text 1504 displayed in the voice dialogue logo 1501 can be a voice command issued by the father, such as "Xiao A, navigate home."

[0425] Based on the relationship between the audio zones and the display screen, the passenger screen 20 can be responsible for processing the mother's wake-up-free command "Turn on the air conditioner." In this case, the passenger screen 20 needs to process one wake-up-free command and can display the main interface 1510 shown in Figure 15C.

[0426] As shown in Figure 15C, a wake-up-free indicator 1511 may be displayed on the main interface 1510 of the passenger screen 20. This wake-up-free indicator 1511 can be used to inform the user that the passenger screen 20 is receiving and processing a wake-up-free instruction. The specific content of the main interface 1510, the wake-up-free indicator 1511, and the display process can be found in the description of the embodiments shown in Figures 6B to 6D above and will not be repeated here. Furthermore, the text 1512 displayed on the wake-up-free indicator 1511 may be a wake-up-free instruction issued by the mother, such as "Turn on the air conditioner."

[0427] Based on the relationship between the audio zones and the display screen, the right rear projection screen 30 can be responsible for processing the son's wake-up command "I want to watch cartoon XX." In this case, the right rear projection screen 30 is responsible for processing the wake-up command and can display the main interface 1520 shown in Figure 15D.

[0428] As shown in FIG15D , a wake-up-free indicator 1521 may be displayed in the main interface 1520 of the right rear projection screen 30. The wake-up-free indicator 1521 can be used to prompt the user that the right rear projection screen 30 is receiving and processing the wake-up-free instruction. The specific content of the main interface 1520, the wake-up-free indicator 1521, and the display process can be referred to the relevant description of the embodiments shown in FIG6B to FIG6D above, and will not be repeated here. In addition, the text 1512 displayed in the wake-up-free indicator 1511 can be feedback information based on the voice command issued by the mother, such as "I want to watch cartoon XX."

[0429] After recognizing the son's wake-up-free instruction "I want to watch cartoon XX" and obtaining the video data of cartoon XX, the right rear projection screen 30 can stop displaying the wake-up-free indicator 1521 and display the video playback interface 1530 shown in FIG15E based on the video data. The video playback interface 1530 can be used to play cartoon XX.

[0430] It can be understood that the embodiment shown in Figures 15A to 15E above is only an example of a scenario. The voice interaction method provided in the embodiment of the present application can also be applied to more application scenarios different from the above embodiment, and the present application does not limit it here.

[0431] The following introduces the functional modules of a smart car 100 provided in an embodiment of the present application.

[0432] FIG16 shows a schematic diagram of functional modules of a smart car 100 provided in an embodiment of the present application.

[0433] As shown in Figure 16, the smart car 100 may include a voice receiving module 1601, a voice zone recognition module 1602, a voice recognition module 1603, a voice assistant module 1604, and an execution module 1605. Optionally, the smart car 100 may also include any one or more of the following: a voice zone locking module 1606 and a voiceprint recognition module 1607.

[0434] The voice receiving module 1601 can collect voices uttered by users within the smart car 100, including but not limited to wake-up words, wake-up-free commands, and non-wake-up-free commands. The voice receiving module 1601 can include one or more submodules, each of which can correspond to a sound zone and be used to collect voices within that sound zone. After receiving the user's voice, the voice receiving module 1601 can send the collected voice to the voice recognition module 1603 and the sound zone recognition module 1602.

[0435] The vocal range identification module 1602 can determine the vocal range corresponding to the speech based on the speech. The specific method by which the vocal range identification module 1602 determines the vocal range corresponding to the speech can be found in the description of step S1302 shown in FIG. 13 , and will not be further described here. After determining the vocal range of the speech, the vocal range identification module 1602 can send the vocal range corresponding to the speech to the voice assistant module 1604.

[0436] The speech recognition module 1603 can recognize the text content of the speech. After receiving the speech sent by the speech receiving module 1601, the speech recognition module 1603 can recognize the text content of the speech and make a judgment based on the text content of the speech to determine whether the speech contains a wake-up word, whether it is a wake-up-free instruction, or whether it contains a non-wake-up-free instruction.

[0437] When the voice recognition module 1603 determines that the user's voice contains a wake-up word, the voice recognition module 1603 can send instruction 1 to the voice assistant module 1604. Instruction 1 is used to instruct the voice assistant module 1604 to determine the main display screen and turn on the voice assistant function of the main display screen.

[0438] When the speech recognition module 1603 determines that the user's voice contains a wake-up-free instruction, the speech recognition module 1603 can send instruction 2 to the voice assistant module 1604. Instruction 2 may include the text content of the speech (or the wake-up-free instruction contained in the speech). Instruction 2 can be used to instruct the voice assistant module 1604 to process the wake-up-free instruction.

[0439] In some embodiments, the speech recognition module 1603 may also receive a judgment result sent by the voiceprint recognition module 1607. When it is determined that the user's voice contains a non-wake-up-free command, and based on the judgment result sent by the voiceprint recognition module 1607, it is determined that the user is a user who uses the wake-up word to turn on the voice assistant function of the main display screen, the speech recognition module 1603 may send instruction 3 to the voice assistant module 1604, and instruction 3 is used to instruct the voice assistant module 1604 to process the non-wake-up-free command.

[0440] The voice assistant module 1604 may include one or more submodules, each of which may correspond to a display screen in the smart car 100, and each display screen may use the corresponding submodule function to implement the voice assistant function of the display screen. Upon receiving instruction 1 sent by the voice recognition module 1603, the voice assistant module 1604 may determine the display screen corresponding to the sound zone based on the sound zone sent by the sound zone recognition module 1602, and determine the display screen as the main display screen. The voice assistant function of the display screen may be activated through the corresponding submodule of the display screen in the voice assistant module 1604, and the main display screen may be controlled to display a voice dialogue logo. Optionally, corresponding feedback information may also be output.

[0441] When the voice assistant module 1604 receives instruction 2 sent by the voice recognition module 1603, the voice assistant module 1604 can determine the display screen corresponding to the sound zone based on the sound zone sent by the sound zone recognition module 1602, and turn on the voice assistant function of the display screen through the corresponding sub-module of the display screen in the voice assistant module 1604, and control the display screen to display the wake-up-free mark. Optionally, it can also output corresponding feedback information, etc.

[0442] When the voice assistant module 1604 receives instruction 3 sent by the voice recognition module 1603, the voice assistant module 1604 can process the non-wake-up-free instruction through the corresponding sub-module of the main display screen in the voice assistant module 1604, for example, controlling the main display screen to output the text content of the non-wake-up-free instruction, or output corresponding feedback information, etc.

[0443] In some embodiments, when the voice assistant module 1604 receives the text content of the voice instruction sent by the voice recognition module 1604, such as a wake-up-free instruction or a non-wake-up-free instruction, it can determine the operation to be performed based on the text content of the received voice instruction and send the execution instruction to the execution module 1605. The execution instruction is used to instruct the execution module 1605 to perform the specified operation (such as closing the car window, turning on the navigation, turning on the air conditioner, etc.).

[0444] The execution module 1605 can receive the execution instruction sent by the voice assistant module 1604 and execute the operation specified by the execution instruction (such as closing the car window, turning on the navigation, turning on the air conditioner, etc.).

[0445] The voice zone lock module 1606 can assist the voice receiving module 1601 in performing voice collection. When the voice receiving module 1601 collects the user's voice in a certain voice zone through one of its submodules, the voice zone lock module 1606 can suppress voice outside the corresponding voice zone while enhancing the voice in the corresponding voice zone. This ensures that the voice receiving module 1601 can collect voice in different voice zones.

[0446] The voiceprint recognition module 1607 can extract voiceprint information based on the voice, and determine whether two voices come from the same user based on the voiceprint information. In an embodiment of the present application, the voiceprint recognition module 1607 can receive the voice sent by the voice receiving module 1601, and determine whether the voice comes from the user who activated the voice assistant function of the main display screen by using the wake-up word based on the voiceprint information, and send the judgment result to the voice recognition module 1603.

[0447] It is understandable that the embodiment shown in FIG16 is merely an example. In the embodiment of the present application, the smart car 100 may also include more, fewer, or different functional modules than those in the above embodiment, and the present application does not limit this.

[0448] The following describes the voice interaction method provided in the embodiments of the present application.

[0449] FIG17 shows a flow chart of a voice interaction method provided in an embodiment of the present application.

[0450] As shown in FIG17 , the specific process of the vehicle executing the voice interaction method may include the following steps:

[0451] S1701: The vehicle receives a first voice message from a first user, where the first voice message includes a wake-up word.

[0452] The vehicle includes a first display screen. The vehicle may be the smart car 100 in the above embodiment.

[0453] For example, reference may be made to the embodiments shown in FIG. 5A to FIG. 5F . In the embodiments, the first display screen may be the central control screen 10 , the first user may be user A, and the first voice may include a wake-up word issued by user A.

[0454] As another example, reference may be made to the embodiment shown in FIG. 12A to FIG. 12B . In the embodiment, the first display screen may be the central control screen 10 , the first user may be user A, and the first voice may include a wake-up word issued by user A.

[0455] In one possible implementation, the first voice message also includes a first instruction for instructing the vehicle to perform a first operation. The method further includes: executing the first operation in response to the first voice message. Thus, after receiving the first voice message, the vehicle can execute the operation corresponding to the first instruction in the first voice message, such as opening navigation, playing music, or opening a window.

[0456] In one possible implementation, after the first dialogue identifier is displayed on the first display screen, the method further includes: receiving a third voice from the first user, the third voice being used to instruct the vehicle to perform the first operation; and performing the first operation in response to the third voice. In this way, the first user may also issue a voice command after turning on the voice assistant function of the first display screen through the wake-up word. In this case, the first display screen has already started voice interaction with the first user in response to the first user's wake-up word. Therefore, the first display screen can continue to receive voice commands issued by the user and perform the operation corresponding to the voice command, such as turning on navigation, playing music, opening the window, etc.

[0457] S1702, the vehicle displays a first dialogue identifier on the first display screen in response to the first voice, and the first dialogue identifier is used to prompt the user that the first display screen is performing voice interaction with the first user.

[0458] Exemplarily, the first conversation identifier may be the voice conversation identifier 501 shown in FIG. 5B .

[0459] As another example, the first dialogue identifier may be the voice dialogue identifier 1201 shown in FIG. 12A .

[0460] In one possible implementation, the method further includes: after displaying the first dialogue identifier on the first display screen, outputting a first feedback, the first feedback being used to prompt the user that the vehicle has received the first voice; after displaying the voice sub-identifier on the first display screen, outputting a second feedback, the second feedback being used to prompt the user that the vehicle has received the second voice.

[0461] The vehicle may output feedback information (e.g., first feedback, second feedback, etc.) using one or more methods, such as display screen display, voice announcement, flashing indicator light, vibration, etc. The first feedback and the second feedback may be output in different ways, or they may be output in the same way. In this way, the user can be notified by the output of feedback information that the vehicle has received the user's voice.

[0462] In one possible implementation, the first feedback is also used to inform the user whether the first voice-directed operation has been executed, and the second feedback is also used to inform the user whether the second voice-directed operation has been executed. In this way, the feedback information can be used to inform the user of the vehicle's execution of the user's voice command, such as prompting the user that the command has been completed.

[0463] The specific content and output method of the feedback information can also refer to the relevant description in Figure 6D and step S1311 shown in Figure 13, which will not be repeated here.

[0464] S1703: The vehicle receives a second voice message from a second user.

[0465] For example, reference may be made to the embodiment shown in FIG. 5A to FIG. 5B . In the embodiment, the second user may be user B.

[0466] S1704, when the second voice includes a wake-up-free instruction, the vehicle displays a voice sub-identifier on the first display screen in response to the second voice, and the voice sub-identifier is used to prompt the user that the first display screen is processing multiple voices.

[0467] For example, reference may be made to the embodiments shown in FIG. 5A to FIG. 5F . In the embodiments, the second voice may be a wake-up-free instruction issued by user B, and the voice sub-identifier may be the voice sub-identifier 530 shown in FIG. 5E .

[0468] In this way, the vehicle can process the second user's wake-up instructions through the first display screen while conducting voice interaction with the first user on the first display screen, thereby improving the efficiency of voice interaction and providing users with better voice interaction services.

[0469] In one possible implementation, a vehicle interior includes a first audio zone and a second audio zone, a first user is located in the first audio zone, and a second user is located in the second audio zone; a first display screen is used to respond to speech from the first audio zone and the second audio zone; displaying a first conversation identifier on the first display screen specifically includes: determining that the first speech originates from the first audio zone based on the first speech; displaying the first conversation identifier on the first display screen based on the first audio zone; and displaying a voice sub-identifier on the first display screen specifically includes: determining that the second speech originates from the second audio zone based on the second speech; and displaying the voice sub-identifier on the first display screen based on the second audio zone. In this way, when the vehicle receives a user's speech, it can determine the display screen used to display the speech identifier (e.g., the first conversation identifier, the speech sub-identifier, etc.) based on the correspondence between the audio zone from which the speech originates and the display screen.

[0470] For example, referring to the embodiment shown in 3E above, if the first display screen is the central control screen 10, the first sound zone may be the main driving sound zone, and the second sound zone may be the second row left sound zone.

[0471] S1705, when the second voice includes a wake-up word, the vehicle stops displaying the first dialogue logo and displays the second dialogue logo on the first display screen, and the second dialogue logo is used to prompt the user that the first display screen is performing voice interaction with the second user.

[0472] For example, referring to the embodiment shown in FIG. 12A to FIG. 12B , the first conversation identifier may be the voice conversation identifier 1201 shown in FIG. 12A , and the second conversation identifier may be the voice conversation identifier 1211 shown in FIG. 12B .

[0473] In this way, when the vehicle receives the wake-up word from the second user, it can end the voice interaction between the first display and the first user and continue the voice interaction with the second user through the first display. This ensures that at the same time, the vehicle only needs to process the voice interaction initiated by the wake-up word on one route.

[0474] In one possible implementation, a vehicle interior includes a first audio zone and a second audio zone, a first user is located in the first audio zone, and a second user is located in the second audio zone; a first display screen is used to respond to speech from the first audio zone and the second audio zone; displaying a first conversation identifier on the first display screen specifically includes: determining that the first speech originates from the first audio zone based on the first speech; displaying the first conversation identifier on the first display screen based on the first audio zone; displaying the second conversation identifier on the first display screen specifically includes: determining that the second speech originates from the second audio zone based on the second speech; and displaying the second conversation identifier on the first display screen based on the second audio zone. In this way, when the vehicle receives a user's speech, it can determine the display screen used to display the speech identifier (e.g., the first conversation identifier, the second conversation identifier, etc.) based on the correspondence between the audio zone from which the speech originates and the display screen.

[0475] For example, referring to the embodiment shown in 3E above, if the first display screen is the central control screen 10, the first sound zone can be the main driving sound zone, and the second sound zone can be the third row left sound zone.

[0476] In one possible implementation, the method further includes: when the second voice includes a wake-up word, outputting an interruption prompt, where the interruption prompt is used to prompt that the voice interaction between the first user and the first display screen has been interrupted. In one possible implementation, the vehicle may output the interruption prompt in one or more ways, such as display screen display, flashing indicator light, voice broadcast, vibration, etc. In this way, the interruption prompt can be used to indicate that the voice interaction between the first user and the first display screen has ended. Exemplarily, the interruption prompt may be the interruption prompt 1103 shown in FIG. 11C above.

[0477] In one possible implementation, the method also includes: in response to the first voice, displaying a first indicator on the first display screen, the first indicator being used to indicate the position of the first sound zone relative to the first display screen; when the second voice includes a wake-up-free instruction, in response to the second voice, replacing the first indicator with a second indicator, the second indicator being used to indicate the positions of the first sound zone and the second sound zone relative to the display screen; when the second voice includes a wake-up word, in response to the second voice, replacing the first indicator with a third indicator, the third indicator being used to indicate the position of the second sound zone relative to the display screen.

[0478] In this way, the user can be prompted with the sound zone indicator (eg, the first indicator, the second indicator, the third indicator, etc.) of the source sound zone of the speech currently being processed on the display screen.

[0479] Exemplarily, the first indicator may be the music zone indicator 505 shown in FIG5B above, and the second indicator may be the music zone indicator 511 shown in FIG5D above; and again exemplary, the first indicator may be the music zone indicator 1202 shown in FIG12A above, and the third indicator may be the music zone indicator 1212 shown in FIG12B above.

[0480] In one possible implementation, the voice interaction mode between the first display and the first user is full-duplex interaction; when the second voice includes a wake-up-free command, the voice interaction mode between the first display and the second user is single-round interaction; when the second voice includes a wake-up word, the voice interaction mode between the first display and the second user is full-duplex interaction.

[0481] This ensures that voice interaction initiated by a wake-up word is full-duplex, while voice interaction initiated without a wake-up command is single-round. This ensures that timely voice interaction services are provided to more users.

[0482] FIG18 shows a flow chart of another voice interaction method provided in an embodiment of the present application.

[0483] As shown in FIG18 , the specific process of the vehicle executing the voice interaction method may include the following steps:

[0484] S1801, the vehicle receives a fourth voice message from a first user, where the fourth voice message includes a wake-up-free instruction.

[0485] The vehicle may be the smart car 100 in the above embodiment.

[0486] For example, reference may be made to the embodiments shown in FIG. 8A to FIG. 8C . In the embodiments, the first user may be user A, and the fourth voice may be a wake-up-free instruction issued by user A.

[0487] As another example, reference may be made to the embodiments shown in FIG. 10A to FIG. 10C . In the embodiments, the first user may be user A, and the fourth voice may be a wake-up-free instruction issued by user A.

[0488] S1802: The vehicle responds to the fourth voice and displays a first wake-up-free indicator on the first display screen. The first wake-up-free indicator is used to prompt the user that the first display screen is processing the wake-up-free instruction.

[0489] For example, the first display screen may be the central control screen 10 in the embodiment shown in Figures 8A to 8C above, or the central control screen 10 in the embodiment shown in Figures 10A to 10C above. The first wake-up-free indicator may be the wake-up-free indicator 801 shown in Figure 8A above, or the wake-up-free indicator 1011 shown in Figure 10A above.

[0490] S1803: The vehicle receives a fifth voice message from the second user.

[0491] For example, reference may be made to the embodiments shown in FIG. 8A to FIG. 8C , or FIG. 10A to FIG. 10C . In the above embodiments, the second user may be user B.

[0492] S1804, when the fifth voice includes a wake-up-free instruction, the vehicle responds to the fifth voice, stops displaying the first wake-up-free indicator, and displays a composite wake-up-free indicator on the first display screen, where the composite wake-up-free indicator is used to prompt the user that the first display screen is processing multiple wake-up-free instructions.

[0493] Illustratively, the composite wake-up-free flag may be the composite wake-up-free flag 810 shown in FIG. 8B .

[0494] In this way, the vehicle can process wake-up-free commands issued by multiple users through the first display screen, thereby improving the efficiency of voice interaction and providing users with better voice interaction services.

[0495] In one possible implementation, the interior of a vehicle includes a first audio zone and a second audio zone, a first user is located in the first audio zone, and a second user is located in the second audio zone; a first display screen is used to respond to voices from the first audio zone and the second audio zone; a first wake-up-free indicator is displayed on the first display screen, specifically including: determining that the fourth voice comes from the first audio zone based on a fourth voice; displaying the first wake-up-free indicator on the first display screen based on the first audio zone; and displaying a composite wake-up-free indicator on the first display screen, specifically including: determining that the fifth voice comes from the second audio zone based on a fifth voice; and displaying the composite wake-up-free indicator on the first display screen based on the second audio zone. In this way, when the vehicle receives a user's wake-up-free command, it can determine the display screen used to display the wake-up-free indicator based on the correspondence between the audio zone from which the voice originates and the display screen.

[0496] S1805. When the fifth voice includes a wake-up word, the vehicle responds to the fifth voice, stops displaying the first wake-up-free logo, and displays a third dialogue logo and a voice sub-logo on the first display screen. The third dialogue logo is used to prompt the user that the first display screen is conducting voice interaction with the second user, and the voice sub-logo is used to prompt the user that the first display screen is processing multiple voices.

[0497] For example, reference may be made to the embodiments shown in Figures 10A to 10C above. In the above embodiments, the first wake-up-free identifier may be the wake-up-free identifier 1011 shown in Figure 10A, the third conversation identifier may be the voice conversation identifier 1021 shown in Figure 10C, and the voice sub-identifier may be the voice sub-identifier 1022 shown in Figure 10C.

[0498] In this way, the vehicle can process the wake-up command issued by the first user through the first display screen, receive the wake-up word issued by the second user, and conduct voice interaction with the second user through the first display screen.

[0499] In one possible implementation, a vehicle interior includes a first audio zone and a second audio zone, a first user is located in the first audio zone, and a second user is located in the second audio zone; a first display screen is configured to respond to voices from the first and second audio zones; a first wake-up-free indicator is displayed on the first display screen, specifically comprising: determining, based on a fourth audio zone, that the fourth audio zone originates from the first audio zone; displaying the first wake-up-free indicator on the first display screen based on the first audio zone; and displaying a third conversation indicator and a voice sub-identifier on the first display screen, specifically comprising: determining, based on a fifth audio zone, that the fifth audio zone originates from the second audio zone; and displaying the third conversation indicator and voice sub-identifier on the first display screen based on the second audio zone. In this way, when the vehicle receives a user's wake-up-free command or wake-up word, it can determine the display screen to display the wake-up-free indicator based on the correspondence between the audio zone from which the audio source originates and the display screen.

[0500] In one possible implementation, the method further includes: when the fifth voice includes a wake-up word, displaying a fourth indicator on the first display screen in response to the fifth voice, the fourth indicator being used to indicate positions of the first sound zone and the second sound zone relative to the display screen. In this way, the fourth indicator can be displayed on the first display screen when the wake-up word is received.

[0501] Exemplarily, the fourth indicator may be the audio zone indicator 1023 shown in FIG. 10B .

[0502] In one possible implementation, the method also includes: in response to the fourth voice, displaying a fifth indicator on the first display screen, the fifth indicator being used to indicate the position of the first sound zone relative to the first display screen; when the fifth voice includes a wake-up-free instruction, in response to the fifth voice, replacing the fifth indicator with a sixth indicator, the sixth indicator being used to indicate the positions of the first sound zone and the second sound zone relative to the display screen; when the fifth voice includes a wake-up word, in response to the fifth voice, replacing the fifth indicator with a seventh indicator, the seventh indicator being used to indicate the positions of the first sound zone and the second sound zone relative to the display screen.

[0503] In this way, the user can be prompted with the sound zone indicator (eg, the first indicator, the second indicator, the fourth indicator, etc.) of the source sound zone of the speech currently being processed on the display screen.

[0504] For example, the fifth indicator may be the music zone indicator 802 shown in FIG. 8A , the sixth indicator may be the music zone indicator 805 shown in FIG. 8B , and the seventh indicator may be the music zone indicator 1023 shown in FIG. 10B .

[0505] In one possible implementation, the voice interaction mode between the first display and the first user is single-round interaction; when the fifth voice includes a wake-up-free command, the voice interaction mode between the first display and the second user is single-round interaction; when the fifth voice includes a wake-up word, the voice interaction mode between the first display and the second user is full-duplex interaction.

[0506] This ensures that voice interaction initiated by a wake-up word is full-duplex, while voice interaction initiated without a wake-up command is single-round. This ensures that timely voice interaction services are provided to more users.

[0507] FIG19 shows a flow chart of another voice interaction method provided in an embodiment of the present application.

[0508] As shown in FIG19 , the specific process of the vehicle executing the voice interaction method may include the following steps:

[0509] S1901, the vehicle receives a sixth voice emitted by a first user from a first sound zone, the sixth voice includes a wake-up word, the vehicle includes a first display screen and a second display screen, and the interior of the vehicle includes the first sound zone and the second sound zone; the first display screen is used to respond to the voice from the first sound zone, and the second display screen is used to respond to the voice from the second sound zone.

[0510] The vehicle may be the smart car 100 in the above embodiment.

[0511] For example, referring to the embodiments shown in Figures 6A to 6D above, in which the first display screen may be the central control screen 10, and the second display screen may be the passenger screen 20. The first user may be user A, and the first audio zone may be audio zone a (e.g., the driver's seat audio zone, the second row left audio zone, etc.). The second user may be user B, and the second audio zone may be audio zone b (e.g., the passenger seat audio zone). The sixth voice may include the wake-up word uttered by user A.

[0512] For example, referring to the embodiment shown in Figures 11A to 11D above, in the above embodiment, the first display screen may be the central control screen 10, and the second display screen may be the right rear projection screen 30. The first user may be user A, and the first audio zone may be audio zone a (e.g., the main driver audio zone, the second row left audio zone, etc.). The second user may be user B, and the second audio zone may be audio zone b (e.g., the second row right audio zone).

[0513] S1902 , the vehicle responds to the sixth voice and determines based on the sixth voice that the sixth voice comes from the first sound zone.

[0514] S1903, the vehicle displays a third dialogue identifier on the first display screen based on the first sound zone, and the third dialogue identifier is used to prompt the user that the first display screen is performing voice interaction with the first user.

[0515] Illustratively, the third conversation identifier may be the voice conversation identifier 501 shown in FIG. 6A , or the voice conversation identifier 1101 shown in FIG. 11A .

[0516] S1904: The vehicle receives a seventh voice from a second user in a second voice zone.

[0517] S1905 , when the seventh voice includes a wake-up-free instruction, the vehicle responds to the seventh voice and determines based on the seventh voice that the seventh voice comes from the second sound zone.

[0518] S1906, the vehicle displays a second wake-up-free indicator on the second display screen based on the second sound zone, and the second wake-up-free indicator is used to prompt the user that the second display screen is processing the wake-up-free instruction.

[0519] Exemplarily, the second wake-up-free flag may be the wake-up-free flag 601 shown in FIG. 6C .

[0520] In this way, based on the correspondence between the sound zones and the display screens, voice interaction can be performed with different users through different display screens, thereby improving the efficiency of voice interaction and providing users with better voice interaction services.

[0521] S1907, when the seventh voice includes a wake-up word, the vehicle stops displaying the third dialogue identifier on the first display screen in response to the seventh voice.

[0522] Illustratively, the third dialogue identifier may be the voice dialogue identifier 1101 shown in FIG. 11A .

[0523] S1908: The vehicle determines, based on the seventh speech, that the seventh speech is from the second sound zone.

[0524] S1909, the vehicle displays a fourth dialogue identifier on the second display screen based on the second audio zone, and the fourth dialogue identifier is used to prompt the user that the second display screen is performing voice interaction with the second user.

[0525] Exemplarily, the fourth dialogue identifier may be the voice dialogue identifier 1111 shown in FIG. 11D .

[0526] In this way, when the vehicle receives the wake-up word from the second user, it can end the voice interaction between the first display and the first user and continue the voice interaction with the second user through the second display. This ensures that at the same time, the vehicle only needs to process the voice interaction initiated by the wake-up word on one route.

[0527] In one possible implementation, the voice interaction mode between the first display and the first user is full-duplex interaction; when the seventh voice includes a wake-up-free command, the voice interaction mode between the second display and the second user is single-round interaction; when the seventh voice includes a wake-up word, the voice interaction mode between the second display and the second user is full-duplex interaction.

[0528] This ensures that voice interaction initiated by a wake-up word is full-duplex, while voice interaction initiated without a wake-up command is single-round. This ensures that timely voice interaction services are provided to more users.

[0529] FIG20 shows a flow chart of another voice interaction method provided in an embodiment of the present application.

[0530] As shown in FIG20 , the specific process of the vehicle executing the voice interaction method may include the following steps:

[0531] S2001, the vehicle receives an eighth voice sent by a first user from a first sound zone, the eighth voice includes a wake-up-free instruction, the vehicle includes a first display screen and a second display screen, and the interior of the vehicle includes the first sound zone and the second sound zone; the first display screen is used to respond to the voice from the first sound zone, and the second display screen is used to respond to the voice from the second sound zone.

[0532] The vehicle may be the smart car 100 in the above embodiment.

[0533] For example, referring to the embodiments shown in Figures 7A to 7D above, in the above embodiments, the first display screen may be the central control screen 10, and the second display screen may be the right rear projection screen 30. The first user may be user A, and the first audio zone may be audio zone a (e.g., the main driver audio zone, the second row left audio zone, etc.). The second user may be user B, and the second audio zone may be audio zone b (e.g., the second row right audio zone). The eighth voice may include a wake-up-free command issued by user A.

[0534] For example, referring to the embodiment shown in Figures 9A to 9D above, in the above embodiment, the first display screen may be the right rear projection screen 30, and the second display screen may be the central control screen 10. The first user may be user A, and the first audio zone may be audio zone a (e.g., the second row right audio zone). The second user may be user B, and the second audio zone may be audio zone b (e.g., the main driving audio zone).

[0535] S2002 : The vehicle responds to the eighth voice and determines based on the eighth voice that the eighth voice comes from the first sound zone.

[0536] S2003: The vehicle displays a third wake-up-free indicator on the first display screen based on the first sound zone. The third wake-up-free indicator is used to prompt the user that the first display screen is processing a wake-up-free instruction.

[0537] Exemplarily, the third wake-up-free flag may be the wake-up-free flag 701 shown in FIG. 7C .

[0538] S2004: The vehicle receives a ninth voice message emitted by a second user from a second voice zone.

[0539] S2005 , when the ninth voice includes a wake-up-free instruction, the vehicle responds to the ninth voice and determines based on the ninth voice that the ninth voice comes from the second sound zone.

[0540] S2006: The vehicle displays a fourth wake-up-free indicator on the second display screen based on the second sound zone. The fourth wake-up-free indicator is used to prompt the user that the second display screen is processing a wake-up-free instruction.

[0541] Exemplarily, the fourth wake-up-free flag may be the wake-up-free flag 711 shown in FIG. 7D .

[0542] In this way, based on the correspondence between the sound zones and the display screens, the wake-up-free commands of different users can be processed through different display screens, thereby improving the efficiency of voice interaction and providing users with better voice interaction services.

[0543] S2007 , when the ninth voice includes a wake-up word, the vehicle responds to the ninth voice and determines based on the ninth voice that the ninth voice is from the second sound zone.

[0544] S2008, the vehicle displays a fifth dialogue identifier on the second display screen based on the second audio zone, and the fifth dialogue identifier is used to prompt the user that the second display screen is performing voice interaction with the second user.

[0545] Exemplarily, the fifth dialogue identifier may be the voice dialogue identifier 911 shown in FIG. 9C .

[0546] In this way, when the vehicle receives the wake-up word issued by the second user, it can process the first user's wake-up command through the first display screen while performing voice interaction with the second user through the second display screen.

[0547] In one possible implementation, the voice interaction mode between the first display and the first user is single-round interaction; when the ninth voice includes a wake-up-free command, the voice interaction mode between the second display and the second user is single-round interaction; when the ninth voice includes a wake-up word, the voice interaction mode between the second display and the second user is full-duplex interaction.

[0548] This ensures that voice interaction initiated by a wake-up word is full-duplex, while voice interaction initiated without a wake-up command is single-round. This ensures that timely voice interaction services are provided to more users.

[0549] The various implementation modes of this application can be combined arbitrarily to achieve different technical effects.

[0550] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described herein are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0551] Those skilled in the art will appreciate that all or part of the process steps in the above-described method embodiments can be implemented by a computer program instructing the relevant hardware. The program can be stored in a computer-readable storage medium, and when executed, the program can include the process steps in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

[0552] In short, the above description is only an embodiment of the technical solution of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made based on the disclosure of the present invention should be included in the scope of protection of the present invention.

Claims

1. A voice interaction method, characterized in that: Applied to a vehicle, the vehicle includes a first display screen; the method includes: Receiving a first voice sent by a first user, where the first voice includes a wake-up word; In response to the first voice, displaying a first dialogue mark on the first display screen, wherein the first dialogue mark is used to prompt the user that the first display screen is performing voice interaction with the first user; receiving a second voice message from a second user; When the second voice includes a wake-up-free instruction, a voice sub-identifier is displayed on the first display screen in response to the second voice, and the voice sub-identifier is used to prompt the user that the first display screen is processing multi-channel voices.

2. The method according to claim 1, characterized in that The first voice also includes a first instruction, and the first instruction is used to instruct the vehicle to perform the first operation; the method also includes: In response to the first voice, the first operation is performed.

3. The method according to claim 1, characterized in that After displaying the first dialog identifier on the first display screen, the method further includes: receiving a third voice issued by the first user, wherein the third voice is used to instruct the vehicle to perform a first operation; In response to the third voice, the first operation is performed.

4. The method according to any one of claims 1 to 3, characterized in that The method further comprises: After the first dialogue mark is displayed on the first display screen, a first feedback is output, where the first feedback is used to prompt a user that the vehicle has received the first voice; After the voice sub-identity is displayed on the first display screen, a second feedback is output, where the second feedback is used to prompt a user that the vehicle has received the second voice.

5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: When the second voice includes the wake-up word, the first dialogue identifier is stopped from being displayed, and a second dialogue identifier is displayed on the first display screen, where the second dialogue identifier is used to prompt the user that the first display screen is performing voice interaction with the second user.

6. The method according to claim 5, characterized in that The method further comprises: When the second voice includes the wake-up word, an interruption prompt is output, where the interruption prompt is used to prompt that the voice interaction between the first user and the first display screen is interrupted.

7. The method according to any one of claims 1 to 6, characterized in that The interior of the vehicle includes a first sound zone and a second sound zone, the first user is located in the first sound zone, and the second user is located in the second sound zone; The first display screen is used to respond to speech from the first sound zone and the second sound zone; The step of displaying the first dialog identifier on the first display screen specifically includes: determining, based on the first voice, that the first voice comes from the first sound zone; displaying a first dialogue identifier on the first display screen based on the first sound zone; The displaying the voice sub-identification on the first display screen specifically includes: determining, based on the second voice, that the second voice comes from the second sound zone; The voice sub-identity is displayed on the first display screen based on the second sound zone.

8. The method according to claim 5 or 6, characterized in that: The interior of the vehicle includes a first sound zone and a second sound zone, the first user is located in the first sound zone, and the second user is located in the second sound zone; The first display screen is used to respond to speech from the first sound zone and the second sound zone; The step of displaying the first dialog identifier on the first display screen specifically includes: determining, based on the first voice, that the first voice comes from the first sound zone; displaying a first dialogue identifier on the first display screen based on the first sound zone; The displaying the second dialogue identifier on the first display screen specifically includes: determining, based on the second voice, that the second voice comes from the second sound zone; The second dialogue identifier is displayed on the first display screen based on the second audio zone.

9. The method according to claim 7 or 8, characterized in that: The method further comprises: In response to the first voice, displaying a first indicator on the first display screen, the first indicator being used to indicate a position of the first sound zone relative to the first display screen; When the second voice includes a wake-up-free instruction, in response to the second voice, the first indicator is replaced with a second indicator, wherein the second indicator is used to indicate positions of the first sound zone and the second sound zone relative to the display screen; When the second voice includes the wake-up word, in response to the second voice, the first indicator is replaced with a third indicator, where the third indicator is used to indicate a position of the second sound zone relative to the display screen.

10. The method according to any one of claims 1 to 9, characterized in that The voice interaction mode between the first display screen and the first user is full-duplex interaction; when the second voice includes a wake-up-free instruction, the voice interaction mode between the first display screen and the second user is single-round interaction; when the second voice includes the wake-up word, the voice interaction mode between the first display screen and the second user is full-duplex interaction.

11. A voice interaction method, characterized in that: Applied to a vehicle, the vehicle includes a first display screen; the method includes: Receiving a fourth voice sent by the first user, wherein the fourth voice includes a wake-up-free instruction; In response to the fourth voice, a first wake-up-free mark is displayed on the first display screen, where the first wake-up-free mark is used to prompt a user that the first display screen is processing a wake-up-free instruction; receiving a fifth voice message from a second user; When the fifth voice includes a wake-up-free instruction, in response to the fifth voice, the first wake-up-free logo is stopped from being displayed, and a composite wake-up-free logo is displayed on the first display screen, where the composite wake-up-free logo is used to prompt the user that the first display screen is processing multiple wake-up-free instructions.

12. The method according to claim 11, characterized in that The method further comprises: When the fifth voice includes a wake-up word, in response to the fifth voice, the first wake-up-free logo is stopped from being displayed, and a third dialogue logo and a voice sub-logo are displayed on the first display screen, the third dialogue logo is used to prompt the user that the first display screen is performing voice interaction with the second user, and the voice sub-logo is used to prompt the user that the first display screen is processing multiple voices.

13. The method according to claim 11 or 12, characterized in that: The interior of the vehicle includes a first sound zone and a second sound zone, the first user is located in the first sound zone, and the second user is located in the second sound zone; the first display screen is used to respond to voices from the first sound zone and the second sound zone; The displaying a first wake-up-free mark on the first display screen specifically includes: Determining based on the fourth voice that the fourth voice comes from the first sound zone; Displaying a first wake-up-free mark on the first display screen based on the first sound zone; The displaying of the composite wake-up-free mark on the first display screen specifically includes: determining, based on the fifth voice, that the fifth voice comes from the second sound area; A composite wake-up-free logo is displayed on the first display screen based on the second sound zone.

14. The method according to claim 12, characterized in that The interior of the vehicle includes a first sound zone and a second sound zone, the first user is located in the first sound zone, and the second user is located in the second sound zone; the first display screen is used to respond to voices from the first sound zone and the second sound zone; The displaying a first wake-up-free mark on the first display screen specifically includes: Determining based on the fourth voice that the fourth voice comes from the first sound zone; Displaying a first wake-up-free mark on the first display screen based on the first sound zone; The displaying of the third dialogue identifier and the voice sub-identifier on the first display screen specifically includes: determining, based on the fifth voice, that the fifth voice comes from the second sound area; A third dialogue identifier and a voice sub-identifier are displayed on the first display screen based on the second voice zone.

15. The method according to claim 13 or 14, characterized in that The method further comprises: When the fifth voice includes the wake-up word, in response to the fifth voice, a fourth indicator is displayed on the first display screen, wherein the fourth indicator is used to indicate positions of the first sound zone and the second sound zone relative to the display screen.

16. A voice interaction method, characterized in that: Applied to a vehicle, the vehicle comprises a first display screen and a second display screen, and the interior of the vehicle comprises a first sound zone and a second sound zone; The first display screen is used to respond to speech from the first sound zone, and the second display screen is used to respond to speech from the second sound zone; the method comprises: receiving a sixth voice emitted by a first user from the first voice zone, wherein the sixth voice includes a wake-up word; In response to the sixth voice, determining based on the sixth voice that the sixth voice comes from the first sound zone; Displaying a third dialogue mark on the first display screen based on the first sound zone, wherein the third dialogue mark is used to prompt the user that the first display screen is performing voice interaction with the first user; receiving a seventh voice emitted by a second user from the second voice zone; When the seventh voice includes a wake-up-free instruction, in response to the seventh voice, determining that the seventh voice comes from the second sound zone based on the seventh voice; A second wake-up-free mark is displayed on the second display screen based on the second sound zone, where the second wake-up-free mark is used to prompt the user that the second display screen is processing a wake-up-free instruction.

17. The method according to claim 16, characterized in that The method further comprises: When the seventh voice includes the wake-up word, in response to the seventh voice, stopping displaying the third dialogue identifier on the first display screen; Determining, based on the seventh voice, that the seventh voice comes from the second voice area; A fourth dialogue identifier is displayed on the second display screen based on the second sound zone, and the fourth dialogue identifier is used to prompt the user that the second display screen is performing voice interaction with the second user.

18. A voice interaction method, characterized in that: Applied to a vehicle, the vehicle comprises a first display screen and a second display screen, and the interior of the vehicle comprises a first sound zone and a second sound zone; The first display screen is used to respond to speech from the first sound zone, and the second display screen is used to respond to speech from the second sound zone; the method comprises: receiving an eighth voice emitted by a first user from the first sound zone, wherein the eighth voice includes a wake-up-free instruction; In response to the eighth voice, determining based on the eighth voice that the eighth voice comes from the first sound zone; Displaying a third wake-up-free mark on the first display screen based on the first sound zone, wherein the third wake-up-free mark is used to prompt the user that the first display screen is processing a wake-up-free instruction; receiving a ninth voice emitted by a second user from the second voice zone; When the ninth voice includes a wake-up-free instruction, in response to the ninth voice, determining based on the ninth voice that the ninth voice comes from the second sound zone; A fourth wake-up-free mark is displayed on the second display screen based on the second sound zone, and the fourth wake-up-free mark is used to prompt the user that the second display screen is processing a wake-up-free instruction.

19. The method according to claim 18, characterized in that The method further comprises: When the ninth voice includes a wake-up word, in response to the ninth voice, determining based on the ninth voice that the ninth voice comes from the second sound zone; A fifth dialogue identifier is displayed on the second display screen based on the second sound zone, and the fifth dialogue identifier is used to prompt the user that the second display screen is performing voice interaction with the second user.

20. A vehicle, characterized in that: include: One or more processors, one or more memories; the one or more memories are coupled to the one or more processors, the one or more memories are used to store computer program code, the computer program code includes computer instructions, when the one or more processors execute the computer instructions, the vehicle executes the method described in any one of claims 1-19.

21. A computer-readable storage medium comprising computer instructions, characterized in that: When the computer instructions are executed on a vehicle, the vehicle is caused to execute the method according to any one of claims 1 to 19.

Citation Information

Patent Citations

  • Voice interaction method, device and equipment and computer readable storage medium

    CN115424623A

  • Voice interaction configuration method, electronic equipment and computer readable medium

    CN115705844A

  • Intelligent voice interaction processing method and mobile terminal

    CN115910052A

  • Voice control method, device, system and equipment for vehicle-mounted equipment and storage medium

    CN115985295A

  • Multi-tone-area voice dialogue flow interaction method, device and equipment, and vehicle

    CN116129892A