Voice interaction method, apparatus, and vehicle
By determining the location and legitimacy of voice commands within the vehicle, the security risks of external voice wake-up are resolved, achieving a balance between security and cost-effectiveness.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- YINWANG INTELLIGENT TECHNOLOGIES CO LTD
- Filing Date
- 2024-11-27
- Publication Date
- 2026-06-04
AI Technical Summary
Existing external voice wake-up technology poses security risks, as voiceprint recognition models are easily compromised, potentially leading to vehicle safety and user property security issues.
The system determines whether to execute an operation by judging whether the vehicle's location is within a preset safe area and whether the voice command is within a preset set of voice commands, thus avoiding the need to configure a voiceprint recognition model.
It reduces the safety risks when vehicles respond to voice commands, protects users' property, and reduces manpower and material resources, thus lowering vehicle costs.
Smart Images

Figure CN2024135034_04062026_PF_FP_ABST
Abstract
Description
Voice interaction methods, devices and vehicles Technical Field
[0001] This application relates to the field of intelligent vehicles, and more specifically, to a voice interaction method, device, and vehicle. Background Technology
[0002] External voice wake-up is a current hot topic. There are two main approaches: one is to omits external wake-up protection in the vehicle, allowing users to wake it up and interact with it by issuing voice commands from outside. This poses a significant risk to vehicle security and user property safety. The other approach involves installing a voiceprint recognition model in the vehicle, enabling it to respond to voice commands when the voiceprint is recognized. This method, too, may be vulnerable to voiceprint compromise, posing a further risk to vehicle security and user property safety. Summary of the Invention
[0003] This application provides a voice interaction method, device, and vehicle that helps reduce safety risks when a vehicle responds to voice commands and also helps protect the user's property.
[0004] In a first aspect, this application provides a voice interaction method, which includes: acquiring a first voice command; and executing the operation corresponding to the first voice command when the vehicle is located within a first preset safety area and the first voice command is an command within a first preset set of voice commands.
[0005] Based on the above technical solution, when a vehicle executes a voice command, it only needs to determine whether the vehicle is within a preset safe area and whether the voice command is within a preset set of voice commands. When the vehicle is within the preset safe area and the voice command is within the preset set of voice commands, the operation corresponding to the first voice command can be executed. This helps ensure that the vehicle's execution of the voice command will not pose a risk to the vehicle's safety or the user's property. At the same time, there is no need to configure a voiceprint recognition model in the vehicle, avoiding the manpower and material resources required to train such a model, and also helping to reduce the cost of the vehicle.
[0006] In some possible implementations, when the vehicle is located within a first preset safety area and the first voice command is an instruction within a first preset voice command set, the operation corresponding to the first voice command is executed, including: when the first voice command originates from outside the cabin, the vehicle is located within a first preset safety area, and the first voice command is an instruction within a first preset voice command set, the operation corresponding to the first voice command is executed.
[0007] Based on the above technical solution, when the vehicle is located within a preset safe area and the voice command is within a preset set of voice commands, the operation corresponding to the voice command outside the cabin can be executed. This helps to ensure that the vehicle will not pose a risk to the safety of the vehicle or the safety of the user's property after the vehicle executes the voice command outside the cabin. At the same time, there is no need to configure a voiceprint recognition model in the vehicle, avoiding the consumption of manpower and material resources for training the voiceprint model, and also helping to reduce the cost of the vehicle.
[0008] In some possible implementations, when the vehicle is located within a first preset safety area and the first voice command is an instruction within a first preset voice command set, the operation corresponding to the first voice command is executed, including: when the moment the first voice command is obtained is within a preset time period, the vehicle is located within the first preset safety area and the first voice command is an instruction within a first preset voice command set, the operation corresponding to the first voice command is executed.
[0009] For example, the preset time period is from 7 a.m. to 10 p.m.
[0010] In some possible implementations, acquiring the first voice command includes: acquiring a voice signal collected by a microphone in the cockpit; and determining the first voice command based on the voice signal.
[0011] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: when the vehicle is not located within the first preset safety area, determining the response to the first voice command based on the source of the first voice command.
[0012] Based on the above technical solution, when the vehicle is not located in the preset safe area, it can be determined whether to respond to the first voice command based on the source of the first voice command.
[0013] In some possible implementations, the vehicle's location not being within the preset safety zone can be understood as the vehicle being located outside the preset safety zone.
[0014] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: when the first voice command originates from inside the cockpit, executing the operation corresponding to the first voice command; or, when the first voice command originates from outside the cockpit, ignoring the first voice command; or, prompting the user to refuse to execute the operation corresponding to the first voice command.
[0015] Based on the above technical solution, when the vehicle is not located within a preset safe area, if the voice command is detected as originating from a user inside the cabin, the operation corresponding to the first voice command can be executed; if the voice command is detected as originating from a user outside the cabin, the first voice command can be ignored or the user can be prompted not to execute the operation corresponding to the first voice command. This helps ensure that executing the voice command does not pose a risk to the vehicle's safety or the user's property.
[0016] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: when the vehicle is located within the first preset safety area and the first voice command is not an command within the first preset voice command set, determining the response to the first voice command based on the source of the first voice command.
[0017] Based on the above technical solution, when the vehicle is located within a preset safe area and the first voice command is not a command in the preset voice command set, it can be determined whether to respond to the first voice command based on the source of the first voice command.
[0018] In conjunction with the first aspect, in some implementations of the first aspect, determining the response to the first voice command based on its source includes: executing the operation corresponding to the first voice command when the first voice command originates from inside the cockpit; or ignoring the first voice command when the first voice command originates from outside the cockpit; or prompting the user to refuse to execute the operation corresponding to the first voice command.
[0019] In conjunction with the first aspect, in some implementations of the first aspect, before executing the operation corresponding to the first voice command, the method further includes: determining that the voiceprint information of the first voice command matches one of the one or more voiceprint information stored in the vehicle.
[0020] Based on the above technical solution, when the vehicle is within a preset safe area, the voice command is from a preset set of voice commands, and the voiceprint information of the voice command matches one of the one or more voiceprint information stored in the vehicle, the operation corresponding to the first voice command can be executed. This helps to further ensure that the vehicle's execution of the voice command will not pose a risk to the vehicle's safety or the user's property.
[0021] In conjunction with the first aspect, in some implementations of the first aspect, the first preset safety area includes one or more sub-areas. When the vehicle is located within the first preset safety area and the first voice command is an instruction within the first preset voice command set, the operation corresponding to the first voice command is executed, including: when the vehicle is located within a first sub-area of the first preset safety area, obtaining the first preset voice command set corresponding to the first sub-area according to a mapping relationship, wherein the mapping relationship includes the association between each sub-area of one or more sub-areas and the preset voice command set corresponding to each sub-area, and the one or more sub-areas include the first sub-area; and when the first voice command is an instruction within the first preset command set, executing the operation corresponding to the first voice command.
[0022] Based on the above technical solution, the vehicle stores the association between preset safe areas and preset voice command sets. Thus, when the vehicle is in a first preset safe area, the first preset voice command set corresponding to that first preset safe area can be obtained based on this mapping relationship. When a first voice command is determined to be an instruction in the first preset voice command set, the operation corresponding to that first voice command can be executed. In this way, different preset voice command sets can be available for different preset safe areas, which helps improve the flexibility of voice interaction between the user and the vehicle, and can also meet the user's needs when the vehicle is in different preset safe areas, thus improving the user's human-machine interaction experience.
[0023] In conjunction with the first aspect, in some implementations of the first aspect, when the vehicle is located within a first preset safety area and the first voice command is an instruction within a first preset voice command set, before executing the operation corresponding to the first voice command, the method further includes: controlling a display device to display information about the first preset safety area and multiple components; in response to a user's input to select at least one component from the multiple components, establishing an association between the first preset safety area and the first voice command set, the first voice command set including operation instructions for the at least one component.
[0024] Based on the above technical solution, users can select controllable components within a first preset safety area via a display device in the vehicle. Taking the first preset safety area as the user's company and at least one component including the trunk and seats as an example, after determining that the vehicle's location is the company and the user issues a voice command to control the trunk or seats, the vehicle can execute the operation corresponding to the voice command.
[0025] In conjunction with the first aspect, in some implementations of the first aspect, when the vehicle is located within a first preset safety area and the first voice command is an instruction within a first preset instruction set, before executing the operation corresponding to the first voice command, the method further includes: a control prompting device prompting the user to input a voice command to be executed within the first preset safety area; and in response to one or more voice commands input by the user, establishing an association between the first preset safety area and the first voice command set, wherein the first voice command set includes the one or more voice commands.
[0026] Based on the above technical solution, a prompting device in the vehicle prompts the user to input voice commands executed within a first preset safety zone, thereby establishing a correlation between the first preset safety zone and the user-input voice commands. In this way, users can "encrypt" their voice interactions with the vehicle by inputting personalized voice commands. That is, when the vehicle is within the first preset safety zone, it can respond to the user's personalized voice commands, further ensuring that executing voice commands outside the cabin will not pose a risk to the vehicle's safety or the user's property.
[0027] In conjunction with the first aspect, in some implementations of the first aspect, before executing the operation corresponding to the first voice instruction, the method further includes: executing the operation corresponding to the first voice instruction when the first text content corresponding to the first voice instruction matches the second text content corresponding to the second voice instruction among the one or more voice instructions.
[0028] Based on the above technical solution, when the first voice command acquired by the vehicle matches the previously set personalized voice command, the operation corresponding to the first voice command can be executed, which helps to further ensure that the vehicle will not pose a risk to the safety of the vehicle or the safety of the user's property after executing the voice command outside the cabin.
[0029] In conjunction with the first aspect, in some implementations of the first aspect, before executing the operation corresponding to the first voice command, the method further includes: determining, based on data collected by sensors outside the cockpit, that a user exists within the first preset safe area.
[0030] Based on the above technical solution, when a user is present within the first preset safety zone, the operation corresponding to the voice command can be executed. This avoids the risk of the vehicle responding to a voice command issued by a user outside the first preset safety zone.
[0031] In some possible implementations, a microphone is installed outside the vehicle's cabin. Before executing the operation corresponding to the first voice command, the method further includes: determining, based on data collected by sensors outside the cabin, that a user exists within the first preset safety area.
[0032] In conjunction with the first aspect, in some implementations of the first aspect, before determining the response to the first voice command based on its source, the method further includes: inputting the first voice command into a prediction model to obtain the source of the first voice command, wherein the source of the first voice command indicates that the first voice command originates from inside or outside the vehicle's cabin.
[0033] Based on the above technical solution, the source of the first voice command can be predicted using a prediction model.
[0034] In conjunction with the first aspect, in some implementations of the first aspect, before determining the response to the first voice command based on its source, the method further includes: acquiring a first voice signal corresponding to the first voice command collected by a microphone in the cockpit; mapping the energy of the first voice signal within a preset energy range to obtain a second voice signal; filtering the second voice signal to obtain a third voice signal; and determining the source of the first voice command based on the third voice signal, wherein the source of the first voice command indicates that the first voice command originates from inside or outside the vehicle's cockpit.
[0035] Based on the above technical solution, the first voice signal corresponding to the first voice command can be processed by energy normalization and filtering, so as to determine the source of the first voice command based on the obtained third voice signal.
[0036] Secondly, this application provides a voice interaction device, which includes: an acquisition unit for acquiring a first voice command; and an instruction execution unit for executing the operation corresponding to the first voice command when the vehicle is located within a first preset safety area and the first voice command is an instruction within a first preset set of voice commands.
[0037] In conjunction with the second aspect, in some implementations of the second aspect, the instruction execution unit is used to determine the response to the first voice instruction based on the source of the first voice instruction when the vehicle is not located within the first preset safety area.
[0038] In conjunction with the second aspect, in some implementations of the second aspect, the instruction execution unit is used to determine the response to the first voice instruction based on the source of the first voice instruction when the vehicle is located within the first preset safety area and the first voice instruction is not an instruction within the first preset voice instruction set.
[0039] In conjunction with the second aspect, in some implementations of the second aspect, the device further includes a determining unit, which is configured to determine, before the instruction execution unit executes the operation corresponding to the first voice instruction, that the voiceprint information of the first voice instruction matches one of the one or more voiceprint information stored in the vehicle.
[0040] In conjunction with the second aspect, in some implementations of the second aspect, the first preset safety zone includes one or more sub-regions. The acquisition unit is further configured to, when the vehicle is located within a first sub-region of the first preset safety zone, acquire the first preset voice command set corresponding to the first sub-region according to a mapping relationship. The mapping relationship includes the association between each sub-region of the one or more sub-regions and the preset voice command set corresponding to each sub-region. The one or more sub-regions include the first sub-region. The command execution unit is configured to, when the first voice command is an command within the first preset command set, execute the operation corresponding to the first voice command.
[0041] In conjunction with the second aspect, in some implementations of the second aspect, the device further includes a control unit and a mapping relationship establishment unit. The control unit is configured to control the display device to display information about the first preset safe area and multiple components before the instruction execution unit executes the operation corresponding to the first voice instruction. The mapping relationship establishment unit is configured to establish an association between the first preset safe area and the first voice instruction set in response to the user's input of selecting at least one component from the multiple components. The first voice instruction set includes operation instructions for the at least one component.
[0042] In conjunction with the second aspect, in some implementations of the second aspect, the device further includes a control unit and a mapping relationship establishment unit. The control unit is used to control the prompting device to prompt the user to input a voice command to be executed in the first preset safe area before the instruction execution unit executes the operation corresponding to the first voice command. The mapping relationship establishment unit is used to establish an association between the first preset safe area and the first voice command set in response to one or more voice commands input by the user. The first voice command set includes the one or more voice commands.
[0043] In conjunction with the second aspect, in some implementations of the second aspect, the instruction execution unit is configured to: execute the operation corresponding to the first voice instruction when the first text content corresponding to the first voice instruction matches the second text content corresponding to the second voice instruction among the one or more voice instructions.
[0044] In conjunction with the second aspect, in some implementations of the second aspect, the device further includes a determining unit, which is used to determine, based on data collected by sensors outside the cockpit, that a user exists within the first preset safe area before the instruction execution unit executes the operation corresponding to the first voice instruction.
[0045] In conjunction with the second aspect, in some implementations of the second aspect, the device further includes: a prediction unit, configured to input the first voice command into a prediction model to obtain the source of the first voice command, wherein the source of the first voice command indicates that the first voice command originates from inside or outside the vehicle's cabin.
[0046] In conjunction with the second aspect, in some implementations of the second aspect, the device further includes an energy mapping unit, a filtering unit, and a determining unit. The acquiring unit is used to acquire a first voice signal corresponding to the first voice command collected by a microphone in the cockpit. The energy mapping unit is used to map the energy of the first voice signal within a preset energy range to obtain a second voice signal. The filtering unit is used to filter the second voice signal to obtain a third voice signal. The determining unit is used to determine the source of the first voice command based on the third voice signal, wherein the source of the first voice command indicates that the first voice command originates from inside or outside the vehicle's cockpit.
[0047] In conjunction with the second aspect, in some implementations of the second aspect, the instruction execution unit is configured to: execute the operation corresponding to the first voice instruction when the first voice instruction originates from inside the cockpit; or, ignore the first voice instruction when the first voice instruction originates from outside the cockpit; or, prompt the user to refuse to execute the operation corresponding to the first voice instruction.
[0048] Thirdly, this application provides a voice interaction device, which includes a processor and a memory, wherein the memory is used to store instructions, and the processor executes the instructions stored in the memory to cause the device to perform any of the possible methods in the first aspect.
[0049] Fourthly, this application provides a voice interaction system, which includes a computing platform and a microphone, wherein the computing platform includes any of the possible devices in the second or third aspect.
[0050] Fifthly, this application provides a vehicle that includes any of the possible devices of the second or third aspect, or includes the system described in the fourth aspect.
[0051] In a sixth aspect, this application provides a computer program product comprising: computer program code, which, when executed on a computer, causes the computer to perform any of the possible methods described in the first aspect above.
[0052] It should be noted that the above-mentioned computer program code can be stored in whole or in part on the first storage medium, wherein the first storage medium can be packaged together with the processor or packaged separately from the processor. This application embodiment does not specifically limit this.
[0053] In a seventh aspect, this application provides a computer-readable medium storing program code that, when run on a computer, causes the computer to perform any of the possible methods described in the first aspect above.
[0054] Eighthly, this application provides a chip system including a processor for calling a computer program or computer instructions stored in a memory to cause the processor to perform any of the possible methods in the first aspect above.
[0055] In conjunction with the eighth aspect, in one possible implementation, the processor is coupled to the memory via an interface.
[0056] In conjunction with the eighth aspect, in one possible implementation, the chip system also includes a memory in which computer programs or computer instructions are stored.
[0057] Ninthly, this application provides a chip system including circuitry for performing any of the possible methods described in the first aspect above. Attached Figure Description
[0058] Figure 1 is a functional block diagram of the vehicle provided in an embodiment of this application.
[0059] Figure 2 is a schematic flowchart of the voice interaction method provided in an embodiment of this application.
[0060] Figures 3A-3C are a set of GUIs provided in the embodiments of this application.
[0061] Figures 4A-4F show another set of GUIs provided in the embodiments of this application.
[0062] Figures 5A-5C show another set of GUIs provided in the embodiments of this application.
[0063] Figures 6A-6B are another set of GUIs provided in the embodiments of this application.
[0064] Figures 7A-7C show another set of GUIs provided in the embodiments of this application.
[0065] Figure 8 is a schematic flowchart of the voice interaction method provided in the embodiments of this application.
[0066] Figure 9 is a schematic flowchart of the voice interaction method provided in an embodiment of this application.
[0067] Figure 10 is a schematic block diagram of the voice interaction device provided in an embodiment of this application. Detailed Implementation
[0068] The technical solutions in this application will now be described with reference to the accompanying drawings.
[0069] Figure 1 is a functional block diagram of a vehicle provided in an embodiment of this application. As shown in Figure 1, the vehicle 100 may include a computing platform 110, a microphone 120, and a display device 130. The display device 130 in the cabin is mainly divided into two categories: the first is an in-vehicle display screen; the second is a projection display screen, such as a head-up display (HUD). An in-vehicle display screen is a physical display screen and an important component of the in-vehicle infotainment system. Multiple displays can be installed in the cabin, such as a digital instrument cluster display screen, a central control screen, a display screen in front of the front passenger (also known as the front row passenger), a display screen in front of the left rear passenger, and a display screen in front of the right rear passenger; even the car window can be used as a display screen. A head-up display, also known as a head-up display system, is mainly used to display driving information such as speed and navigation on a display device (e.g., the windshield) in front of the driver. This reduces the driver's eye movement time, avoids pupil changes caused by eye movement, and improves driving safety and comfort. HUDs include, for example, combiner-HUD (C-HUD) systems, windshield-HUD (W-HUD) systems, and augmented reality HUD (AR-HUD) systems.
[0070] Some or all of the functions of vehicle 100 can be controlled by computing platform 110. Computing platform 110 may include processors 111 to 11n. A processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction read and execute capabilities, such as a central processing unit (CPU), microprocessor, graphics processing unit (GPU) (which can be understood as a type of microprocessor), or digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. These logical relationships are fixed or reconfigurable. For example, the processor may be a hardware circuit implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as a field-programmable gate array (FPGA). In reconfigurable hardware circuits, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. Furthermore, the processor can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a neural network processing unit (NPU), tensor processing unit (TPU), deep learning processing unit (DPU), etc. In addition, the computing platform 110 may also include a memory for storing instructions. Some or all of the processors 111 to 11n can call the instructions in the memory to implement the corresponding functions.
[0071] Optionally, the structure of the vehicle 100 described above is merely illustrative. In actual applications, various components of the vehicle 100 may be added or removed as needed.
[0072] As mentioned earlier, external voice wake-up is a current hot topic. There are two current solutions: one is that the vehicle is not equipped with external wake-up prevention functionality, allowing the user to wake up the vehicle by issuing voice commands from outside and interact with it. This method poses a significant risk to vehicle security and the user's property. The other is to install a voiceprint recognition model in the vehicle, allowing the vehicle to respond to voice commands when the voiceprint recognition is successful. This method also carries the risk of voiceprint compromise, posing a certain risk to vehicle security and the user's property.
[0073] This application provides a voice interaction method, device, and vehicle. Based on whether the vehicle is located within a preset safe area and whether the voice command is part of a preset set of voice commands, the response to a voice command can be determined. This helps reduce the safety risks associated with the vehicle responding to voice commands and also helps protect the user's property.
[0074] Figure 2 shows a schematic flowchart of a voice interaction method 200 provided in an embodiment of this application. The method 200 includes:
[0075] S210, obtain the first voice command.
[0076] Optionally, acquiring the first voice command includes: acquiring the first voice command based on data collected by a microphone in the vehicle's cabin.
[0077] S220, when the vehicle is located within a first preset safety area and the first voice command is a command within a first preset voice command set, the operation corresponding to the first voice command is executed.
[0078] Optionally, when the vehicle is located within a first preset safety area and the first voice command is an instruction within a first preset voice command set, the operation corresponding to the first voice command is executed, including: when the moment the first voice command is obtained is within a preset time period, the vehicle is located within the first preset safety area and the first voice command is an instruction within a first preset voice command set, the operation corresponding to the first voice command is executed.
[0079] For example, the preset time period is from 7 a.m. to 10 p.m.
[0080] For example, the above preset time period can be customized by the user.
[0081] Optionally, before executing the operation corresponding to the first voice command, the method 200 further includes: obtaining user input for setting the first preset security area.
[0082] For example, when a user selects home as the preset safe zone in the frequently used location management section of a map application, the area corresponding to home can be set as part of the first preset safe zone.
[0083] For example, when a user in the cockpit issues a voice command "set home as a safe zone" and that user has the authority to set safe zones, the area corresponding to the home location can be set as part of the first preset safe zone.
[0084] For example, when a user inputs on a mobile terminal (e.g., a mobile phone) that their home is set as a safe zone, the phone can send a command to the vehicle instructing that the area corresponding to the home be designated as a safe zone. In response to receiving this command, the vehicle can set the area corresponding to the home as part of a first preset safe zone.
[0085] Taking the area corresponding to a home as an example, the scope of this area can be obtained from map information, or it can be constructed from data collected by the vehicle through sensors (such as one or more of cameras, lidar, or millimeter-wave radar).
[0086] Optionally, the first preset security area may include multiple sub-areas. For example, the first preset security area includes sub-area 1 and sub-area 2, where sub-area 1 is the area corresponding to home and sub-area 2 is the area corresponding to company.
[0087] Optionally, before executing the operation corresponding to the first voice command, the method 200 further includes: obtaining user input on setting the first preset voice command set.
[0088] For example, when a user inside the cabin issues a voice command to "add the seat adjustment command and window adjustment command to the set of voice commands outside the cabin" and the user has the authority to set voice commands outside the cabin, the seat adjustment command and window adjustment command can be added to the first preset voice command set.
[0089] For example, Table 1 shows the first preset security area and the first preset voice command set provided in the embodiments of this application.
[0090] Table 1
[0091] For example, if the vehicle's current location is detected to be within the area corresponding to the home and the user issues voice command 1 "open the window," the vehicle can respond to the voice command 1 and execute the operation of opening the window. Thus, even if the voice command 1 is issued by the user outside the cabin, because the vehicle is located within the first preset safety area and the voice command 1 is within the first preset voice command set, the vehicle can execute the operation corresponding to the voice command 1. This ensures both vehicle safety and the user's property safety without requiring the vehicle to be equipped with a voiceprint recognition model.
[0092] In this embodiment, when a vehicle executes a voice command, it only needs to determine whether the vehicle is within a preset safe area and whether the voice command is within a preset set of voice commands. When the vehicle is located within the preset safe area and the voice command is within the preset set of voice commands, the operation corresponding to the first voice command can be executed. This helps ensure that the vehicle's execution of the voice command will not pose a risk to the vehicle's safety or the user's property. At the same time, there is no need to configure a voiceprint recognition model in the vehicle, avoiding the manpower and material resources required to train such a model, and also helping to reduce the cost of the vehicle.
[0093] Optionally, the method 200 further includes: when the vehicle is not located within the first preset safety area, determining the response to the first voice command based on the source of the first voice command.
[0094] The fact that the vehicle is not located within the first preset safety area can also be understood as the vehicle being located outside the first preset safety area.
[0095] Optionally, determining the response to the first voice command based on its source includes: executing the operation corresponding to the first voice command when the first voice command originates from inside the cockpit; or ignoring the first voice command when the first voice command originates from outside the cockpit; or prompting the user to refuse to execute the operation corresponding to the first voice command.
[0096] In this embodiment, when the vehicle is not located within the first preset safety area, the source of the first voice command can be determined. If the first voice command originates from outside the cabin, it can be ignored or the user can be prompted to refuse to execute the operation corresponding to the first voice command; or, if the first voice command originates from inside the cabin, the operation corresponding to the first voice command can be executed.
[0097] Optionally, the method 200 further includes: when the vehicle is located within the first preset safety area and the first voice command is not an instruction within the first preset voice command set, determining the response to the first voice command based on the source of the first voice command.
[0098] Optionally, before determining the response to the first voice command based on its source, the method 200 further includes: inputting the first voice command into a prediction model to obtain the source of the first voice command, wherein the source of the first voice command indicates that the first voice command originates from inside or outside the vehicle's cabin.
[0099] For example, the prediction model can be trained from a training dataset that includes sample speech signals and corresponding label information, the label information including the source of the sample speech signals.
[0100] Optionally, before determining the response to the first voice command based on its source, the method 200 further includes: acquiring a first voice signal corresponding to the first voice command collected by a microphone in the cabin; mapping the energy of the first voice signal within a preset energy range to obtain a second voice signal; filtering the second voice signal to obtain a third voice signal; and determining the source of the first voice command based on the third voice signal, wherein the source of the first voice command indicates that the first voice command originates from inside or outside the vehicle cabin.
[0101] For example, when the vehicle is not located within the first preset safety area, or when the vehicle is located within the first preset safety area and the first voice command is not a command in the first preset voice command set, the first voice signal corresponding to the first voice command can be subjected to energy normalization processing and filtering processing. By analyzing the processed third voice signal, it can be determined whether the first voice command originates from inside or outside the cabin.
[0102] For example, taking the acquisition of voice signals by a microphone in the vehicle cabin as an example, the energy normalization process can be as follows:
[0103] (1) Real-time detection of voice signals received by the microphone in the cockpit 1.
[0104] (2) Calculate the actual energy of the speech signal 1 and map the actual energy to a preset energy range to obtain the speech signal 2 after energy mapping.
[0105] For example, this energy normalization process can map energy from [5dB, 100dB] to the range of [20dB, 60dB]. For instance, if the actual collected energy is 100dB, the mapped energy could be 60dB. Similarly, if the actual collected energy is 5dB, the mapped energy could be 20dB.
[0106] Taking a parked vehicle as an example, with all windows closed, the vehicle can filter out the high-frequency components of external sounds, retaining only the low-frequency components of external sounds and the sounds inside the cabin. This energy normalization process aims to prevent users outside the cabin from speaking loudly and forcing the vehicle to respond.
[0107] For example, the filtering process can be as follows:
[0108] (1) Filter out noise and echo in the speech signal 2 after energy mapping.
[0109] (2) Low-frequency (e.g., 200-1100Hz) signals are filtered by bandpass filtering to obtain the filtered speech signal 3.
[0110] By selecting a certain bandpass filter range, the frequency band of human voices inside the vehicle can be preserved while filtering out most of the human voice energy outside the cabin.
[0111] Through the above energy normalization and filtering processes, it is possible to determine whether the voice signal 1 comes from inside or outside the cockpit based on the filtered voice signal 3.
[0112] For example, voice signal 3 can be input into the wake-up model to obtain the source of voice signal 1. If voice signal 1 originates from outside the cockpit, it can be ignored or the user can be prompted not to perform the operation corresponding to voice signal 1; if voice signal 1 originates from inside the cockpit, the operation corresponding to voice signal 1 can be performed.
[0113] For example, the wake-up model can be trained using a training dataset, which may include sample data. For instance, the sample data may include multiple sample speech signals carrying a wake-up word and label information corresponding to each sample speech signal, which is used to indicate whether each sample speech signal originates from inside or outside the cockpit.
[0114] Optionally, before determining the response to the first voice command based on its source, the method 200 further includes: inputting the first voice command into a prediction model to obtain the source of the first voice command.
[0115] Optionally, taking a vehicle that includes microphones outside the cabin and microphones inside the cabin as an example, before determining the response to the first voice command based on its source, the method 200 further includes: determining the source of the first voice signal based on voice signals collected by the microphones inside and outside the cabin.
[0116] The above methods for determining the source of the first voice command are merely exemplary, and the methods for determining the source of the first voice command in this application embodiment are not specifically limited.
[0117] Optionally, the first preset safety zone includes one or more sub-regions. When the vehicle is located within the first preset safety zone and the first voice command is an instruction within the first preset voice command set, the operation corresponding to the first voice command is executed, including: when the vehicle is located within a first sub-region of the first preset safety zone, obtaining the first preset voice command set corresponding to the first sub-region according to a mapping relationship, wherein the mapping relationship includes the association between each sub-region of one or more sub-regions and the preset voice command set corresponding to each sub-region, and the one or more sub-regions include the first sub-region; and when the first voice command is an instruction within the first preset command set, executing the operation corresponding to the first voice command.
[0118] For example, the first preset security zone may consist of one or more sub-zones. For instance, sub-zone 1 is the zone corresponding to home, and sub-zone 2 is the zone corresponding to company.
[0119] For example, Table 2 shows the mapping relationships provided in the embodiments of this application.
[0120] Table 2
[0121] For example, a vehicle can determine its current location within a sub-region corresponding to its home based on data collected by positioning sensors. The vehicle can then determine its corresponding preset voice command set, which includes voice commands for controlling the doors, windows, seats, air conditioning, and trunk. After acquiring a first voice command via a microphone in the cabin, the vehicle can determine whether the first voice command is for controlling the doors, windows, seats, air conditioning, or trunk. If it is, the vehicle can respond to the voice command; otherwise, the vehicle can continue to determine the source of the first voice command. If the first voice command originates from outside the cabin, the operation corresponding to the first voice command is not executed; or, if the first voice command originates from inside the cabin, the operation corresponding to the first voice command can be executed.
[0122] For example, a vehicle can determine its current location within a sub-region corresponding to its home based on data collected by positioning sensors. The vehicle can then determine its corresponding preset voice command set, which includes voice commands for controlling the doors, windows, seats, air conditioning, and trunk. After receiving a first voice command via the microphone, the vehicle can first determine whether the command originates from inside or outside the cabin. If it originates from inside the cabin, the vehicle directly executes the operation corresponding to the first voice command. If it originates from outside the cabin, the vehicle can determine whether the first voice command controls the doors, windows, seats, air conditioning, or trunk. If it does, the vehicle can respond to the command; otherwise, the vehicle can either not execute the operation corresponding to the first voice command or prompt the user to refuse to execute it.
[0123] The above describes the process by which the vehicle determines whether to execute the corresponding operation after receiving a voice command. The following section, using an example, describes the process of establishing a safe zone set by the user and the relationship between preset voice commands.
[0124] Optionally, the vehicle stores the association between the first preset safe area and the first set of voice commands. When the vehicle is located within the first preset safe area and the first voice command is a command within the first preset set of voice commands, the operation corresponding to the first voice command is executed, including: when it is determined based on the association that the vehicle is currently located within the first preset safe area and the first voice command is a command within the first preset set of voice commands, the operation corresponding to the first voice command is executed.
[0125] Optionally, when the vehicle is located within a first preset safety area and the first voice command is an instruction within a first preset voice command set, before executing the operation corresponding to the first voice command, the method 200 further includes: controlling a display device to display information about the first preset safety area and multiple components; and in response to a user's input of selecting at least one component from the multiple components, establishing an association between the first preset safety area and the first voice command set, wherein the first voice command set includes operation instructions for the at least one component.
[0126] For example, Figures 3A-3C illustrate a set of graphical user interfaces (GUIs) provided in embodiments of this application.
[0127] As shown in Figure 3A, the vehicle can display the map application interface on the central control screen, which includes information about the vehicle's current location, such as "xx shopping mall charging area". In response to a user's voice command, "Hey A, Hey A, set up the external wake-up function here," the vehicle can display the GUI shown in Figure 3B on the central control screen.
[0128] As shown in Figure B, in response to the voice command, the vehicle can display a prompt box 301, information on multiple components (e.g., air conditioning, doors, windows, trunk, and seats), a confirmation control 302, a cancellation control, and a memory switch on the central control screen. The prompt box 301 includes the message "Currently in the charging area of xx shopping mall. Please select a component that can be controlled from outside the cabin." When the system detects that the user has selected the trunk from multiple components and clicks the confirmation control 302, the vehicle can establish an association between the area corresponding to that location (xx shopping mall charging area) and the voice command related to controlling the trunk.
[0129] As shown in Figure C, when the vehicle is detected to be in the charging area of xx shopping mall and the user issues the voice command "Xiao A Xiao A, open the trunk", the vehicle can control the trunk to open.
[0130] Optionally, when the memory switch is in the on state, the vehicle can save the association between the location (xx shopping mall charging area) and the voice commands related to controlling the trunk. Thus, when the user drives the vehicle into the xx shopping mall charging area again, if the vehicle detects that it is in the xx shopping mall charging area and detects the user issuing the voice command "Xiao A, Xiao A, open the trunk," the vehicle can control the trunk to open.
[0131] Optionally, when the memory switch is in the off state, after detecting that the vehicle has left the location, the association between the location (xx shopping mall charging area) and the voice commands related to controlling the trunk can be deleted.
[0132] Optionally, the first preset safe area includes a first sub-area and a second sub-area. The control display device displays information about the first preset safe area and multiple components, including: the control display device displays information about the multiple sub-areas and multiple components; in response to a user's input to select at least one component from the multiple components, an association relationship is established between the first preset safe area and the first voice command set, including: in response to a user's input to associate the first sub-area with the first component set and to associate the second sub-area with the second component set, an association relationship is established between the first sub-area and the second preset voice command set, and an association relationship is established between the second sub-area and the third preset voice command set, wherein the second preset voice command set includes operation instructions for components in the first component set, and the third preset voice command set includes operation instructions for components in the second component set.
[0133] For example, Figures 4A-4F illustrate another set of GUIs provided in embodiments of this application.
[0134] As shown in Figure 4A, the vehicle can display the map application's display interface 401 through the central control screen. The display interface 401 includes a function bar 402, which includes vehicle-related controls, sharing controls, and settings controls 403.
[0135] As shown in Figure 4B, in response to the user clicking the control 403, the vehicle can display the settings interface of the map application on the central control screen. The settings interface includes voice navigation, frequently used location management, skins, and my vehicle.
[0136] As shown in Figure 4C, in response to the user clicking the control 404 corresponding to the frequently used location management, the vehicle can display the frequently used location management interface on the central control screen. The display interface includes information on multiple frequently used locations set by the user, a control for adding frequently used locations, a control 405 for setting the vehicle external wake-up function, and a commuting setting control. Among them, multiple frequently used locations may include home (xx residential area) and company (xx research institute underground parking lot).
[0137] For example, the first sub-region can be the region corresponding to a home, and the second sub-region can be the region corresponding to a company.
[0138] As shown in Figure 4D, in response to the user's click on the control 405, the vehicle can display the interface for setting the external wake-up function on the central control screen. This interface is used to set the association between each frequently used location and the corresponding component.
[0139] For example, when it is detected that the user has selected the components corresponding to home as air conditioning, car doors, car windows, trunk and seats, and the components corresponding to company as air conditioning and car windows, the relationship shown in Table 3 can be established.
[0140] Table 3
[0141] For example, the second preset voice command set includes voice commands for controlling the air conditioner, doors, windows, trunk, and seats, and the third preset voice command set includes voice commands for controlling the seats and windows.
[0142] As shown in Figures 4E and 4F, when the vehicle's current location is detected to be within the company's corresponding safe zone and the user's voice command "Xiao A, Xiao A, open the window" is received, the vehicle can perform the operation of opening the window.
[0143] Optionally, when the vehicle is located within a first preset safety area and the first voice command is an instruction within a first preset instruction set, before executing the operation corresponding to the first voice command, the method 200 further includes: a control prompting device prompting the user to input a voice command to be executed in the first preset safety area; and in response to one or more voice commands input by the user, establishing an association between the first preset safety area and the first voice command set, wherein the first voice command set includes the one or more voice commands.
[0144] For example, Figures 5A-5C illustrate another set of GUIs provided in embodiments of this application.
[0145] As shown in Figure 5A, the vehicle can display an interface for setting the external wake-up function on the central control screen. This interface is used to set the association between the corresponding areas and components of the vehicle. These components include the driver's side window, passenger side window, driver's side door, and passenger side door. Each component can also include controls for setting personalized voice commands and controls for managing voice command sets. In response to the user clicking the control 501 for setting personalized voice commands corresponding to the driver's side window, the vehicle can receive the voice commands issued by the user.
[0146] As shown in Figure 5B, when the vehicle receives the user's voice command "Xiao A Xiao A, open the queen-side window", it can establish a relationship between the company's corresponding region, the personalized voice command "open the queen-side window", and the operation of opening the driver's side window.
[0147] As shown in Figure 5C, the vehicle can also respond to the user with the voice command, "Okay, the personalized voice command for the driver's side window has been saved."
[0148] For example, when the vehicle is detected to be within the company's designated area and the user issues the voice command "Xiao A, Xiao A, open the window in the queen's seat", the vehicle can control the driver's side window to open.
[0149] For example, in response to a user clicking the control 501 for setting personalized voice commands on the passenger-side window, the vehicle can receive the user's voice command. Upon receiving the user's voice command "Makka Pakka," the vehicle can establish an association between the company's corresponding region, the personalized voice command "Makka Pakka," and the operation of opening the passenger-side window. For example, when the vehicle's location is detected to be within the company's corresponding region and the user has issued the voice command "Makka Pakka," the vehicle can control the passenger-side window to open.
[0150] In this embodiment, the user can be prompted by a prompting device in the vehicle to input voice commands executed within a first preset safe zone, thereby establishing a connection between the first preset safe zone and the user-input voice commands. This allows the user to "encrypt" the voice interaction between themselves and the vehicle by inputting personalized voice commands. In other words, the vehicle can respond to the user's personalized voice commands when it is within the first preset safe zone, further ensuring that the execution of voice commands outside the cabin will not pose a risk to the vehicle's safety or the user's property.
[0151] Optionally, before executing the operation corresponding to the first voice command, the method 200 further includes: executing the operation corresponding to the first voice command when the first text content corresponding to the first voice command matches the second text content corresponding to the second voice command among the one or more voice commands.
[0152] For example, when the vehicle is detected to be within the company's corresponding area and the user issues the voice command "Open the Queen's window", the vehicle can determine that the matching degree between the voice command and the previously saved voice command "Open the Queen's window" meets the preset conditions, and at this time the vehicle can also control the driver's side window to open.
[0153] Optionally, the preset condition includes an overlap rate of 80% or greater between the first text content corresponding to the user's voice command and the second text content of the voice command stored by the vehicle. This overlap rate can be equal to the number of characters in the first text content that overlap with the second text content divided by the number of characters in the second text content.
[0154] Optionally, before executing the operation corresponding to the first voice command, the method 200 further includes: determining, based on data collected by sensors outside the cockpit, that a user exists within the first preset safe area.
[0155] For example, Figures 6A-6B illustrate another set of GUIs provided in embodiments of this application.
[0156] As shown in Figure 6A, the vehicle can display the map application interface on the central control screen. When the user's voice command "Xiao A Xiao A, customize safe area" is received, the vehicle can display the GUI shown in Figure 6B on the central control screen.
[0157] As shown in Figure 6B, the vehicle can display a bird's-eye view of its location via the central control screen. Upon detecting a user's finger swiping input on the central control screen, the vehicle can display area 601. When the user clicks the confirmation control, this area 601 can be set as part of a first preset safe area.
[0158] Optionally, if the vehicle previously saved the association between the charging area of the shopping mall and the preset voice command set, then after setting the area 601 as part of the first preset safety area, the association can be updated to the association between the area 601 and the preset voice command set.
[0159] Optionally, if the vehicle has not previously saved the association between the charging area of the shopping mall and the preset voice command set, then after setting the area 601 as part of the first preset safety area, the vehicle can prompt the user to set the preset voice command set corresponding to the area 601.
[0160] For example, the bird's-eye view may be a 3D map of the charging area of the xx shopping mall stored in the map application based on the vehicle's current location, or the bird's-eye view may be a 3D map of the charging area of the xx shopping mall constructed by the vehicle based on the vehicle's current location and data collected by the vehicle's sensors.
[0161] Optionally, before executing the operation corresponding to the first voice command, the method 200 further includes: determining that the voiceprint information of the first voice command matches one of the one or more voiceprint information stored in the vehicle.
[0162] For example, the vehicle may store one or more voiceprint information entries, which may be voiceprint information that the user has registered with the vehicle. Before executing the operation corresponding to the first voice command, the vehicle may first determine that the voiceprint information of the first voice command matches one or more voiceprint information entries stored in the vehicle.
[0163] For example, the vehicle may store the voiceprint information of one or more users. When the vehicle is located within the first preset safe area and the first voice command is an instruction within the first preset voice command set, and the voiceprint information corresponding to the first voice command matches one of the voiceprint information of one or more users, the vehicle can execute the operation corresponding to the first voice command.
[0164] Optionally, the vehicle stores the association between the first preset safe area, the voiceprint information of the first user, and the first set of voice commands. When the vehicle is located within the first preset safe area and the first voice command is a command within the first preset set of voice commands, the operation corresponding to the first voice command is executed, including: when it is determined based on the association that the vehicle is currently located within the first preset safe area, the first voice command is a command within the first preset set of voice commands, and the voiceprint information of the first voice command matches the voiceprint information of the first user, the operation corresponding to the first voice command is executed.
[0165] For example, Figure 7 illustrates another set of GUIs provided in an embodiment of this application.
[0166] As shown in Figure 7A, the vehicle can display the interface for setting the external wake-up function on the central control screen. This interface is used to set the association between frequently used locations, users, and the components corresponding to those locations.
[0167] For example, when the system detects that a user has selected inputs that associate the company, user A, and the air conditioner, doors, windows, trunk, and seats, as well as inputs that associate the company, user B, and the air conditioner and trunk, the association relationships shown in Table 4 can be established.
[0168] Table 4
[0169] The biometric information of the above users may include the user's voiceprint information and / or facial features.
[0170] As shown in Figure 7B, when the vehicle's current location is detected within the company's corresponding safe zone and the user's voice command "Hey A, open the window" is received, and the voiceprint information of this voice command matches the voiceprint information of user B, based on the correlation shown in Table 4, it can be determined that the voice command is not in the corresponding preset voice command set. At this time, the vehicle can determine the source of the voice command; if the voice command originates from outside the cabin, it can be ignored.
[0171] As shown in Figure 7C, when the vehicle's current location is detected to be within the company's corresponding safe area and the user's voice command "Xiao A Xiao A, open the car window" is obtained, and the voiceprint information of the voice command matches the voiceprint information of user A, the operation of opening the car window can be executed based on the correlation shown in Table 4.
[0172] Figure 8 shows a schematic flowchart of a voice interaction method 800 provided in an embodiment of this application. The method 800 includes:
[0173] S801, acquires voice signal.
[0174] For example, the voice signal includes a wake-up voice signal and a command voice signal.
[0175] For example, the wake-up voice signal includes a wake-up word, which can be the name of the voice assistant, such as "Xiao A Xiao A".
[0176] S802, the wake-up voice signal is input into the wake-up module for detection, and the confidence level of the wake-up word in the wake-up voice signal is obtained.
[0177] For example, the wake-up voice signal is input into the wake-up module for detection to obtain the confidence level of the wake-up word in the wake-up voice signal, including: inputting the wake-up voice signal into the wake-up model to obtain the confidence level of the wake-up word. If the confidence level of the wake-up word is greater than or equal to a preset confidence level, it can be determined that the vehicle has obtained the wake-up word used by the user to wake up the voice assistant, and can continue to receive command voice signals; otherwise, S810 can be executed.
[0178] S803, determine whether the confidence level of the wake word is greater than or equal to the preset confidence level.
[0179] If the confidence level of the wake word is greater than or equal to the preset confidence level, then execute S804; otherwise, execute S810.
[0180] S804, acquire command voice signal.
[0181] For example, the voice signal can be used to determine the voice command. For instance, the voice command corresponding to the voice signal can be "open the car window".
[0182] S805 determines whether the vehicle is currently in a preset safe zone.
[0183] For example, a vehicle can determine whether it is currently within a first preset safety zone based on data collected by a positioning sensor and information about a first preset safety zone stored in the vehicle.
[0184] If the vehicle is currently within the preset safe area, execute S806; otherwise, execute S808.
[0185] S806, determine whether the voice command corresponding to the command voice signal is an instruction in the first preset voice instruction set.
[0186] If the voice command corresponding to the command voice signal is an instruction in the first preset voice instruction set, then S807 can be executed; otherwise, S808 can be executed.
[0187] S807 executes the operation corresponding to the voice command.
[0188] For example, the first preset security zone includes the area corresponding to the company, and the first preset semantic command set includes voice commands for controlling the car windows. When it is detected that the vehicle is currently in the area corresponding to the company and the obtained voice command is "open the car window", the vehicle can perform the operation of opening the car window.
[0189] S808 determines the source of the voice command.
[0190] The process of determining the source of the voice signal in S808 above can refer to the process of determining the source of the first voice command in method 200 above, and will not be repeated here.
[0191] If the voice command originates from inside the cockpit, execute S809; otherwise, execute S810.
[0192] S809, execute the operation corresponding to this voice command.
[0193] S810, ignore the voice signal, or the control prompt device prompts the user to refuse to perform the corresponding operation.
[0194] Figure 9 shows a schematic flowchart of a voice interaction method 900 provided in an embodiment of this application. The method 900 includes:
[0195] S910, obtain the source of the first voice command.
[0196] The process of obtaining the source of the first voice command in S910 above can be referred to the description in method 200 above, and will not be repeated here.
[0197] S920, when the first voice command originates from outside the cabin, the vehicle is located within a first preset safety area, and the first voice command is an instruction within a first preset voice command set, the operation corresponding to the first voice command is executed.
[0198] Optionally, the method 900 further includes: when the first voice command originates from inside the cockpit, performing the operation corresponding to the first voice command.
[0199] Optionally, the method 900 further includes: ignoring the first voice command when the first voice command originates from outside the cockpit and the vehicle is not located within the first preset safety area, or prompting the user to refuse to perform the operation corresponding to the first voice command.
[0200] Optionally, the method 900 further includes: ignoring the first voice command when the first voice command originates from outside the cockpit, the vehicle is located within the first preset safety area, and the first voice command is not an instruction within the first preset voice command set; or prompting the user to refuse to execute the operation corresponding to the first voice command.
[0201] Optionally, before executing the operation corresponding to the first voice command, the method 900 further includes: determining that the voiceprint information of the first voice command matches one of the one or more voiceprint information stored in the vehicle.
[0202] Optionally, the first preset safety zone includes one or more sub-regions. When the first voice command originates from outside the cabin, the vehicle is located within the first preset safety zone, and the first voice command is an instruction within a first preset voice command set, the operation corresponding to the first voice command is executed, including: when the first voice command originates from outside the cabin and the vehicle is located within a first sub-region of the first preset safety zone, obtaining the first preset voice command set corresponding to the first sub-region according to a mapping relationship, wherein the mapping relationship includes the association between each sub-region within one or more sub-regions and the preset voice command set corresponding to each sub-region, and the one or more sub-regions include the first sub-region; and when the first voice command is an instruction within a first preset command set, executing the operation corresponding to the first voice command.
[0203] The settings for the preset voice command sets for different sub-regions can be found in the description of method 200 above, and will not be repeated here.
[0204] Optionally, before executing the operation corresponding to the first voice command, the method 900 further includes: controlling the display device to display information about the first preset safe area and multiple components; and in response to the user's input of selecting at least one component from the multiple components, establishing an association between the first preset safe area and the first voice command set, the first voice command set including operation instructions for the at least one component.
[0205] The process of establishing the association between the first preset security area and the first set of voice commands can be illustrated in Figure 4.
[0206] Optionally, before executing the operation corresponding to the first voice command, the method 900 further includes: controlling the prompting device to prompt the user to input a voice command to be executed in the first preset safe area; and in response to one or more voice commands input by the user, establishing an association between the first preset safe area and the first voice command set, wherein the first voice command set includes the one or more voice commands.
[0207] The process of establishing the association between the first preset security area and the first set of voice commands can be illustrated in Figure 5.
[0208] Optionally, before executing the operation corresponding to the first voice command, the method 900 further includes: executing the operation corresponding to the first voice command when the first text content corresponding to the first voice command matches the second text content corresponding to the second voice command among the one or more voice commands.
[0209] Optionally, before executing the operation corresponding to the first voice command, the method 900 further includes: determining, based on data collected by sensors outside the cockpit, that a user exists within the first preset safe area.
[0210] Methods 200, 800, and 900 described above can be executed by the vehicle 100; or, methods 200, 800, and 900 can be executed by the computing platform 110; or, methods 200, 800, and 900 can be executed by the processor, chip, or circuit in the computing platform.
[0211] Figure 10 shows a schematic block diagram of a voice interaction device 1000 provided in an embodiment of this application. The device 1000 includes: an acquisition unit 1010 for acquiring a first voice command; and an instruction execution unit 1020 for executing the operation corresponding to the first voice command when the vehicle is located within a first preset safety area and the first voice command is an instruction within a first preset set of voice commands.
[0212] Optionally, the instruction execution unit is configured to determine the response to the first voice instruction based on the source of the first voice instruction when the vehicle is not located within the first preset safety area.
[0213] Optionally, the instruction execution unit 1020 is configured to determine the response to the first voice instruction based on the source of the first voice instruction when the vehicle is located within the first preset safety area and the first voice instruction is not an instruction within the first preset voice instruction set.
[0214] Optionally, the device further includes a determining unit, which is used to determine, before the instruction execution unit performs the operation corresponding to the first voice instruction, that the voiceprint information of the first voice instruction matches one of the one or more voiceprint information stored in the vehicle.
[0215] Optionally, the first preset safety zone includes one or more sub-regions. The acquisition unit 1010 is further configured to, when the vehicle is located within the first sub-region of the first preset safety zone, acquire the first preset voice command set corresponding to the first sub-region according to the mapping relationship. The mapping relationship includes the association between each sub-region and the preset voice command set corresponding to each sub-region in the one or more sub-regions, and the one or more sub-regions include the first sub-region. The instruction execution unit 1020 is configured to execute the operation corresponding to the first voice command when the first voice command is an instruction within the first preset instruction set.
[0216] Optionally, the device 1000 further includes a control unit and a mapping relationship establishment unit. The control unit is configured to control the display device to display information about the first preset safe area and multiple components before the instruction execution unit executes the operation corresponding to the first voice instruction. The mapping relationship establishment unit is configured to establish an association between the first preset safe area and the first voice instruction set in response to the user's input of selecting at least one component from the multiple components. The first voice instruction set includes operation instructions for the at least one component.
[0217] Optionally, the device 1000 further includes a control unit and a mapping relationship establishment unit. The control unit is used to control the prompting device to prompt the user to input a voice command to be executed in the first preset safe area before the instruction execution unit executes the operation corresponding to the first voice command. The mapping relationship establishment unit is used to establish an association between the first preset safe area and the first voice command set in response to one or more voice commands input by the user. The first voice command set includes the one or more voice commands.
[0218] Optionally, the instruction execution unit 1020 is configured to: execute the operation corresponding to the first voice instruction when the first text content corresponding to the first voice instruction matches the second text content corresponding to the second voice instruction among the one or more voice instructions.
[0219] Optionally, the device 1000 further includes a determining unit, which is used to determine, based on data collected by sensors outside the cockpit, that a user exists within the first preset safe area before the instruction execution unit executes the operation corresponding to the first voice instruction.
[0220] Optionally, the device 1000 further includes: a prediction unit, configured to input the first voice command into a prediction model to obtain the source of the first voice command, wherein the source of the first voice command indicates that the first voice command originates from inside or outside the vehicle's cabin.
[0221] Optionally, the device further includes an energy mapping unit, a filtering unit, and a determining unit. The acquisition unit 1010 is used to acquire a first voice signal corresponding to the first voice command collected by a microphone in the cockpit. The energy mapping unit is used to map the energy of the first voice signal within a preset energy range to obtain a second voice signal. The filtering unit is used to filter the second voice signal to obtain a third voice signal. The determining unit is used to determine the source of the first voice command based on the third voice signal, wherein the source of the first voice command indicates that the first voice command originates from inside or outside the vehicle's cockpit.
[0222] Optionally, the instruction execution unit 1020 is configured to: execute the operation corresponding to the first voice instruction when the first voice instruction originates from inside the cockpit; or, ignore the first voice instruction when the first voice instruction originates from outside the cockpit; or, prompt the user to refuse to execute the operation corresponding to the first voice instruction.
[0223] It should be understood that the division of units in the above device is only a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, the units in the device can be implemented by a processor calling software; for example, the device includes a processor connected to memory, which stores instructions. The processor calls the instructions stored in memory to implement any of the above methods or to implement the functions of each unit in the device. The processor can be, for example, a general-purpose processor, such as a CPU or microprocessor, and the memory can be internal or external to the device. Alternatively, the units in the device can be implemented as hardware circuits. The functions of some or all units can be implemented through the design of the hardware circuits, which can be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all units are implemented through the design of the logical relationships between the components within the circuit. In another implementation, the hardware circuit can be implemented using a PLD, such as an FPGA, which can include a large number of logic gates. The connection relationships between the logic gates are configured through configuration files, thereby implementing the functions of some or all units. All units of the above devices can be implemented entirely through processor calling software, or entirely through hardware circuits, or partially through processor calling software with the remaining parts implemented through hardware circuits.
[0224] In this application embodiment, a processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and execution capabilities, such as a CPU, microprocessor, GPU, or DSP. In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits. These logical relationships are fixed or reconfigurable. For example, the processor may be a hardware circuit implemented as an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and configuring the hardware circuit can be understood as the processor loading instructions to implement the functions of some or all of the above units. Furthermore, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as an NPU, TPU, or DPU.
[0225] As can be seen, each unit in the above device can be one or more processors (or processing circuits) configured to implement the above methods, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.
[0226] Furthermore, the units in the above devices can be integrated in whole or in part, or they can be implemented independently. In one implementation, these units are integrated together as a System-on-a-Chip (SoC). The SoC may include at least one processor for implementing any of the above methods or implementing the functions of the units in the device. The at least one processor may be of different types, such as CPU and FPGA, CPU and AI processor, CPU and GPU, etc.
[0227] This application also provides a voice interaction device, which includes a processing unit and a storage unit. The storage unit is used to store instructions, and the processing unit executes the instructions stored in the storage unit to cause the device to perform the methods or steps described in the above embodiments.
[0228] Optionally, if the voice interaction device is located in a vehicle, the processing unit may be the processor 111-11n shown in FIG1.
[0229] This application also provides a voice interaction system, which may include a computing platform and a microphone. The computing platform may include the aforementioned voice interaction device 1000.
[0230] This application also provides a vehicle that may include the aforementioned voice interaction device 1000 or voice interaction system.
[0231] This application also provides a computer program product, which includes computer program code that, when run on a computer, causes the computer to perform the methods described in the above embodiments.
[0232] This application also provides a computer-readable medium storing program code that, when run on a computer, causes the computer to perform the methods described in the above embodiments.
[0233] This application also provides a chip, which includes a circuit for performing the methods described in the above embodiments.
[0234] In implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software. The method disclosed in the embodiments of this application can be directly implemented by a hardware processor, or by a combination of hardware and software modules within the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, power-on erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, detailed descriptions are omitted here.
[0235] It should be understood that in the embodiments of this application, the memory may include read-only memory and random access memory, and provides instructions and data to the processor.
[0236] It should also be understood that, in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0237] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0238] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0239] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0240] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0241] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0242] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0243] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be covered. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A voice interaction method, characterized in that, include: Obtain the first voice command; When the vehicle is located within a first preset safety area and the first voice command is a command within a first preset voice command set, the operation corresponding to the first voice command is executed.
2. The method according to claim 1, characterized in that, The method further includes: When the vehicle is not located within the first preset safety area, the response to the first voice command is determined based on the source of the first voice command.
3. The method according to claim 1, characterized in that, The method further includes: When the vehicle is located within the first preset safety area and the first voice command is not a command within the first preset voice command set, the response to the first voice command is determined based on the source of the first voice command.
4. The method according to any one of claims 1 to 3, characterized in that, Before executing the operation corresponding to the first voice command, the method further includes: The voiceprint information of the first voice command is determined to match one of the one or more voiceprint information stored in the vehicle.
5. The method according to any one of claims 1 to 4, characterized in that, The first preset safety zone includes one or more sub-regions. When the vehicle is located within the first preset safety zone and the first voice command is an instruction within the first preset voice command set, executing the operation corresponding to the first voice command includes: When the vehicle is located within a first sub-region of the first preset safety area, the first preset voice command set corresponding to the first sub-region is obtained according to the mapping relationship. The mapping relationship includes the association between each sub-region in the one or more sub-regions and the preset voice command set corresponding to each sub-region. The one or more sub-regions include the first sub-region. When the first voice command is an instruction within the first preset instruction set, the operation corresponding to the first voice command is executed.
6. The method according to any one of claims 1 to 5, characterized in that, When the vehicle is located within a first preset safety area and the first voice command is a command within a first preset voice command set, before executing the operation corresponding to the first voice command, the method further includes: The control display device displays information about the first preset safety area and multiple components; In response to a user's input to select at least one component from the plurality of components, an association is established between the first preset security area and the first voice command set, the first voice command set including operation instructions for the at least one component.
7. The method according to any one of claims 1 to 5, characterized in that, When the vehicle is located within a first preset safety area and the first voice command is an instruction within a first preset instruction set, before executing the operation corresponding to the first voice command, the method further includes: The control prompting device prompts the user to input a voice command to be executed in the first preset safe area; In response to one or more voice commands input by the user, an association is established between the first preset security area and the first set of voice commands, wherein the first set of voice commands includes the one or more voice commands.
8. The method according to claim 7, characterized in that, Before executing the operation corresponding to the first voice command, the method further includes: When the first text content corresponding to the first voice command matches the second text content corresponding to the second voice command among the one or more voice commands, the operation corresponding to the first voice command is executed.
9. The method according to any one of claims 1 to 8, characterized in that, Before executing the operation corresponding to the first voice command, the method further includes: Based on data collected by sensors outside the cockpit, it is determined that a user exists within the first preset safety area.
10. The method according to claim 2 or 3, characterized in that, Before determining the response to the first voice command based on its source, the method further includes: The first voice command is input into the prediction model to obtain the source of the first voice command. The source of the first voice command indicates that the first voice command originates from inside or outside the vehicle's cabin.
11. The method according to claim 2 or 3, characterized in that, Before determining the response to the first voice command based on its source, the method further includes: Acquire the first voice signal corresponding to the first voice command collected by the microphone in the cockpit; The energy of the first speech signal is mapped within a preset energy range to obtain the second speech signal; The second speech signal is filtered to obtain the third speech signal; Based on the third voice signal, the source of the first voice command is determined, and the source of the first voice command indicates that the first voice command originates from inside or outside the vehicle's cabin.
12. The method according to claim 10 or 11, characterized in that, Determining the response to the first voice command based on its source includes: When the first voice command originates from inside the cockpit, execute the operation corresponding to the first voice command; or... When the first voice command originates from outside the cockpit, the first voice command is ignored, or the user is prompted to refuse to execute the operation corresponding to the first voice command.
13. A voice interaction device, characterized in that, include: Acquisition unit, used to acquire the first voice command; The instruction execution unit is used to execute the operation corresponding to the first voice instruction when the vehicle is located within a first preset safety area and the first voice instruction is an instruction within a first preset set of voice instructions.
14. The apparatus according to claim 13, characterized in that, The instruction execution unit is configured to determine the response to the first voice instruction based on the source of the first voice instruction when the vehicle is not located within the first preset safety area.
15. The apparatus according to claim 13, characterized in that, The instruction execution unit is configured to determine the response to the first voice instruction based on the source of the first voice instruction when the vehicle is located within the first preset safety area and the first voice instruction is not an instruction within the first preset voice instruction set.
16. The apparatus according to any one of claims 13 to 15, characterized in that, The device also includes a determining unit. The determining unit is configured to determine, before the instruction execution unit executes the operation corresponding to the first voice instruction, that the voiceprint information of the first voice instruction matches one of the one or more voiceprint information stored in the vehicle.
17. The apparatus according to any one of claims 13 to 16, characterized in that, The first preset security area includes one or more sub-areas. The acquisition unit is further configured to, when the location of the vehicle is within a first sub-region of the first preset safety area, acquire the first preset voice command set corresponding to the first sub-region according to the mapping relationship, wherein the mapping relationship includes the association between each sub-region in the one or more sub-regions and the preset voice command set corresponding to each sub-region, and the one or more sub-regions include the first sub-region; The instruction execution unit is used to execute the operation corresponding to the first voice instruction when the first voice instruction is an instruction within a first preset instruction set.
18. The apparatus according to any one of claims 13 to 17, characterized in that, The device also includes a control unit and a mapping relationship establishment unit. The control unit is configured to control the display device to display information about the first preset safety area and multiple components before the instruction execution unit executes the operation corresponding to the first voice instruction; The mapping relationship establishment unit is used to establish an association between the first preset security area and the first voice command set in response to the user's input of selecting at least one component from the plurality of components. The first voice command set includes operation instructions for the at least one component.
19. The apparatus according to any one of claims 13 to 17, characterized in that, The device also includes a control unit and a mapping relationship establishment unit. The control unit is configured to control the prompting device to prompt the user to input a voice command to be executed in the first preset safe area before the instruction execution unit executes the operation corresponding to the first voice command. The mapping relationship establishment unit is used to establish an association between the first preset security area and the first voice command set in response to one or more voice commands input by the user, wherein the first voice command set includes the one or more voice commands.
20. The apparatus according to claim 19, characterized in that, The instruction execution unit is used for: When the first text content corresponding to the first voice command matches the second text content corresponding to the second voice command among the one or more voice commands, the operation corresponding to the first voice command is executed.
21. The apparatus according to any one of claims 13 to 20, characterized in that, The device also includes a determining unit. The determining unit is used to determine, based on data collected by sensors outside the cockpit, that a user exists within the first preset safe area before the instruction execution unit executes the operation corresponding to the first voice instruction.
22. The apparatus according to claim 14 or 15, characterized in that, The device further includes: The prediction unit is used to input the first voice command into the prediction model to obtain the source of the first voice command, wherein the source of the first voice command indicates that the first voice command originates from inside or outside the vehicle cabin.
23. The apparatus according to claim 14 or 15, characterized in that, The device also includes an energy mapping unit, a filtering unit, and a determination unit. The acquisition unit is used to acquire the first voice signal corresponding to the first voice command collected by the microphone in the cockpit; The energy mapping unit is used to map the energy of the first speech signal within a preset energy range to obtain the second speech signal; The filtering unit is used to filter the second speech signal to obtain the third speech signal; The determining unit is configured to determine the source of the first voice command based on the third voice signal, wherein the source of the first voice command indicates that the first voice command originates from inside or outside the vehicle's cabin.
24. The apparatus according to claim 22 or 23, characterized in that, The instruction execution unit is used for: When the first voice command originates from inside the cockpit, execute the operation corresponding to the first voice command; or... When the first voice command originates from outside the cockpit, the first voice command is ignored, or the user is prompted to refuse to execute the operation corresponding to the first voice command.
25. A voice interaction device, characterized in that, include: Memory, used to store computer programs; A processor for executing a computer program stored in the memory to cause the apparatus to perform the method as described in any one of claims 1 to 12.
26. A voice interaction system, characterized in that, It includes a microphone and a computing platform, the computing platform including the means as described in any one of claims 13 to 25.
27. A vehicle, characterized in that, Includes the apparatus as described in any one of claims 13 to 25, or the system as described in claim 26.
28. A computer-readable storage medium, characterized in that, It stores instructions that, when executed by a processor, cause the processor to implement the method as described in any one of claims 1 to 12.
29. A computer program product, characterized in that, The computer program product includes computer program code that, when run on a computer, causes the computer to perform the method as described in any one of claims 1 to 12.
30. A chip, characterized in that, The chip includes circuitry for performing the method as described in any one of claims 1 to 12.