Voice interaction-based instruction recognition method, storage medium and electronic device
By recognizing the behavior of people inside the vehicle and extracting passenger voice features, the problem of inaccurate passenger voice recognition in the vehicle has been solved, thus improving the accuracy of vehicle system operation and driving safety.
Patent Information
- Application Number
- CN202110625185.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-04
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2041-06-04
AI Technical Summary
Existing technology cannot accurately recognize the voice commands of passengers in the vehicle, leading to erroneous triggering of the vehicle's infotainment system and affecting driving safety.
By recognizing the behavioral information of people inside the vehicle, they are divided into voice interaction objects and non-voice interaction objects. The voice features of non-voice interaction objects are extracted and their voice features are masked in the in-vehicle voice information to generate voice commands to control the vehicle.
It effectively prevents passengers from triggering inappropriate vehicle functions with their voice, improving driving safety, especially by blocking passenger voice during entertainment applications, thus ensuring the accuracy of driver voice control.
Smart Images

Figure CN115440204B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of voice control, and relates to a voice instruction recognition method, in particular to a voice interaction-based instruction recognition method, a storage medium and an electronic device. BACKGROUND
[0002] At present, the intelligent degree of automobiles is getting higher and higher, which brings a lot of convenience to people. Liquid crystal panels are used as display devices, and their application on automobiles is becoming more and more popular. On some vehicle models, there are not only the instrument panel of the driver, but also the center control panel and several liquid crystal panels used by passengers on corresponding seats. In addition to the instrument panel of the driver, the other panels also have touch functions, which can control the playing of videos and audios, and can also be used for users to play games and other entertainment operations.
[0003] However, sometimes the passengers in the car will make various sounds when using entertainment application software, for example, the passengers other than the driver will usually make various voices when playing games in the car. In particular, for racing and racing games, the players will shout due to being too involved, and when the vehicle-mounted system cannot recognize the voice of the specific passenger, it will trigger inappropriate voice interaction, which may adversely affect driving safety.
[0004] Therefore, how to provide a voice interaction-based instruction recognition method, a storage medium and an electronic device to solve the problem that the prior art cannot accurately recognize vehicle machine control instructions from many voices in the car, and to prevent vehicle machine operation from being triggered by mistake, has become a technical problem to be solved by those skilled in the art. SUMMARY
[0005] In view of the above-mentioned shortcomings of the prior art, the purpose of the present application is to provide a voice interaction-based instruction recognition method, a storage medium and an electronic device, which has the advantage that it can accurately recognize vehicle machine control instructions from many voices in the car, and thus prevent vehicle machine operation from being triggered by mistake.
[0006] Another purpose of the present application is to provide a voice interaction-based instruction recognition method, a storage medium and an electronic device, which has the advantage that it can recognize the voice of a specific passenger in the car, effectively avoiding the situation that the sound emitted by some passengers in the car triggers inappropriate vehicle machine voice interaction functions, which in turn adversely affects driving safety.
[0007] Another purpose of the present application is to provide a voice interaction-based instruction recognition method, a storage medium and an electronic device, which has the advantage that it can shield the voice of a user when the user uses entertainment applications such as games and KTV, and restore the function of triggering vehicle machine control by the voice of the user when the user exits the entertainment applications such as games and KTV.
[0008] Another object of the present application is to provide a voice interaction-based instruction recognition method, storage medium and electronic device, which has the advantage that the driver can actively trigger the vehicle to record the voice features of a specific user and set the time period for which the voice of the user is shielded.
[0009] Another object of the present application is to provide a voice interaction-based instruction recognition method, storage medium and electronic device, which has the advantage that the behavior of a passenger in the vehicle can be identified through a video image to shield the voice of the passenger when the passenger is performing an entertainment operation.
[0010] Another object of the present application is to provide a voice interaction-based instruction recognition method, storage medium and electronic device, which has the advantage that the voice features of a passenger in the vehicle can be extracted from a video image, and then the voice of the passenger is determined to be shielded or not in combination with the behavior of the passenger in the vehicle.
[0011] To achieve the above object and other related objects, one aspect of the present application provides a voice interaction-based instruction recognition method, which comprises: dividing a person in a vehicle into a voice interaction object and a non-voice interaction object according to the behavior information of the person in the vehicle; identifying the voice features of the non-voice interaction object; shielding the voice features of the non-voice interaction object in the acquired voice information in the vehicle to generate a voice instruction; and using the voice instruction to perform voice control on the vehicle.
[0012] To achieve the above object and other related objects, another aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement the steps of the voice interaction-based instruction recognition method.
[0013] To achieve the above object and other related objects, the last aspect of the present application provides an electronic device, which comprises: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to make the electronic device execute the steps of the voice interaction-based instruction recognition method. BRIEF DESCRIPTION OF DRAWINGS
[0014] Figure 1 A principle flowchart of the voice interaction-based instruction recognition method of the present application in an embodiment is shown.
[0015] Figure 2 An object division flowchart of the voice interaction-based instruction recognition method of the present application in an embodiment is shown.
[0016] Figure 3 A voice feature identification flowchart of the voice interaction-based instruction recognition method of the present application in an embodiment is shown.
[0017] Figure 4 A voice instruction generation flowchart shown in an embodiment of the voice interaction-based instruction recognition method of the present application.
[0018] Figure 5 A permission setting flowchart shown in an embodiment of the voice interaction-based instruction recognition method of the present application.
[0019] Figure 6 An entertainment APP voice shielding flowchart shown in an embodiment of the voice interaction-based instruction recognition method of the present application.
[0020] Figure 7 An entertainment APP voice shielding flowchart shown in another embodiment of the voice interaction-based instruction recognition method of the present application.
[0021] Figure 8 A structural connection schematic diagram shown in an embodiment of the electronic device of the present application.
[0022] Figure 9 A functional module interaction schematic diagram shown in an embodiment of the electronic device of the present application.
[0023] Figure 10 A functional module interaction schematic diagram shown in another embodiment of the electronic device of the present application.
[0024] Element number explanation
[0025] 1 electronic device
[0026] 11 processor
[0027] 12 memory
[0028] S11-S14 steps
[0029] S111-S113 steps
[0030] S121-S122 steps
[0031] S131-S134 steps
[0032] S1A-S1B steps DETAILED DESCRIPTION
[0033] Following make the specific concrete example explain the implementation of the present application, the person skilled in the art can be easily understood from the present application disclosed in the content of the advantages and efficacy of the present application.It can also be implemented or applied by another different specific implementation, the details in the specification can also be based on different views and applications, without departing from the spirit of the present application, various modifications or changes are made.The need to explain that the following examples and the features in the examples can be combined with each other without conflict.
[0034] Need to explain that the following examples provided in the figure only illustrates the basic concept of the present application, the figure shows only the components related to the present application and not drawn according to the actual implementation of the number of components, shape and size, the actual implementation of each component type, quantity and proportion can be a kind of arbitrary change, and its component layout type may be more complex.
[0035] The voice interaction based instruction recognition method, storage medium and electronic device described in the present application can recognize the voice of a specific passenger in the vehicle, effectively avoiding the sound emitted by some passengers in the vehicle triggering inappropriate car audio voice interaction functions, thereby adversely affecting the driving safety.
[0036] The following will be combined Figures 1 to 10 The principle and implementation of a voice interaction based instruction recognition method, storage medium and electronic device of the present embodiment will be described in detail, so that those skilled in the art can understand the voice interaction based instruction recognition method, storage medium and electronic device of the present embodiment without creative labor.
[0037] Please refer to Figure 1 , which shows the principle flowchart of the voice interaction based instruction recognition method of the present application in an embodiment.As Figure 1 shown, the voice interaction based instruction recognition method specifically includes the following steps:
[0038] S11, according to the behavior information of the vehicle personnel, the vehicle personnel is divided into voice interaction object and non voice interaction object.
[0039] Specifically, the voice interaction object refers to the personnel with voice control car demand, and the non voice interaction object refers to the personnel without voice control car demand.The division of vehicle personnel is based on whether the sound is used to control the car, for example, the driver controls the car through voice instruction during driving, then the driver is the voice interaction object, the rear passengers are playing games or karaoke through the rear display screen, and the game shouting and singing sound is not used to control the car, then the rear passengers are non voice interaction object.
[0040] Please refer to Figure 2, shows the object division flowchart in an embodiment of the voice interaction based instruction recognition method of the present application. As shown in Figure 2 S11 specifically includes the following steps:
[0041] S111, acquiring the type information of the application operated by the in-vehicle person.
[0042] Specifically, the type information of the operated application can be operation information of an entertainment APP (Application, short for application) or other operation information irrelevant to voice control in the vehicle driving process.
[0043] S112, judging whether the type of the application operated by the in-vehicle person is an entertainment application according to the type information of the operated application.
[0044] In an embodiment, the entertainment application includes at least one of a game application or a K song application.
[0045] Specifically, the entertainment application refers to a game APP, a K song APP, a child learning APP and other leisure and entertainment related applications used by other passengers. Control operations used in the driving process, such as voice control sunroof and window, voice control air conditioner internal and external circulation, voice control navigation operation and other vehicle control related applications do not belong to the entertainment application.
[0046] Further, the judgment principle of whether it is a game APP or a K song APP is that: when each APP runs, the application information maintained by the application store (such as APP Store) for the APP is read, the application information includes type information, and the type information is used to determine whether it is an entertainment APP such as a game APP or a K song APP.
[0047] S113, if the type of the application is not an entertainment application, the in-vehicle person is determined as a voice interaction object; if the type of the application is an entertainment application, the in-vehicle person is determined as a non-voice interaction object. Thus, the in-vehicle person is divided by judging whether the application on the vehicle display screen is an entertainment application.
[0048] In another embodiment, S11 includes the following steps:
[0049] (1) Acquiring the in-vehicle image of the in-vehicle person.
[0050] Specifically, the in-vehicle image can be an image of the passenger operating on a different display screen installed on the vehicle, or an image of the passenger operating an application by using a handheld terminal such as a mobile phone or a tablet computer in the vehicle. Thus, in the present embodiment, the division of the voice interaction object and the non-voice interaction object breaks through the detection of the vehicle display screen APP and can identify any non-driving application operation in the in-vehicle space.
[0051] (2) determining whether the in-vehicle person is at the driving seat according to the in-vehicle image.
[0052] Specifically, the in-vehicle image can be an image of the passenger operating on a different display screen installed on the vehicle, or an image of the passenger operating an application by using a handheld terminal such as a mobile phone or a tablet computer in the vehicle. Thus, in the present embodiment, the division of the voice interaction object and the non-voice interaction object breaks through the detection of the vehicle display screen APP and can identify any non-driving application operation in the in-vehicle space.
[0053] (3) If the in-vehicle person is at the driving seat, the in-vehicle person is determined as the voice interaction object; if the in-vehicle person is not at the driving seat, the in-vehicle person is determined as the non-voice interaction object. Thus, the driver at the driving seat is determined as the voice interaction object, and the passengers at other seats are determined as the non-voice interaction object.
[0054] Further, the judgment of the driving seat can be combined with the passenger behavior information to make a judgment. For example, it is determined by image recognition that there is a passenger at the left rear seat, the passenger is a primary school student, and the passenger is practicing reading an article in a textbook by using a tablet computer. It is determined by the behavior of the primary school student that the primary school student is the non-voice interaction object.
[0055] Further, the voice feature of the non-voice interaction object is extracted from the in-vehicle image. For example, the voice feature of the primary school student is extracted from the in-vehicle image. The in-vehicle image is a video in the vehicle, the audio in the in-vehicle video is extracted, the in-vehicle video image and the audio are compared according to the same time axis, if it is identified that the primary school student is reading an article by opening his mouth in the in-vehicle video in a certain time period, the audio in the time period is intercepted, and thus the audio of the primary school student is obtained, that is, the voice feature of the primary school student is obtained.
[0056] S12, identifying the voice feature of the non-voice interaction object.
[0057] Specifically, the sound of the input non-voice interaction object is recorded as a kind of machine-identifiable signal, and the extractable voice feature includes indicators such as sound intensity, loudness, pitch, gene period, gene frequency, and signal-to-noise ratio, which can clearly identify the sound characteristics of different persons.
[0058] Please refer to Figure 3 , which shows a voice feature identification flowchart of the voice interaction-based instruction recognition method in an embodiment of the present application. As shown in Figure 3 , S12 specifically includes the following steps:
[0059] S121, presenting a prompt interface of the non-voice interactive object; the prompt interface is used to prompt the non-voice interactive object to input an audio file.
[0060] In an embodiment, the type of the application is an entertainment application, such as a game application or a karaoke application. A sound input prompt dialog box is displayed on a game operation interface or a karaoke application interface; the prompt dialog box is used to guide an operator to input the audio file before starting the application.
[0061] S122, performing audio processing on the audio file to generate a voice feature of the non-voice interactive object.
[0062] Specifically, MFCC (Mel Frequency Cepstrum Coefficient) or FBank features are extracted from the audio signal. The MFCC feature and the FBank feature are two features widely used in existing speech recognition technologies. Fbank is a front-end processing algorithm that processes audio in a manner similar to the human ear, which can improve the performance of speech recognition.
[0063] S13, shielding the voice feature of the non-voice interactive object in the obtained in-vehicle voice information to generate a voice instruction.
[0064] Specifically, before generating a voice instruction that can control the car machine, the in-vehicle voice is recognized and shielded, for example, when the voice interactive object and the non-voice interactive object jointly emit sound, the voice of the non-voice interactive object is shielded according to the voice feature of the non-voice interactive object, that is, even if the voice of the non-voice interactive object has the same voice content as the voice instruction for controlling the car machine, it is considered invalid, for example, during the shielding process, the non-voice interactive object emits “open window or accelerate” and the like, which are the same voice content as the voice instruction for controlling the car machine, and is considered invalid. When the non-voice interactive object emits sound alone, the voice of the non-voice interactive object is ignored.
[0065] Please refer to Figure 4 , which shows a voice instruction generation flowchart of the voice interaction-based instruction recognition method in an embodiment of the present application. As shown in Figure 4 , S13 includes the following steps:
[0066] S131, determining a preset time period according to the voice interaction prohibition time of the non-voice interactive object.
[0067] Specifically, the driver is in the driving process from 10 o'clock to 11 o'clock in the morning, and the preset time period can be set to 10 o'clock to 11 o'clock in the morning.
[0068] S132, determine the generation mode of the voice instruction according to the preset time period.
[0069] Specifically, between 10am and 11am, the generation mode of the voice instruction is set, specifically, between 10am and 11am, the voice instruction is generated according to the generation mode described in step S133; outside 10am and 11am, the voice instruction is generated according to the generation mode described in step S134.
[0070] S133, within the preset time period, screen the voice features of the non-voice interaction object in the obtained in-vehicle voice information, and generate the voice instruction.
[0071] Specifically, between 10am and 11am, the generated voice instruction only includes the voice generated instruction of the driver, or includes the voice of the driver and the voice of the passenger in the vehicle who does not make sound for entertainment application.
[0072] S134, outside the preset time period, directly generate the voice instruction according to the obtained in-vehicle voice information.
[0073] Specifically, outside 10am and 11am, if the driving trip is over, the voices of all passengers in the vehicle can participate in the generation of the voice instruction.
[0074] S14, use the voice instruction to perform voice control on the vehicle.
[0075] Please refer to Figure 5 , which shows the permission setting flowchart of the voice interaction-based instruction recognition method according to an embodiment of the present application. As Figure 5 shown, in an embodiment, before step S11, the voice interaction-based instruction recognition method further includes the following steps:
[0076] S1A, obtain the interaction permission instruction of the voice interaction object.
[0077] Specifically, the driver sets the generation of the voice instruction through the center control display screen, on the one hand, sets the in-vehicle personnel whose voice needs to be screened, and on the other hand, sets the voice instruction effective and invalid time of the in-vehicle personnel whose voice needs to be screened.
[0078] S1B, determine the non-voice interaction object and the voice interaction prohibition time of the non-voice interaction object according to the interaction permission instruction.
[0079] Specifically, a personnel list is presented on the center control display screen used by the driver, the driver selects the personnel currently in the vehicle by browsing the personnel information, and after the vehicle machine receives the selection instruction, the voice features of the selected in-vehicle personnel are screened.
[0080] Specifically, a control button is presented on the center display screen used by the driver. After the driver clicks the button, the button is highlighted, indicating that the voice shielding function of other passengers is activated. The driver clicks the button again, and the button becomes gray, indicating that the voice shielding function of other passengers is deactivated. Alternatively, a time period setting control is presented on the center display screen used by the driver. The driver can accurately set the start time and end time of the shielding function through the setting control.
[0081] Referring to Figure 6 , a flowchart of an entertainment APP voice shielding process in an embodiment of the voice interaction-based instruction recognition method of the present application is shown. As Figure 6 indicated, the principle of the voice interaction-based instruction recognition method is described in detail taking the use of an entertainment APP by a passenger in the vehicle as an example. After starting execution, the co-driver passenger clicks the entertainment APP on the co-driver display screen. The vehicle machine intercepts the APP activation instruction and determines whether the co-driver passenger clicks the entertainment APP. If not, the determination continues to be monitored. If yes, the voice recognition function is activated, a dialog box is popped up, and the co-driver passenger is guided to input a specific voice segment. The voice characteristics of the co-driver passenger are recorded so that the voice characteristics of the co-driver passenger can be recognized in the process of voice control of the vehicle machine. After the voice characteristics are successfully extracted, the co-driver display screen displays an entertainment interface for the co-driver passenger to use the entertainment APP. Before the entertainment APP ends, the vehicle machine shields the voice interaction function of the co-driver passenger. After the entertainment APP ends, the vehicle machine restores the voice interaction function of the co-driver passenger.
[0082] Referring to Figure 7 , a flowchart of an entertainment APP voice shielding process in another embodiment of the voice interaction-based instruction recognition method of the present application is shown. As Figure 7 indicated, in another embodiment, the driver discovers which seat the passenger is using the entertainment application during the driving process, triggers the display screen corresponding to the passenger, and specifically controls the voice interaction of the specific passenger. The driver operates the vehicle machine and triggers the display screen in front of the specific passenger through voice or the main control screen to activate the audio recording function. The specific passenger on the other seat performs audio recording according to the prompt of the display screen. The vehicle machine analyzes the audio and records the voice characteristics of the audio. The vehicle machine shields the voice interaction of the specific passenger during the driving process. After the specific passenger ends the entertainment application, the driver deactivates the voice interaction function of the specific passenger through voice or the main control screen.
[0083] Thus, Figure 6 and Figure 7 in the embodiments, the execution of the voice interaction-based instruction recognition method can avoid unnecessary safety risks to vehicle driving caused by the shouting of a specific user during the driving process due to playing games or karaoke.
[0084] The protection scope of the voice interaction based instruction recognition method described in the present application is not limited to the step execution order listed in the present embodiment, and any scheme realized by adding, reducing or replacing steps according to the principle of the present application is included in the protection scope of the present application.
[0085] The present embodiment provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the voice interaction based instruction recognition method.
[0086] Those skilled in the art can understand that all or part of the steps of the above-mentioned method embodiments can be completed by computer program related hardware. The aforementioned computer program can be stored in a computer readable storage medium. The program is executed to perform the steps of the above-mentioned method embodiments; and the aforementioned computer readable storage medium includes ROM, RAM, magnetic disc or optical disc and various computer storage media that can store program codes.
[0087] Please refer to Figure 8 , which shows the structure connection schematic diagram of the electronic device in an embodiment of the present application. As shown in Figure 8 , the present embodiment provides an electronic device 1, which specifically includes a processor 11 and a memory 12; the memory 12 is used to store a computer program, and the processor 11 is used to execute the computer program stored in the memory 12, so that the electronic device 1 performs each step of the voice interaction based instruction recognition method.
[0088] The aforementioned processor 11 can be a general processor, including a central processing unit (CPU), a network processor (NP) and the like; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.
[0089] The aforementioned memory 12 can include a random access memory (RAM), and can also include a non-volatile memory, for example, at least one disk memory.
[0090] In practical applications, the electronic device can be a computer including a memory, a storage controller, one or more processing units (CPU), a peripheral interface, RF circuitry, audio circuitry, a speaker, a microphone, an input / output (I / O) subsystem, a display screen, other output or control devices, and an external port, etc. The computer includes, but is not limited to, a personal computer such as a desktop computer, a notebook computer, a tablet computer, a smart phone, a smart television, a personal digital assistant (PDA), etc. The electronic device can also be a car terminal or smart glasses, a smart watch, or other wearable devices. In other embodiments, the electronic device can also be a server, which can be arranged on one or more physical servers according to functions, loads, and other factors, or can be a cloud server composed of distributed or centralized server clusters, and the present embodiment is not limited thereto.
[0091] Further, the electronic device can be a car terminal, and each step of the instruction recognition method based on voice interaction can be performed by the car terminal. The electronic device can also be a mobile phone, a computer, or the like other than the car terminal, and the mobile phone, the computer, or the like other than the car terminal is in communication connection with the car terminal. The behavior information of the person in the vehicle is acquired by a detection device such as a camera in wireless communication, and each step of the instruction recognition method based on voice interaction is performed by the mobile phone, the computer, or the like. After the instruction recognition method based on voice interaction is performed, the voice instruction is generated and sent to the car terminal, so as to realize voice control of the vehicle by using the voice instruction.
[0092] Referring to Figure 9 , a functional module interaction diagram of the electronic device of the present application in an embodiment is shown. As Figure 9 indicated, in combination with the example shown in Figure 6 , a functional module interaction process in which the car terminal automatically identifies an entertainment APP is presented. The car terminal is a T-BOX (Telematics BOX), and the audio recording module can be a module provided by the car terminal or a module provided by an external device having an audio recording function and in communication with the car terminal.
[0093] Specifically, the car terminal monitors that the APP to be opened is an entertainment APP, triggers an audio recording interface on the screen of the entertainment APP, the audio recording module displays a voice recording interface, completes voice recording, saves the voice in the car terminal, extracts the audio features of the voice, and shields the voice interaction function of the passenger. The car terminal monitors that the entertainment APP is exited, and restores the voice interaction function of the passenger.
[0094] Referring to Figure 10, shows the functional module interaction schematic diagram of the electronic device in another embodiment of the application. As shown in Figure 10 Figure 7 The functional module interaction process identified by the driver-led entertainment APP is presented in the example shown in Figure 7 The car machine interacts with the audio recording module, where the car machine is a vehicle-mounted T-BOX (short for Telematics BOX), and the audio recording module can be a module provided by the car machine or an external module with audio recording function that can communicate with the car machine.
[0095] Specifically, the driver issues an instruction by pressing a button, making a gesture, or speaking, controls the car machine to start the voice feature collection function, and then the audio recording module displays a voice recording interface on the display screen of the entertainment APP, completes voice recording, saves the voice in the car machine, extracts the audio features of the voice, and shields the voice interaction function of the passenger. The voice interaction function of the passenger is restored when the passenger exits the entertainment APP.
[0096] In summary, the instruction recognition method based on voice interaction, storage medium, and electronic device of the application can accurately identify car machine control instructions from many voices in the vehicle, thereby preventing the car machine from being mistakenly triggered. The voice of a specific passenger in the vehicle can be identified, effectively avoiding the voice of a passenger in the vehicle triggering inappropriate car machine voice interaction functions, thereby adversely affecting driving safety. The voice of a user can be shielded when the user uses entertainment applications such as games and karaoke, and the voice of the user can be restored when the user exits the entertainment applications such as games and karaoke. The driver can actively trigger the car machine to record the voice features of a specific user, and can also set the time period for shielding the voice of the user. The behavior of a passenger in the vehicle can be identified through video images to shield the voice of the passenger when the passenger is engaged in entertainment operations. The voice features of a passenger in the vehicle can be extracted from video images, and then combined with the behavior of the passenger in the vehicle to determine whether to shield the voice of the passenger. The application effectively overcomes the shortcomings of the prior art and has high industrial utilization value.
[0097] The above embodiments only exemplarily illustrate the principles and effects of the application, and are not intended to limit the application. Any person skilled in the art can modify or change the above embodiments without departing from the spirit and scope of the application. Therefore, all equivalent modifications or changes made by those skilled in the art without departing from the spirit and technical ideas disclosed by the application should be covered by the claims of the application.
Claims
1.A voice interaction based instruction recognition method, characterized in that, the voice interaction based instruction recognition method comprises: dividing in-car personnel into a voice interaction object and a non-voice interaction object according to behavior information of the in-car personnel; recognizing voice features of the non-voice interaction object; shielding the voice features of the non-voice interaction object in acquired in-car voice information to generate a voice instruction; controlling a vehicle by using the voice instruction; the step of dividing the in-car personnel into the voice interaction object and the non-voice interaction object according to the behavior information of the in-car personnel comprises: acquiring type information of an application operated by the in-car personnel; judging whether the type of the application operated by the in-car personnel is an entertainment application according to the type information of the operated application; if the type of the application is not the entertainment application, determining the in-car personnel as the voice interaction object; if the type of the application is the entertainment application, determining the in-car personnel as the non-voice interaction object. 2.The voice interaction based instruction recognition method according to claim 1, wherein the entertainment application comprises at least one of a game application or a karaoke application. 3.The voice interaction based instruction recognition method according to claim 2, wherein the step of recognizing the voice features of the non-voice interaction object comprises: presenting a prompt interface of the non-voice interaction object; the prompt interface is used to prompt the non-voice interaction object to input an audio file; performing audio processing on the audio file to generate the voice features of the non-voice interaction object. 4.The voice interaction based instruction recognition method according to claim 3, wherein the step of presenting the prompt interface of the non-voice interaction object comprises: displaying a sound input prompt dialog box on a game operation interface or a karaoke application interface; the prompt dialog box is used to guide an operator to input the audio file before starting the application. 5.The voice interaction based instruction recognition method according to claim 1, wherein the step of dividing the in-car personnel into the voice interaction object and the non-voice interaction object according to the behavior information of the in-car personnel comprises: acquiring an in-car image of the in-car personnel; judging whether the in-car personnel is at a driving seat according to the in-car image; if the in-car personnel is at the driving seat, determining the in-car personnel as the voice interaction object; if the in-car personnel is not at the driving seat, determining the in-car personnel as the non-voice interaction object. 6.The voice interaction based instruction recognition method according to claim 5, wherein the step of recognizing the voice features of the non-voice interaction object comprises: extracting the voice features of the non-voice interaction object from the in-car image. 7.The voice interaction based instruction recognition method according to claim 1, wherein before the step of dividing the in-car personnel into the voice interaction object and the non-voice interaction object according to the behavior information of the in-car personnel, the voice interaction based instruction recognition method further comprises: acquiring an interaction permission instruction of the voice interaction object. determine the non-voice interactive object and a voice interaction prohibition time of the non-voice interactive object according to the interactive permission instruction. 8.The voice interaction based instruction recognition method of claim 7, wherein the step of screening the voice feature of the non-voice interactive object from the acquired in-vehicle voice information to generate a voice instruction comprises the following steps: determining a preset time period according to the voice interaction prohibition time of the non-voice interactive object; determining a generation mode of the voice instruction in combination with the preset time period; screening the voice feature of the non-voice interactive object from the acquired in-vehicle voice information to generate the voice instruction within the preset time period; generating the voice instruction directly according to the acquired in-vehicle voice information outside the preset time period. 9.A computer readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the steps of the voice interaction based instruction recognition method of any one of claims 1 to 8. 10.An electronic device, comprising: a processor and a memory; the memory is configured to store a computer program, and the processor is configured to execute the computer program stored in the memory to enable the electronic device to perform the steps of the voice interaction based instruction recognition method of any one of claims 1 to 8.
Citation Information
Patent Citations
Voice control method and voice control system used for vehicles
CN106887232A
On-vehicle voice recognition method and device thereof
CN108831462A
Interaction method and device of automobile intelligent terminal, computer equipment and storage medium
CN112802468A