Vehicle control method and device based on voice, storage medium, and electronic device
By using a preset voiceprint model to identify the voice of different people in the car in the car, and combining semantic information and sound intensity to generate vehicle control instructions, the accuracy of vehicle control in a mixed voice environment for multiple people is solved, and higher personalization and recognition accuracy are achieved, improving the driving experience.
Patent Information
- Application Number
- CN202310259974.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-16
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2043-03-16
AI Technical Summary
The existing vehicle voice interaction system is difficult to accurately identify specific voices in a mixed voice environment for multiple people, resulting in redundant information and confusing control permissions, and personalized vehicle control cannot be achieved.
By obtaining the voice signals collected by the vehicle microphone, a preset voiceprint model is used to identify the pronunciation object, and combining semantic information and sound intensity, corresponding vehicle control instructions are generated to achieve flexible control of the vehicle.
It improves the personalization of in-vehicle voice interaction and the accuracy of recognition of voice commands, enhances the user's driving experience, and provides a smarter, more pleasant and safer vehicle control method.
Smart Images

Figure CN116259320B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of vehicle control, and in particular to a voice-based vehicle control method and device, a storage medium, and an electronic device. Background Art
[0002] In related technologies, cars are evolving from traditional means of transportation to the "third space" of intelligent mobile travel. More and more intelligent devices are gradually and deeply integrated into human travel activities with cars as carriers, providing convenience from different dimensions such as driving experience, audio-visual entertainment, and assisted driving, creating a sense of technology. Among them, the in-vehicle voice interaction system, as the ears of the car, bears the responsibility of accepting the voice information input of the driver and passengers, identifying and translating it into codes that can be operated and executed by the car computer, realizing functions such as navigation route planning, entertainment equipment control, and in-vehicle electrical appliance coordination, covering a large number of daily driving needs of car owners, and at the same time replacing some auxiliary button operations (button / touch screen) also indirectly improves driving safety.
[0003] In actual use, the application scenarios of automobiles include not only the main driver, but also the co-driver and rear passengers. The voice information of different people in the car may overlap or even contradict each other. At the same time, with the improvement of the electrification level of the whole vehicle and the transition of the electronic and electrical architecture from distributed to domain control and then to central control, the richness of in-car electrical functions will increase explosively, including intelligent navigation, driving mode switching, air conditioning adjustment, wire-controlled chassis, seat ventilation / heating, etc. There is a high correlation between the main driver, co-driver, and passenger's demand for electrical functions and their role attributes and spatial positions in the car, but the current in-car voice interaction system usually adopts non-specific speaker-independent training SI non-specific speech model (speaker independent training), resulting in a large amount of information redundancy in the information input of the in-vehicle voice interaction system, which places high demands on the voice interaction algorithm and hardware capabilities. The current SD-specific voice interaction model (speaker dependent training), namely the voiceprint system, is limited by the voice recognition algorithm and hardware computing power. Its scope of use is limited to user identity confirmation when getting on the bus. It can realize personalized functions such as seat adjustment and act like a power-on password. It is currently unable to identify specific timbres from a variety of mixed audio sources with human voices.
[0004] With respect to the above-mentioned problems existing in the related technologies, no efficient and accurate solutions have been found yet. Summary of the invention
[0005] The present invention provides a voice-based vehicle control method and device, a storage medium, and an electronic device to solve technical problems in related technologies.
[0006] According to one embodiment of the present invention, a voice-based vehicle control method is provided, comprising: acquiring a voice signal collected by a vehicle microphone; using a preset voiceprint model to identify a pronunciation object of the voice signal, and identifying semantic information and sound intensity of the voice signal; searching for a control authority type that matches the sound intensity; and generating a vehicle control instruction of the control authority type based on the pronunciation object and the semantic information.
[0007] Furthermore, using a preset voiceprint model to identify the pronunciation object of the voice signal includes one of the following: using a preset voiceprint model to identify that the pronunciation object of the voice signal is the main driver user; using a preset voiceprint model to identify that the pronunciation object of the voice signal is the front-driver user; using a preset voiceprint model to identify that the pronunciation object of the voice signal is the rear passenger user.
[0008] Furthermore, using a preset voiceprint model to identify the pronunciation object of the voice signal includes: using a preset voiceprint model to identify a first pronunciation object of the voice signal; calculating the transmission time of the voice signal to each microphone in a vehicle microphone array, wherein the vehicle microphone array includes multiple microphones, and the layout position of each microphone corresponds to a driving position of the vehicle; selecting a target microphone with the shortest transmission time, and determining whether the layout position of the target microphone matches the driving position of the first pronunciation object; if the layout position of the target microphone matches the driving position of the first pronunciation object, outputting the first pronunciation object as the recognition result of the preset voiceprint model; if the layout position of the target microphone does not match the driving position of the first pronunciation object, outputting a second pronunciation object corresponding to the driving position of the target microphone as the recognition result of the preset voiceprint model.
[0009] Further, obtaining the speech signal collected by the vehicle microphone includes: determining the sound arrival time t of each microphone in the vehicle microphone array and the sound amplitude Lp received by each microphone; and calculating the double factor coefficient of each microphone using the following formula: D i =a(t i -t min )+b(Lp i -Lp min ), where D i is the double factor coefficient of the i-th microphone, a is the time factor coefficient, t i is the sound arrival time of the i-th microphone, t min is the minimum sound arrival time of all microphones, b is the sound pressure decibel factor coefficient, Lp iis the sound amplitude received by the i-th microphone, Lp min is the minimum sound amplitude of all microphones; select the sound source microphone with the largest double-factor coefficient, and obtain the voice signal collected by the sound source microphone.
[0010] Further, searching for a control authority type that matches the sound intensity includes: locating a target intensity interval in which the sound intensity is located; if the target intensity interval is in a first interval, matching the sound intensity with a first control authority type; if the target intensity interval is in a second interval, matching the sound intensity with a second control authority type, wherein the minimum value of the second interval is greater than the maximum value of the first interval, the first control authority type includes one of the following: vehicle safety control authority, ride comfort adjustment authority, audio and video entertainment control authority, and the second control authority type includes: emergency control authority.
[0011] Furthermore, there are multiple pronunciation objects, and generating a vehicle control instruction of the control authority type based on the pronunciation objects and the semantic information includes: extracting the voiceprint components of each pronunciation object from the voice signal respectively to obtain multiple voiceprint components; searching for component weight distribution information matching the control authority type; calculating the control priority of each pronunciation object based on the component weight distribution information and the multiple voiceprint components, and outputting the target voiceprint component with the highest priority; and generating a vehicle control instruction of the control authority type according to the semantic information of the target voiceprint component.
[0012] Furthermore, identifying the semantic information of the speech signal includes: discretizing the speech signal into frames through a moving window function to obtain multiple discrete segments; performing fast Fourier transform on the waveforms of the multiple discrete segments respectively to convert the time domain signals of the multiple discrete segments into frequency domain signals to obtain an observation sequence matrix, wherein the observation sequence matrix includes multiple observation sequences, and each observation sequence corresponds to a discrete segment; using a pre-constructed hidden Markov chain model to perform semantic analysis on the observation sequence to obtain the semantic information of the observation sequence matrix, wherein the hidden Markov chain model includes observation probability, transition probability and language probability, the observation probability represents the corresponding probability of each observation sequence and each state, the transition probability is used to describe the probability of the state of each observation sequence transferring to itself or to the next state, and the language probability is used to describe the probability of each observation sequence obtained according to the statistical law of language.
[0013] According to another embodiment of the present invention, a voice-based vehicle control device is provided, comprising: an acquisition module for acquiring a voice signal collected by a vehicle microphone; a recognition module for identifying a pronunciation object of the voice signal using a preset voiceprint model, and identifying semantic information and sound intensity of the voice signal; a search module for searching for a control authority type that matches the sound intensity; and a generation module for generating a vehicle control instruction of the control authority type based on the pronunciation object and the semantic information.
[0014] Furthermore, the recognition module includes one of the following: a first recognition unit, used to use a preset voiceprint model to identify that the pronunciation object of the voice signal is the main driving user; a second recognition unit, used to use a preset voiceprint model to identify that the pronunciation object of the voice signal is the co-pilot user; a third recognition unit, used to use a preset voiceprint model to identify that the pronunciation object of the voice signal is the rear passenger user.
[0015] Furthermore, the recognition module includes: a fourth recognition unit, used to identify the first pronunciation object of the voice signal by using a preset voiceprint model; a calculation unit, used to calculate the transmission time of the voice signal to each microphone in the vehicle microphone array, wherein the vehicle microphone array includes multiple microphones, and the layout position of each microphone corresponds to a driving position of the vehicle; a judgment unit, used to select a target microphone with the shortest transmission time, and judge whether the layout position of the target microphone matches the driving position of the first pronunciation object; an output unit, used to output the first pronunciation object as the recognition result of the preset voiceprint model if the layout position of the target microphone matches the driving position of the first pronunciation object; if the layout position of the target microphone does not match the driving position of the first pronunciation object, output the second pronunciation object of the driving position corresponding to the target microphone as the recognition result of the preset voiceprint model.
[0016] Furthermore, the acquisition module includes: a determination unit for determining the sound arrival time t of each microphone in the vehicle microphone array and the sound amplitude Lp received by each microphone; a calculation unit for calculating the double factor coefficient of each microphone using the following formula: D i =a(t i -t min )+b(Lp i -Lp min ), where D i is the double factor coefficient of the i-th microphone, a is the time factor coefficient, t i is the sound arrival time of the i-th microphone, t min is the minimum sound arrival time of all microphones, b is the sound pressure decibel factor coefficient, Lp i is the sound amplitude received by the i-th microphone, Lpmin is the minimum sound amplitude of all microphones; an acquisition unit is used to select the sound source microphone with the largest double-factor coefficient and acquire the voice signal collected by the sound source microphone.
[0017] Further, the search module includes: a positioning unit, used to locate the target intensity interval of the sound intensity; a matching unit, used to match the sound intensity with a first control authority type if the target intensity interval is in a first interval; and to match the sound intensity with a second control authority type if the target intensity interval is in a second interval, wherein the minimum value of the second interval is greater than the maximum value of the first interval, the first control authority type includes one of the following: vehicle safety control authority, ride comfort adjustment authority, audio and video entertainment control authority, and the second control authority type includes: emergency control authority.
[0018] Furthermore, there are multiple pronunciation objects, and the generation module includes: an extraction unit, used to extract the voiceprint components of each pronunciation object from the speech signal respectively, to obtain multiple voiceprint components; a search unit, used to search for component weight distribution information matching the control authority type; a calculation unit, used to calculate the control priority of each pronunciation object based on the component weight distribution information and the multiple voiceprint components, and output the target voiceprint component with the highest priority; a generation unit, used to generate a vehicle control instruction of the control authority type according to the semantic information of the target voiceprint component.
[0019] Furthermore, the recognition module includes: a discrete unit, which is used to discretize the speech signal into frames through a moving window function to obtain multiple discrete segments; a transformation unit, which is used to perform fast Fourier transform on the waveforms of the multiple discrete segments respectively to convert the time domain signals of the multiple discrete segments into frequency domain signals to obtain an observation sequence matrix, wherein the observation sequence matrix includes multiple observation sequences, and each observation sequence corresponds to a discrete segment; a parsing unit, which is used to perform semantic parsing on the observation sequence using a pre-constructed hidden Markov chain model to obtain semantic information of the observation sequence matrix, wherein the hidden Markov chain model includes observation probability, transition probability and language probability, the observation probability represents the corresponding probability of each observation sequence and each state, the transition probability is used to describe the probability of the state of each observation sequence transferring to itself or to the next state, and the language probability is used to describe the probability of each observation sequence obtained according to the statistical law of language.
[0020] According to another aspect of an embodiment of the present application, a storage medium is further provided, which includes a stored program, and the above steps are executed when the program is run.
[0021] According to another aspect of an embodiment of the present application, there is also provided an electronic device, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; wherein: the memory is used to store computer programs; and the processor is used to execute the steps in the above method by running the program stored in the memory.
[0022] The embodiment of the present application also provides a computer program product including instructions, which, when executed on a computer, enables the computer to execute the steps in the above method.
[0023] Through the embodiments of the present invention, a voice signal collected by a vehicle microphone is obtained, a preset voiceprint model is used to identify the pronunciation object of the voice signal, and the semantic information and sound intensity of the voice signal are identified, the control authority type matching the sound intensity is found, and a vehicle control instruction of the control authority type is generated based on the pronunciation object and the semantic information. By identifying the pronunciation object, semantic information and sound intensity of the voice signal, a vehicle control instruction of the control authority type corresponding to the sound intensity is generated based on the semantic information of the pronunciation object, thereby realizing a flexible generation method of vehicle control instructions, which can flexibly control the vehicle by voice, solves the technical problem of related technologies that only authenticated users are allowed to control the vehicle by voice, improves the personalization of in-vehicle voice interaction and the recognition accuracy of voice instructions, and thus brings users a smarter, more pleasant and safer driving experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0025] Figure 1 is a hardware structure block diagram of a vehicle-mounted terminal according to an embodiment of the present invention;
[0026] Figure 2 is a flow chart of a voice-based vehicle control method according to an embodiment of the present invention;
[0027] Figure 3 is a schematic diagram of the distribution of microphones and driving seats in an implementation scenario of the present invention;
[0028] Figure 4 is a schematic diagram of multi-microphone array positioning in an embodiment of the present invention;
[0029] Figure 5 is a flowchart of voiceprint recognition and multi-microphone array calibration according to an embodiment of the present invention;
[0030] Figure 6is a flowchart of the operation of semantic recognition of sound signals in an embodiment of the present invention;
[0031] Figure 7 A logic diagram of vehicle-mounted voice decision-making in an embodiment of the present invention;
[0032] Figure 8 4 is a structural block diagram of a voice-based vehicle control device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0033] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only embodiments of a part of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work should fall within the scope of protection of the present application. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0035] Example 1
[0036] The method embodiment provided in the first embodiment of the present application can be executed in a vehicle-mounted terminal, a vehicle control module, a voice control module, or a similar processing device. Taking running on a vehicle-mounted terminal as an example, Figure 1 FIG. 1 is a hardware structure diagram of a vehicle-mounted terminal according to an embodiment of the present invention. Figure 1 As shown, the vehicle terminal may include one or more ( Figure 1 Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Optionally, the vehicle terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1The structure shown is only for illustration and does not limit the structure of the vehicle-mounted terminal. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.
[0037] The memory 104 can be used to store vehicle terminal programs, for example, software programs and modules of application software, such as a vehicle terminal program corresponding to a voice-based vehicle control method in an embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the vehicle terminal program stored in the memory 104, that is, to implement the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the vehicle terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0038] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the vehicle terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0039] In this embodiment, a vehicle control method based on voice is provided. Figure 2 is a flow chart of a voice-based vehicle control method according to an embodiment of the present invention, such as Figure 2 As shown, the process includes the following steps:
[0040] Step S202, obtaining a voice signal collected by a vehicle microphone;
[0041] The vehicle microphone of this embodiment may be one or more microphones arranged on the vehicle, and the voice signal may be a single person's voice emitted by a single object (driver or passenger) or a mixed voice of multiple persons.
[0042] Figure 3This is a schematic diagram of the distribution of microphones and driver's seats in an implementation scenario of the present invention. Microphones are arranged at different locations in the car to collect original voice signals, including the main driver's seat (microphone 2), the co-driver's seat (microphone 1), and the left and right rear passenger seats a / b (left: microphone 3; right: microphone 4). At the same time, sound-absorbing materials are arranged symmetrically in the car to effectively suppress the reverberation phenomenon in the car and reduce its negative effect on voice recognition.
[0043] Step S204, using a preset voiceprint model to identify the pronunciation object of the voice signal, and to identify the semantic information and sound intensity of the voice signal;
[0044] Step S206, searching for a control authority type that matches the sound intensity;
[0045] Optionally, the vehicle's control authority type can be classified according to control priority, control scope, etc., such as dividing the control authority type into core information, important information, auxiliary information and emergency information. Core information is information that directly affects vehicle safety, including engine ignition, hybrid / pure electric vehicle power-on, body stability system EPS / ECS activation, adaptive cruise control ACC, automatic driving, etc. Important information may include information that affects ride comfort, such as seat adjustment, air conditioning adjustment, window lifting, seat ventilation, etc. Auxiliary information can be for in-car audio and video entertainment needs, including music playback, news information and other multimedia functions. Emergency information includes the vehicle's braking system, early warning system, etc.
[0046] Step S208: generating a vehicle control instruction of a control authority type based on the pronunciation object and the semantic information.
[0047] Optionally, vehicle control instructions are used to control the vehicle's power system, computer system, audio and video system and other vehicle components, such as controlling vehicle start and stop, adjusting the air conditioner, adjusting the speakers, braking, accelerating, shifting gears, changing lanes, etc.
[0048] Through the above steps, the voice signal collected by the vehicle microphone is obtained, the preset voiceprint model is used to identify the pronunciation object of the voice signal, and the semantic information and sound intensity of the voice signal are identified, the control authority type matching the sound intensity is found, and the vehicle control instruction of the control authority type is generated based on the pronunciation object and the semantic information. By identifying the pronunciation object, semantic information and sound intensity of the voice signal, a vehicle control instruction of the control authority type corresponding to the sound intensity is generated based on the semantic information of the pronunciation object, thereby realizing a flexible generation method of vehicle control instructions, which can flexibly control the vehicle by voice, solves the technical problem that the related technology only allows authenticated users to control the vehicle by voice, improves the personalization of the in-vehicle voice interaction and the recognition accuracy of the voice instructions, and thus brings users a smarter, more pleasant and safer driving experience.
[0049] In one implementation of this embodiment, obtaining the speech signal collected by the vehicle microphone includes: determining the sound arrival time t of each microphone in the vehicle microphone array and the sound amplitude Lp received by each microphone; and calculating the double factor coefficient of each microphone using the following formula: D i =a(t i -t min )+b(Lp i -Lp min ), where D i is the double factor coefficient of the ith microphone, a is the time factor coefficient, t i is the sound arrival time of the i-th microphone, t min is the minimum sound arrival time of all microphones, b is the sound pressure decibel factor coefficient, Lp i is the sound amplitude received by the i-th microphone, Lp min is the minimum sound amplitude of all microphones; select the sound source microphone with the largest double factor coefficient, and obtain the speech signal collected by the sound source microphone.
[0050] Figure 4 is a schematic diagram of multi-microphone array positioning in an embodiment of the present invention, and the judgment parameter is the time t at which different microphones receive sound signals. i , and the received sound signal amplitude Lp i , where the subscript i represents the microphone code, t 0 Lp is the time when the sound signal is emitted. 0 is the original amplitude of the sound signal (without considering the time and space attenuation of the signal). Since the time when the sound signal is emitted cannot be obtained in actual engineering operations, 0 and the original amplitude Lp 0 Before locating the sound source, t 1 ,t 2 ,t 3 ,t 4 and Lp 1 ,Lp 2 ,Lp 3 ,Lp 4 Sort and compare, and use the relative offset to infer the spatial orientation of the sound source. The specific implementation method is as follows: find t 1 ,t 2 ,t 3 ,t 4 The minimum value t in min and Lp 1 ,Lp 2 ,Lp 3 ,Lp 4 The minimum value Lp in minAccording to the propagation formula of sound waves (1), the propagation speed c of sound waves is only related to the temperature and pressure of the propagation medium. The speed of sound waves under the conditions of 1 standard atmospheric pressure and 15°C is about 340m / s. Therefore, the sound source is at t 0 After the sound is emitted at time t, the sound detected by different microphones is detected at time t according to the spatial distance from the sound source. i With distance L i There is a linear negative correlation. The closer the microphone is to the sound source, the earlier it detects it.
[0051]
[0052] c represents the speed of sound wave propagation, γ is the adiabatic index of the propagation medium, which is related to temperature, P is the pressure of the propagation medium, and ρ is the density of the propagation medium.
[0053] The original sound amplitude decibel number Lp can be calculated according to formula (2), where original sounds less than 40dB are considered as interference noise and do not enter the subsequent signal processing and vehicle computer decision-making process. Original sounds greater than 40db but less than 70db are considered core / important / auxiliary information. Original sounds greater than 70db are considered urgent information. After calculation, the decibel number corresponding to the main voiceprint is defined as Lp 1 , the decibel number corresponding to the secondary voiceprint is defined as Lp 2 , the decibel level of the guest’s voiceprint is defined as Lp 3 .
[0054]
[0055] Where Lp represents the decibel number after the original sound is converted, P rms is the RMS value of the amplitude of the sound sampling point, P ref It is the reference value of sound amplitude.
[0056] According to the propagation attenuation law of sound waves in space, if the distance from the sound source is L 1 The sound pressure level at Lp 1 , then at a distance L from the sound source 2 The sound pressure level at Lp 2 It can be calculated based on formula (3).
[0057]
[0058] Where Lp i is the sound pressure decibel detected by the i-th microphone, L i is the spatial distance between the ith microphone and the sound source. The closer the microphone is to the sound source, the greater the decibel value of the sound pressure detected.
[0059] Based on formula (4), the double-factor sound source detection model is defined, and the double-factor coefficient Di In order to comprehensively consider the theoretical basis of propagation time and sound pressure decibel number, D i The larger the coefficient is, the greater the probability that a sound source is close to the i-th microphone. The corresponding logical authority can be given to the sound source according to the microphone position to input and control the vehicle system.
[0060] D i =a(t i -t min )+b(Lp i -Lp min ) (4)
[0061] Where D i is the double factor coefficient of the i-th microphone, a is the time factor coefficient, t i is the arrival time of the sound from the ith microphone, b is the sound pressure decibel factor coefficient, Lp i is the sound amplitude received by the i-th microphone. The farther the microphone is from the sound source, the farther the sound arrives, and the more serious the sound pressure amplitude attenuation is, that is, the time factor coefficient a is a negative number, and the sound pressure decibel factor coefficient b is a positive number. The a coefficient and b coefficient of different models are calibrated by the OEM before leaving the factory, which is related to the microphone selection and the spatial layout of the microphone. The specific calibration process will not be repeated here.
[0062] Finally, select the double factor coefficient D i The largest microphone is used as the sound source microphone closest to the sound source, and the sound information recorded by this microphone is used as the original signal of the sound source recognized by the vehicle system. The sound source signals recorded by other microphones are ignored to prevent cross-interference of repeated signals and affect subsequent semantic analysis and function execution. The sound characteristics of the sound source can be determined based on the final selected microphone position, and verified and proofread with the corresponding voiceprint signal. For example, the sound selected based on microphone 2 is the main voiceprint, the sound selected based on microphone 1 is the secondary voiceprint, and the sounds selected based on microphones 3 and 4 are not distinguished and are all defined as guest voiceprints.
[0063] The pronunciation object of this embodiment can be but is not limited to the main driver, the co-driver, and the rear-seat guest. In one implementation of this embodiment, the pronunciation object of the voice signal recognized by the preset voiceprint model includes one of the following: the pronunciation object of the voice signal recognized by the preset voiceprint model is the main driver; the pronunciation object of the voice signal recognized by the preset voiceprint model is the co-driver; the pronunciation object of the voice signal recognized by the preset voiceprint model is the rear-seat guest.
[0064] In the scenario of this embodiment, three types of voiceprints are set up: main voiceprint, secondary voiceprint and guest voiceprint. Specifically, the main voiceprint is the owner of the vehicle, usually the car owner or a full-time driver. The daily car use scenario is to control the driving of the vehicle in the main driving position, assume the safety responsibility of the vehicle, and have full access to the equipment in the car. The secondary voiceprint is a fixed passenger of the vehicle, usually the spouse or other family members of the car owner. The daily car use scenario is to sit in the co-pilot seat, which can affect the riding experience of the vehicle. It is not responsible for the safety of the vehicle and has the right to use some of the equipment in the car. The guest voiceprint is a non-fixed passenger of the vehicle, usually the relatives and friends of the car owner and other random passengers. The daily car use scenario is to sit in the back seat and mainly use the multimedia equipment in the car. It is also not responsible for the safety of the vehicle and has the right to use a few devices in the car.
[0065] In this embodiment, a fixed wake-up word is used to activate the in-vehicle voice system, and the in-vehicle microphone is used to collect the original sound signal in the car space. The echo cancellation algorithm is used to suppress the environmental noise of the collected sound signal. The wavelet noise reduction algorithm is used to further denoise the audio frequency signal. The pre-processed voice signal is subjected to a fast Fourier transform according to formula (5) to convert the sound signal from the time domain space to the frequency domain space, thereby extracting the main vector feature of the sound, and matching and comparing it with the voiceprint model. Based on the model similarity, it is determined whether the sound is the main voiceprint, secondary voiceprint or guest voiceprint.
[0066]
[0067] Where t represents time, f(t) represents the function of the sound amplitude changing with time, ω is the sound frequency to be converted, i is the imaginary unit, and F(ω) represents the function of the sound amplitude changing with frequency ω.
[0068] In one example, using a preset voiceprint model to identify the pronunciation object of a speech signal includes: using a preset voiceprint model to identify a first pronunciation object of a speech signal; calculating the transmission time of the speech signal to each microphone in a vehicle microphone array, wherein the vehicle microphone array includes multiple microphones and the layout position of each microphone corresponds to a driving position of the vehicle; selecting a target microphone with the shortest transmission time, and determining whether the layout position of the target microphone matches the driving position of the first pronunciation object; if the layout position of the target microphone matches the driving position of the first pronunciation object, outputting the first pronunciation object as the recognition result of the preset voiceprint model; if the layout position of the target microphone does not match the driving position of the first pronunciation object, outputting the second pronunciation object of the target microphone corresponding to the driving position as the recognition result of the preset voiceprint model.
[0069] In this embodiment, in order to reduce the computing power hardware requirements of the vehicle system, when training the preset voiceprint model, the personalized voiceprint (main voiceprint, secondary voiceprint) can be recorded offline through a smartphone, imported into the vehicle system in the form of online data transmission and updated regularly. Therefore, the vehicle system does not need to bear the deep self-learning ability of the specific voice model, but only needs to have the interpretation and execution ability of the specific voice model that has been recorded offline, avoiding the process of the vehicle system generating a specific voice model, and improving the robustness and stability of the vehicle system. Considering the actual scenario of the owner's change, the offline recorded voice model can be updated on demand. Smartphones generally have the necessary hardware for recording voiceprints, such as microphones and high-computing SoC processors. Therefore, from the perspective of user-side feasibility, there is no need to add additional hardware costs, and the computing power of idle devices can be fully utilized. The fast Fourier transform is used to transform the time domain signal of the sound into the frequency domain signal, the time series features of the sound into frequency features, the main vector of the sound frequency matrix is extracted, and it is mapped to a specific voiceprint model using machine learning. With the help of wireless communication, the specific voiceprint model in the mobile phone communicates with the vehicle system regularly, so that the vehicle system can also recognize the voice characteristics of the main driver and the co-driver, and has the feature of being updated over time.
[0070] Considering that the noise in the car comes from all directions, and the voice signals are usually overlapped and fused in the time domain and frequency domain, coupled with the reverberation phenomenon formed by multiple reflections of sound waves, etc., the difficulty of accurate voice recognition and reliable signal processing is further increased. In order to further improve the recognition accuracy of the specific speech model, the sound collection process is differentiated according to the spatial position of the passengers in the car. Compared with the sound signal collected by a single microphone, the temporal and spatial characteristics of the voice signal are integrated. For cars, since the position of passengers in the car is relatively fixed and passengers cannot move around, the accuracy of sound source positioning is higher, and the calibration model parameter range can be set more aggressively.
[0071] like Figure 3 As shown, due to the one-to-one correspondence between the main voiceprint-main driver's seat, the secondary voiceprint-secondary driver's seat, and the guest voiceprint-guest seat a / b, when the voiceprint recognition model is accurate enough, the received voice signal can be distinguished, located, and enabled with corresponding functions only by the voiceprint model of the vehicle system. However, in actual applications, considering that the recognition accuracy of the vehicle's voiceprint model, especially the specific voiceprint model (such as the main voiceprint and the secondary voiceprint), needs to rely on a rich voice source database, it is difficult to achieve a high accuracy rate in the early stage of vehicle use. The present invention adopts a combination of voiceprint model + multi-microphone array positioning for sound matching. Figure 5This is a flowchart of the voiceprint recognition and multi-microphone array calibration of an embodiment of the present invention. The voiceprint model is used to perform the initial recognition of the voice source, and the multi-microphone array positioning is used to verify the recognition accuracy of the voiceprint model. The multi-microphone array positioning adopts the TDOA (Time Difference Of Arrival) algorithm, that is, the time difference of arrival of each microphone to the sound source is used for diagnosis. Figure 4 As shown, this embodiment newly adds the amplitude difference when the sound source reaches the microphone, that is, to comprehensively judge which microphone is finally selected to collect the sound source for subsequent semantic analysis of the voice source and implementation of the vehicle function. The storage space of the vehicle system is used to record the results of each voiceprint recognition and the verification results of the multi-microphone array positioning, and regularly transmit them back to the smartphone to improve the voice source database of the specific voiceprint model, and continuously iterate the voiceprint model to improve the recognition accuracy of the main voiceprint and the secondary voiceprint.
[0072] After the specific / non-specific voiceprint model built into the vehicle completes the initial recognition of the sound source of the main driver, co-driver and guest positions, in order to ensure the accuracy and reliability of command recognition, before the original sound after recognition is subsequently semantically analyzed and the vehicle function is executed, it is necessary to use the multi-microphone array in the vehicle to perform a secondary verification of the initial voiceprint recognition result. For the results of the initial voiceprint recognition and the secondary multi-microphone array verification that are consistent, such as a sound source is identified as the main driver's voiceprint by the voiceprint model, and the positioning results of the multi-microphone array in the vehicle also show that the sound source is emitted from the main driver's position, then the sound source is identified as the main driver's voiceprint by the vehicle system, and the corresponding permissions of the main voiceprint are granted and enter the subsequent vehicle semantic analysis system to complete the implementation of the corresponding vehicle function. For inconsistent results between the initial voiceprint recognition and the secondary multi-microphone array verification, such as a sound source identified as the main driver's voiceprint by the voiceprint model, but the multi-microphone array positioning result in the car shows that the sound source is emitted from the co-pilot position, the sound source is identified as the co-pilot voiceprint by the car system (based on the multi-microphone array positioning), and the functional permissions corresponding to the co-pilot voiceprint are granted for semantic analysis. The storage module of the car system synchronously transmits the matching results of the initial recognition result of the voiceprint model and the secondary verification result of the multi-microphone array model to the smartphone terminal, continuously enriching the sound source database of the specific voiceprint model and regularly improving the recognition accuracy of the specific voiceprint model of the car system through two-way communication.
[0073] In an example of the present embodiment, the semantic information of the recognized speech signal includes: discretizing the speech signal into frames through a moving window function to obtain multiple discrete segments; performing fast Fourier transform on the waveforms of the multiple discrete segments respectively to convert the time domain signals of the multiple discrete segments into frequency domain signals to obtain an observation sequence matrix, wherein the observation sequence matrix includes multiple observation sequences, and each observation sequence corresponds to a discrete segment; using a pre-built hidden Markov chain model to perform semantic analysis on the observation sequence to obtain the semantic information of the observation sequence matrix, wherein the hidden Markov chain model includes observation probability, transition probability and language probability, the observation probability represents the corresponding probability of each observation sequence and each state, the transition probability is used to describe the probability of the state of each observation sequence transferring to itself or to the next state, and the language probability is used to describe the probability of each observation sequence obtained according to the statistical law of language.
[0074] After determining the signal strength of the main voiceprint, secondary voiceprint and guest voiceprint in the voice signal (i.e. decibel number), it is necessary to perform semantic recognition on the sound signal. After the voiceprint model recognition and multi-microphone array positioning verification, the voice information is randomly delivered to the car computer for semantic analysis, translation and execution of the corresponding car function, to obtain the action intention of the sound source and determine the actual function expected to be realized by the car computer. The specific semantic recognition method is as follows: Figure 6 As shown, Figure 6 This is a flowchart of the operation flow of the semantic recognition of sound signals in an embodiment of the present invention. First, the original signal needs to be discretized into frames through a moving window function, that is, the continuous sound is cut into discrete small segments, and the waveform is fast Fourier transformed to convert the time domain signal into a frequency domain signal, and the original sound is converted into an observation sequence matrix containing mathematical features. A hidden Markov chain model is constructed to describe the state of the observation sequence, and the observation sequence is semantically parsed with reference to the three major indicators of observation probability, transition probability and language probability. Among them, the observation probability represents the corresponding probability of each original frame and each state (i.e., the information parsing of a single frame), the transition probability describes the probability of each state transferring to itself or to the next state (i.e., the information parsing of adjacent frames), and the language probability describes the probability obtained according to the statistical law of language (obtained by machine learning training based on a large number of text sources, i.e., matching sources). After performing the above operations on the original signal, the sound information can be interpreted as actual semantics, and subsequent vehicle computer command judgments can be performed.
[0075] In one implementation of the present embodiment, searching for a control authority type that matches the sound intensity includes: locating a target intensity interval where the sound intensity is located; if the target intensity interval is in a first interval, the sound intensity matches a first control authority type; if the target intensity interval is in a second interval, the sound intensity matches a second control authority type, wherein the minimum value of the second interval is greater than the maximum value of the first interval, the first control authority type includes one of the following: vehicle safety control authority, ride comfort adjustment authority, audio and video entertainment control authority, and the second control authority type includes: emergency control authority.
[0076] In one example, the first interval is 40-70dB, and the second interval is 70dB-∞. The microphone is used to measure and evaluate the volume of the sound. Usually, the human voice is 40-60dB, and the call sound in an emergency can reach 70dB. The in-vehicle voice comprehensive decision information can be obtained by weighting the volume decibel value according to Table 1. Sounds with an original volume greater than 40dB are included in the evaluation system of core information, important information, and auxiliary information, and sounds greater than 70dB are included in the evaluation system of emergency information. Sounds below 40dB are determined to be environmental interference signals and are not included in the evaluation system.
[0077] In this example, the first control authority type includes control authority for core information, important information, and auxiliary information, and the second control authority type includes control authority for emergency information.
[0078] In one implementation scenario, there are multiple pronunciation objects, and generating a vehicle control instruction of a control authority type based on the pronunciation objects and semantic information includes: extracting the voiceprint components of each pronunciation object from the speech signal to obtain multiple voiceprint components; searching for component weight distribution information that matches the control authority type; calculating the control priority of each pronunciation object based on the component weight distribution information and the multiple voiceprint components, and outputting the target voiceprint component with the highest priority; and generating a vehicle control instruction of a control authority type according to the semantic information of the target voiceprint component.
[0079] In this embodiment, the component semantic information of each voiceprint component can also be identified, and it can be determined whether the control targets corresponding to the component semantic information of all pronunciation objects are the same (such as all for controlling the air conditioner, all for controlling the volume, etc.). If they are the same, a vehicle control instruction of the control authority type is generated according to the semantic information of the target voiceprint component.
[0080] In this embodiment, different weight ratios may be configured for each type of pronunciation object, as shown in Table 1.
[0081] Table 1
[0082] Control permission type Main voiceprint Secondary voiceprint Guest Voiceprint Weight Total core 100 0 0 100 important 50 50 0 100 Assistance 50 25 25 100 urgent 34 33 33 100
[0083] According to the weights assigned to each voiceprint in Table 1, the voice information is analyzed. For the core information interpreted as directly affecting vehicle safety, the weights of each voiceprint are (main voiceprint 100*Lp 1 , secondary voiceprint 0*Lp 2 , guest voiceprint 0*Lp 3 ), that is, only the main voiceprint (main driver) can control the core information of the vehicle by voice, and the voices of the co-pilot and the guests cannot interfere with the core information, thereby ensuring the driving safety of the vehicle. For important information that is interpreted as affecting the comfort of the ride, the weights of each voiceprint are (main voiceprint 50*Lp 1 , secondary voiceprint 50*Lp 2 , guest voiceprint 0*Lp 3 ), relying on the specific details of semantic recognition to determine whether it is the same command (such as whether it is to adjust the air conditioning temperature in the car, etc.), if it is not the same command, the car computer will execute the corresponding function operations of the main voiceprint and the secondary voiceprint in the order of acceptance. If it is the same command, first determine 50*Lp 1 and 50*Lp 2 The relationship between the size of the two is that the larger value is input to the car computer to execute the corresponding function, and the smaller value is ignored. For the auxiliary information identified as the in-car audio and video entertainment needs, the final decision can be made and input to the car computer. It is worth mentioning that the influence of the guest's voiceprint on the final result needs to be considered at this time. The specific operation method is consistent with the important information decision-making method described above.
[0084] For sounds identified as emergency information (greater than 70dB), such as "brake", "be careful", "someone", etc., after the weighted value of the information is calculated as described above, it is not input into the car system at the same time, but the execution order is comprehensively judged by combining the recognized semantics and the time for the car to perform the corresponding operation (delay time + execution time). For example, if the semantics of the main voiceprint, the secondary voiceprint and the guest voiceprint are judged to be the same instruction, the instruction is sent to the car and the corresponding operation is quickly executed. If the main voiceprint, the secondary voiceprint and the guest voiceprint are judged to be different instructions, the execution order of the instructions is selected according to the urgency of the instructions and the actual operation speed, with the principle of minimizing the impact of the accident, and the actions of the three voiceprints are executed in sequence, and finally the operation intentions of the main driver, the co-driver and the rear passengers are realized.
[0085] It should be mentioned that the classification method of voiceprints and vehicle information described in this patent, including the corresponding weights of different voiceprints and information (Table 1), and the specific values in the division method of sound volume are introduced for the convenience of expression and cannot limit the actual use scenario of this patent. Any change in the classification method, the corresponding weight of voiceprint / information, and the behavior of calibrating the numerical parameters according to the debugging results can realize the solution of this embodiment.
[0086] Figure 7 The following is a logic diagram of the in-vehicle voice decision-making judgment in the embodiment of the present invention. First, the car computer is initialized, including recording the exclusive specific voiceprints of the main driver and the co-driver by the mobile terminal offline, connecting the car computer and the smart phone by wireless communication, and importing the main voiceprint of the main driver and the secondary voiceprint of the co-driver into the car computer system. The non-specific guest voiceprint is constructed by the OEM based on the voice big data source and pre-installed into the car computer system when leaving the factory. At the same time, the car computer system will collect the matching results of the user's specific voiceprint (main / secondary voiceprint) and the verification results of the voiceprint recognition accuracy of the multi-microphone array positioning system, and the above matching results and verification results data will be regularly transmitted back to the smart phone, and the sound source database of the specific voiceprint will be continuously improved. The high-computing power SoC chip of the smart phone will regularly update the specific voiceprint, improve the accuracy of voice recognition for the main driver and the co-driver, and continuously improve the user experience of core users. The voiceprint recognition of non-core users is continuously improved by the regular OTA update of the manufacturer's non-specific voiceprint model.
[0087] The in-vehicle voice hierarchical interaction method based on voiceprint recognition provided in this embodiment is based on voice recognition algorithm and hardware computing power. Through the hierarchical architecture design and control logic optimization of the system, the in-vehicle application scope of the specific voice interaction model is broadened, the personalization of the in-vehicle voice interaction and the recognition accuracy of the voice commands are improved, thereby bringing users a smarter, more pleasant and safer driving experience.
[0088] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0089] Example 2
[0090] In this embodiment, a voice-based vehicle control device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0091] Figure 8is a structural block diagram of a speech-based vehicle control device according to an embodiment of the present invention, such as Figure 8 As shown, the device comprises:
[0092] An acquisition module 80 is used to acquire a voice signal collected by a vehicle microphone;
[0093] The recognition module 82 is used to recognize the pronunciation object of the voice signal by using a preset voiceprint model, and recognize the semantic information and sound intensity of the voice signal;
[0094] A search module 84, used to search for a control authority type matching the sound intensity;
[0095] The generating module 86 is used to generate a vehicle control instruction of the control authority type based on the pronunciation object and the semantic information.
[0096] Optionally, the recognition module includes one of the following: a first recognition unit, used to use a preset voiceprint model to identify that the pronunciation object of the voice signal is the main driving user; a second recognition unit, used to use a preset voiceprint model to identify that the pronunciation object of the voice signal is the co-pilot user; a third recognition unit, used to use a preset voiceprint model to identify that the pronunciation object of the voice signal is the rear passenger user.
[0097] Optionally, the recognition module includes: a fourth recognition unit, used to identify the first pronunciation object of the voice signal by using a preset voiceprint model; a calculation unit, used to calculate the transmission time of the voice signal to each microphone in the vehicle microphone array, wherein the vehicle microphone array includes multiple microphones, and the layout position of each microphone corresponds to a driving position of the vehicle; a judgment unit, used to select a target microphone with the shortest transmission time, and judge whether the layout position of the target microphone matches the driving position of the first pronunciation object; an output unit, used to output the first pronunciation object as the recognition result of the preset voiceprint model if the layout position of the target microphone matches the driving position of the first pronunciation object; if the layout position of the target microphone does not match the driving position of the first pronunciation object, output the second pronunciation object of the driving position corresponding to the target microphone as the recognition result of the preset voiceprint model.
[0098] Optionally, the acquisition module includes: a determination unit, used to determine the sound arrival time t of each microphone in the vehicle microphone array, and the sound amplitude Lp received by each microphone; a calculation unit, used to calculate the double factor coefficient of each microphone using the following formula: D i =a(t i -t min )+b(Lp i -Lp min ), where D iis the double factor coefficient of the ith microphone, a is the time factor coefficient, t i is the sound arrival time of the i-th microphone, t min is the minimum sound arrival time of all microphones, b is the sound pressure decibel factor coefficient, Lp i is the sound amplitude received by the i-th microphone, Lp min is the minimum sound amplitude of all microphones; an acquisition unit is used to select the sound source microphone with the largest double-factor coefficient and acquire the voice signal collected by the sound source microphone.
[0099] Optionally, the search module includes: a positioning unit, used to locate the target intensity interval of the sound intensity; a matching unit, used to match the sound intensity with a first control authority type if the target intensity interval is in a first interval; and to match the sound intensity with a second control authority type if the target intensity interval is in a second interval, wherein the minimum value of the second interval is greater than the maximum value of the first interval, the first control authority type includes one of the following: vehicle safety control authority, ride comfort adjustment authority, audio and video entertainment control authority, and the second control authority type includes: emergency control authority.
[0100] Optionally, there are multiple pronunciation objects, and the generation module includes: an extraction unit, used to extract the voiceprint components of each pronunciation object from the speech signal respectively, to obtain multiple voiceprint components; a search unit, used to search for component weight distribution information matching the control authority type; a calculation unit, used to calculate the control priority of each pronunciation object based on the component weight distribution information and the multiple voiceprint components, and output the target voiceprint component with the highest priority; a generation unit, used to generate a vehicle control instruction of the control authority type according to the semantic information of the target voiceprint component.
[0101] Optionally, the recognition module includes: a discrete unit, used to discretize the speech signal into frames through a moving window function to obtain multiple discrete segments; a transformation unit, used to perform fast Fourier transform on the waveforms of the multiple discrete segments respectively, so as to convert the time domain signals of the multiple discrete segments into frequency domain signals to obtain an observation sequence matrix, wherein the observation sequence matrix includes multiple observation sequences, and each observation sequence corresponds to a discrete segment; a parsing unit, used to perform semantic parsing on the observation sequence using a pre-constructed hidden Markov chain model to obtain semantic information of the observation sequence matrix, wherein the hidden Markov chain model includes observation probability, transition probability and language probability, the observation probability represents the corresponding probability of each observation sequence and each state, the transition probability is used to describe the probability of the state of each observation sequence transferring to itself or to the next state, and the language probability is used to describe the probability of each observation sequence obtained according to language statistical laws.
[0102] It should be noted that the above modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.
[0103] Example 3
[0104] An embodiment of the present invention further provides a storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when running.
[0105] Optionally, in this embodiment, the storage medium may be configured to store a computer program for performing the following steps:
[0106] S1, obtaining the voice signal collected by the vehicle microphone;
[0107] S2, using a preset voiceprint model to identify the pronunciation object of the voice signal, and to identify the semantic information and sound intensity of the voice signal;
[0108] S3, searching for a control authority type that matches the sound intensity;
[0109] S4: Generate a vehicle control instruction of the control authority type based on the pronunciation object and the semantic information.
[0110] Optionally, in this embodiment, the above-mentioned storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store computer programs.
[0111] An embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0112] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0113] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:
[0114] S1, obtaining the voice signal collected by the vehicle microphone;
[0115] S2, using a preset voiceprint model to identify the pronunciation object of the voice signal, and to identify the semantic information and sound intensity of the voice signal;
[0116] S3, searching for a control authority type that matches the sound intensity;
[0117] S4: Generate a vehicle control instruction of the control authority type based on the pronunciation object and the semantic information.
[0118] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.
[0119] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0120] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0121] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0122] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0123] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0124] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk and other media that can store program codes.
[0125] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A vehicle control method based on voice, It is characterized in that include: Obtain the voice signal collected by the vehicle microphone; Using a preset voiceprint model to identify the pronunciation object of the voice signal, and to identify the semantic information and sound intensity of the voice signal; Searching for a control authority type that matches the sound intensity, wherein the control authority type corresponds to a control priority or a control range of the vehicle; generating a vehicle control instruction of the control authority type based on the pronunciation object and the semantic information; Among them, there are multiple pronunciation objects, and generating a vehicle control instruction of the control authority type based on the pronunciation objects and the semantic information includes: extracting the voiceprint components of each pronunciation object from the voice signal respectively to obtain multiple voiceprint components; searching for component weight distribution information matching the control authority type; calculating the control priority of each pronunciation object based on the component weight distribution information and the multiple voiceprint components, and outputting the target voiceprint component with the highest priority; and generating a vehicle control instruction of the control authority type according to the semantic information of the target voiceprint component.
2. The method according to claim 1, It is characterized in that Using a preset voiceprint model to identify the pronunciation object of the voice signal includes one of the following: Using a preset voiceprint model to identify the pronunciation object of the voice signal as the main driving user; Using a preset voiceprint model to identify the pronunciation object of the voice signal as the front passenger user; A preset voiceprint model is used to identify that the pronunciation object of the voice signal is a back-seat guest user.
3. The method according to claim 1, It is characterized in that The method of using a preset voiceprint model to identify the pronunciation object of the voice signal includes: Using a preset voiceprint model to identify a first pronunciation object of the voice signal; Calculating the transmission time of the speech signal to each microphone in the vehicle microphone array, wherein the vehicle microphone array includes a plurality of microphones, and the layout position of each microphone corresponds to a driving position of the vehicle; Selecting a target microphone with the shortest transmission time, and determining whether the layout position of the target microphone matches the driving position of the first pronunciation object; If the layout position of the target microphone matches the driving position of the first pronunciation object, the first pronunciation object is output as the recognition result of the preset voiceprint model; if the layout position of the target microphone does not match the driving position of the first pronunciation object, the second pronunciation object corresponding to the driving position of the target microphone is output as the recognition result of the preset voiceprint model.
4. The method according to claim 1, It is characterized in that Acquiring the voice signal collected by the vehicle microphone includes: Determine the sound arrival time t of each microphone in the vehicle microphone array and the sound amplitude received by each microphone Lp; The two-factor coefficient for each microphone is calculated using the following formula: ; Among them, D i is the double factor coefficient of the ith microphone, a is the time factor coefficient, t i is the sound arrival time of the i-th microphone, t min is the minimum sound arrival time of all microphones, b is the sound pressure decibel factor coefficient, Lp i is the sound amplitude received by the i-th microphone, Lp min is the minimum sound amplitude of all microphones; A sound source microphone with a maximum dual-factor coefficient is selected, and a speech signal collected by the sound source microphone is obtained.
5. The method according to claim 1, It is characterized in that The types of control permissions that match the sound intensity include: Locating a target intensity interval of the sound intensity; If the target intensity interval is in the first interval, the sound intensity matches the first control authority type; if the target intensity interval is in the second interval, the sound intensity matches the second control authority type, wherein the minimum value of the second interval is greater than the maximum value of the first interval, the first control authority type includes one of the following: vehicle safety control authority, ride comfort adjustment authority, audio and video entertainment control authority, and the second control authority type includes: emergency control authority.
6. The method according to claim 1, It is characterized in that The semantic information of the speech signal is identified including: The speech signal needs to be discretized into frames by a moving window function to obtain a plurality of discrete segments; Performing fast Fourier transform on the waveforms of the plurality of discrete segments respectively, so as to convert the time domain signals of the plurality of discrete segments into frequency domain signals, and obtaining an observation sequence matrix, wherein the observation sequence matrix includes a plurality of observation sequences, each observation sequence corresponding to a discrete segment; A pre-built hidden Markov chain model is used to perform semantic analysis on the observation sequence to obtain the semantic information of the observation sequence matrix, wherein the hidden Markov chain model includes observation probability, transition probability and language probability. The observation probability represents the corresponding probability of each observation sequence and each state. The transition probability is used to describe the probability that the state of each observation sequence transfers to itself or to the next state. The language probability is used to describe the probability of each observation sequence obtained according to the language statistical law.
7. A vehicle control device based on voice, It is characterized in that include: An acquisition module, used to acquire a voice signal collected by a vehicle microphone; A recognition module, used to recognize the pronunciation object of the voice signal by using a preset voiceprint model, and to recognize the semantic information and sound intensity of the voice signal; A search module, used to search for a control authority type that matches the sound intensity, wherein the control authority type corresponds to a control priority or a control range of the vehicle; A generating module, configured to generate a vehicle control instruction of the control authority type based on the pronunciation object and the semantic information; Among them, there are multiple pronunciation objects, and the generation module includes: an extraction unit, used to extract the voiceprint components of each pronunciation object from the speech signal respectively, to obtain multiple voiceprint components; a search unit, used to search for component weight distribution information matching the control authority type; a calculation unit, used to calculate the control priority of each pronunciation object based on the component weight distribution information and the multiple voiceprint components, and output the target voiceprint component with the highest priority; a generation unit, used to generate a vehicle control instruction of the control authority type according to the semantic information of the target voiceprint component.
8. A storage medium, It is characterized in that The storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 6 when executed.
9. An electronic device comprising a memory and a processor, It is characterized in that A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Vehicle control method and device and vehicle-mounted terminal
CN109410938A
Interaction method, vehicle-mounted terminal and computer readable storage medium
CN114582336A