A vehicle voice control method and related device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-18
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]鉴于上述问题,本申请提供了一种车辆语音控制方法及相关装置,以解决现有技术仅依赖声纹识别的方式容易被攻击者利用,存在安全隐患的问题
[0055] Based on this, this application can further generate auxiliary verification factors based on the target subsequence, and then determine whether the voice control command is issued by an authorized user in real time based on voiceprint similarity and the auxiliary verification factors. Since the real user in the target subsequence is highly likely to be the real speaker of the voice control command, the auxiliary verification factors generated based on the target subsequence can be used as a basis for judgment to determine whether the voice control command is issued by an authorized user. Therefore, compared with the judgment method that relies solely on voiceprint similarity, this application combines voiceprint similarity and auxiliary verification factors, which can more accurately determine whether the voice control command is issued by an authorized user in real time. If so, the vehicle is controlled by voice based on the audio waveform signal. Since this application not only considers audio-related features but also integrates video-related features under sound source localization, it effectively prevents voice control commands from being exploited by attackers and improves vehicle security.
Smart Images

Figure CN122551805A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle technology, and in particular to a vehicle voice control method and related device. Background Technology
[0002] In existing in-vehicle voice control systems, user authentication typically relies solely on voiceprint recognition. Specifically, if the voiceprint of the target voice matches the voiceprint of a pre-stored authorized user, it is determined that the target voice comes from an authorized user, and then voice control is performed on the target vehicle based on the target voice.
[0003] However, vehicle voice control methods that rely solely on voiceprint recognition are easily exploited by attackers, posing security risks. Summary of the Invention
[0004] In view of the above problems, this application provides a vehicle voice control method and related device to solve the security risks posed by existing technologies that rely solely on voiceprint recognition, which are easily exploited by attackers. The specific solution is as follows:
[0005] The first aspect of this application provides a vehicle voice control method, including:
[0006] The system acquires audio waveform signals collected by the microphone array on the vehicle in response to voice control commands, the sound source localization results of the voice control commands, and the voiceprint similarity corresponding to the audio waveform signals.
[0007] Based on the sound source localization result, multiple candidate subsequences are determined from the target video sequence. The candidate subsequences consist of the region of interest of the same user in each video frame. The user is located within a preset neighborhood of the sound source localization result. The target video sequence is a video sequence inside and outside the vehicle determined by the timestamp of the voice control command arriving at the microphone array.
[0008] The candidate subsequences are cross-correlated with the audio waveform signal respectively to determine the candidate subsequence with the highest cross-correlation with the audio waveform signal from the candidate subsequences, which is then used as the target subsequence;
[0009] An auxiliary verification factor is generated based on the target sub-sequence. Based on the voiceprint similarity and the auxiliary verification factor, it is determined whether the voice control command is issued by an authorized user in real time.
[0010] If so, the vehicle is controlled by voice based on the audio waveform signal.
[0011] In one possible implementation, before performing cross-correlation calculations between the plurality of candidate sub-sequences and the audio waveform signal, the method further includes:
[0012] The candidate sub-sequences are matched with the registered face images to obtain the face matching degree corresponding to each candidate sub-sequence;
[0013] From the plurality of candidate sub-sequences, determine a number of candidate sub-sequences whose face matching degree is higher than a preset matching degree threshold;
[0014] The aforementioned candidate subsequences are used as multiple candidate subsequences in the cross-correlation calculation step.
[0015] In one possible implementation, the step of using the plurality of candidate subsequences as multiple candidate subsequences in the cross-correlation calculation step includes:
[0016] Facial micro-motion detection is performed on the candidate sub-sequences to determine whether there are candidate sub-sequences in which the user's facial expression is not moving.
[0017] If so, the candidate subsequence with no facial movement of the user is removed from the plurality of candidate subsequences, and the remaining candidate subsequences are used as multiple candidate subsequences in the cross-correlation calculation step.
[0018] In one possible implementation, the step of performing cross-correlation calculations between the plurality of candidate sub-sequences and the audio waveform signal respectively includes:
[0019] Optical flow calculations are performed on the multiple candidate sub-sequences to obtain the motion energy sequences corresponding to the multiple candidate sub-sequences;
[0020] Determine the audio energy sequence of the audio waveform signal;
[0021] The cross-correlation between the motion energy sequences corresponding to the multiple candidate sub-sequences and the audio energy sequence is calculated, and is used as the cross-correlation between the multiple candidate sub-sequences and the audio waveform signal.
[0022] In one possible implementation, generating the auxiliary verification factor based on the target subsequence includes:
[0023] Anti-fraud detection is performed based on the audio waveform signal and the target sub-sequence to obtain an anti-fraud score;
[0024] Determine the positional matching degree between the user's actual location in the target subsequence and the sound source localization result;
[0025] Obtain the time decay memory value of the user's successful authentication within the most recent preset time period in the target subsequence;
[0026] The target subsequence is matched with the registered face image to obtain the face matching degree corresponding to the target subsequence;
[0027] At least one of the anti-fraud score, the location matching degree, the time decay memory value, and the face matching degree corresponding to the target subsequence is used as the auxiliary verification factor.
[0028] In one possible implementation, the anti-fraud detection based on the audio waveform signal and the target subsequence to obtain an anti-fraud score includes:
[0029] Determine the audio spectrum characteristics of the audio waveform signal, and determine a first confidence probability that the audio waveform signal is a waveform signal played by the terminal device based on the audio spectrum characteristics;
[0030] Facial micro-movement detection is performed on the target sub-sequence to determine the second confidence probability that the user's facial movements are not in the target sub-sequence.
[0031] The anti-fraud score is generated based on the first confidence probability and the second confidence probability.
[0032] In one possible implementation, determining whether the voice control command is issued in real time by an authorized user based on the voiceprint similarity and the auxiliary verification factor includes:
[0033] The voiceprint similarity and the auxiliary verification factor are concatenated into an input feature vector, which is then input into a pre-trained sound source reliability prediction model to obtain a reliability prediction probability value. The sound source reliability prediction model is trained using a training input feature vector labeled with a reliability prediction probability value.
[0034] If the reliability prediction probability value is greater than or equal to a preset first probability threshold, then it is determined that the voice control command is issued in real time by the authorized user corresponding to the registered face image;
[0035] If the reliability prediction probability value is less than the preset second probability threshold, it is determined that the voice control command is not issued in real time by the authorized user corresponding to the registered face image, and the second probability threshold is less than the first probability threshold.
[0036] If the reliability prediction probability value is less than the first probability threshold and greater than or equal to the second probability threshold, the authorized user is notified to repeat the voice control command in order to initiate secondary verification based on the repeated voice control command.
[0037] In one possible implementation, the voice control of the vehicle based on the audio waveform signal includes:
[0038] Determine the instruction level corresponding to the audio waveform signal, and determine the effective area of the instruction level;
[0039] If the sound source localization result is located within the effective area, then the vehicle is controlled by voice based on the audio waveform signal.
[0040] In one possible implementation, the process of determining the sound source localization result includes:
[0041] The timestamps of the voice control commands arriving at each microphone in the microphone array are obtained to obtain the audio arrival timestamps corresponding to the multiple microphones in the microphone array.
[0042] Based on the position coordinates of each of the multiple microphones, the audio arrival timestamps corresponding to the multiple microphones, and the sound source location, multiple time delay difference equations are established.
[0043] The multiple time delay difference equations are solved simultaneously to obtain the solution value of the sound source location, which is used as the sound source localization result.
[0044] A second aspect of this application provides a vehicle voice control device, comprising:
[0045] The data acquisition unit is used to acquire audio waveform signals collected by the microphone array on the vehicle in response to voice control commands, the sound source localization results of the voice control commands, and the voiceprint similarity corresponding to the audio waveform signals.
[0046] The subsequence extraction unit is used to determine multiple candidate subsequences from the target video sequence based on the sound source localization result. The candidate subsequences are composed of the region of interest of the same user in each video frame. The user is located within a preset neighborhood of the sound source localization result. The target video sequence is a video sequence inside and outside the vehicle determined by the timestamp of the voice control command arriving at the microphone array.
[0047] The subsequence filtering unit is used to perform cross-correlation calculations between the plurality of candidate subsequences and the audio waveform signal respectively, so as to determine the candidate subsequence with the highest cross-correlation degree with the audio waveform signal from the plurality of candidate subsequences, and use it as the target subsequence;
[0048] The instruction control unit is used to generate an auxiliary verification factor based on the target sub-sequence, and determine whether the voice control command is issued by an authorized user in real time based on the voiceprint similarity and the auxiliary verification factor. If so, the unit performs voice control on the vehicle based on the audio waveform signal.
[0049] A third aspect of this application provides a computer program product including computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the vehicle voice control method of the first aspect or any implementation thereof.
[0050] A fourth aspect of this application provides an electronic device, including at least one processor and a memory connected to the processor, wherein:
[0051] The memory is used to store computer programs;
[0052] The processor is used to execute the computer program so that the electronic device can implement the vehicle voice control method of the first aspect or any implementation thereof.
[0053] The fifth aspect of this application provides a computer storage medium carrying one or more computer programs, which, when executed by an electronic device, enable the electronic device to implement the vehicle voice control method described in the first aspect or any implementation thereof.
[0054] By utilizing the above technical solution, the vehicle voice control method provided in this application takes into account that if a user issues a voice control command, the user's real location is theoretically consistent with the sound source localization result of the voice control command. At the same time, the user's video sequence must have a very strong cross-correlation with the audio waveform signal of the voice control command. Therefore, in order to determine the real speaker of the voice control command, this application can obtain the audio waveform signal collected by the microphone array on the vehicle for the voice control command, the sound source localization result of the voice control command, and the voiceprint similarity corresponding to the audio waveform signal. Based on the sound source localization result, multiple candidate sub-sequences are determined from the target video sequence. Then, the multiple candidate sub-sequences are cross-correlated with the audio waveform signal respectively to determine the candidate sub-sequence with the highest cross-correlation with the audio waveform signal from the multiple candidate sub-sequences as the target sub-sequence. Then, the real user in the target sub-sequence is highly likely to be the real speaker of the voice control command.
[0055] Based on this, this application can further generate auxiliary verification factors based on the target subsequence, and then determine whether the voice control command is issued by an authorized user in real time based on voiceprint similarity and the auxiliary verification factors. Since the real user in the target subsequence is highly likely to be the real speaker of the voice control command, the auxiliary verification factors generated based on the target subsequence can be used as a basis for judgment to determine whether the voice control command is issued by an authorized user. Therefore, compared with the judgment method that relies solely on voiceprint similarity, this application combines voiceprint similarity and auxiliary verification factors, which can more accurately determine whether the voice control command is issued by an authorized user in real time. If so, the vehicle is controlled by voice based on the audio waveform signal. Since this application not only considers audio-related features but also integrates video-related features under sound source localization, it effectively prevents voice control commands from being exploited by attackers and improves vehicle security. Attached Figure Description
[0056] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0057] Figure 1 A flowchart illustrating a vehicle voice control method provided in this application;
[0058] Figure 2 A schematic diagram of the structure of a vehicle voice control device provided in this application;
[0059] Figure 3 This is a schematic diagram of the structure of an electronic device provided in this application. Detailed Implementation
[0060] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.
[0061] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.
[0062] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0063] To enable those skilled in the art to better understand this application, the vehicle voice control method of this application embodiments will be described in detail below with reference to the accompanying drawings.
[0064] Reference Figure 1 , Figure 1 This is a flowchart illustrating a vehicle voice control method provided in an embodiment of this application, as follows: Figure 1 As shown, the vehicle voice control method may include:
[0065] Step S101: Obtain the audio waveform signal collected by the microphone array on the vehicle in response to the voice control command, the sound source localization result of the voice control command, and the voiceprint similarity corresponding to the audio waveform signal.
[0066] Specifically, in this embodiment, multiple microphones can be scientifically deployed on the exterior and interior of the vehicle body, both at the front and rear and around the vehicle, to form a microphone array. The multiple microphones in the microphone array are connected to the vehicle acoustic processing unit through precise clock synchronization and hardware wiring to collect sounds inside and outside the vehicle.
[0067] Therefore, when a user issues a voice control command, the microphone array can collect audio waveform signals for the voice control command. At the same time, it can determine the timestamp of the voice control command arriving at each microphone in the microphone array, and use this to locate the sound source and obtain the sound source localization result of the voice control command.
[0068] In order to identify whether the voice control command is issued by a registered authorized user (such as a car owner), this embodiment can also extract the voiceprint of the audio waveform signal and perform similarity calculation with the registered voiceprint to obtain the voiceprint similarity corresponding to the audio waveform signal.
[0069] Step S102: Determine multiple candidate subsequences from the target video sequence based on the sound source localization results.
[0070] Among them, the candidate subsequence consists of the region of interest of the same user in each video frame, the user is located within the preset neighborhood of the sound source localization result, and the target video sequence is the in-vehicle and out-of-vehicle video sequence determined by the timestamp of the voice control command arriving at the microphone array.
[0071] As described above, this embodiment can record the timestamps of voice control commands arriving at the microphone array, including the timestamps of the voice control commands arriving at each microphone within the microphone array. Since there is a delay in the arrival of voice control commands at the microphone array, the actual timestamp of the user issuing the voice control command is earlier than the first timestamp of the voice control command arriving at the microphone array. Therefore, optionally, a video sequence with a preset duration or preset number of frames preceding the first timestamp (optionally, this could be a video sequence of the interior and / or exterior of the vehicle captured by a camera installed in the vehicle, or a video sequence of the interior and / or exterior of the vehicle captured by a surveillance camera, or a video sequence of the interior and / or exterior of the vehicle captured by a human (such as the vehicle owner) camera, etc.) can be used as the target video sequence.
[0072] Optionally, a video sequence with a preset duration or preset number of frames after the first timestamp can be added to the aforementioned target video sequence to serve as the target video sequence in this step.
[0073] It should be understood that if the voice control command is given by a real person, then theoretically, there must be a real user among the users in the surrounding neighborhood of the sound source localization result within the target video sequence who gave the voice control command. Therefore, this embodiment can search for multiple candidate users (specifically real users in this application, and the same applies below) within the target video sequence based on the sound source localization result, and extract the region of interest for each user in each video frame of the target video sequence, thereby forming a candidate subsequence composed of the region of interest for that user. Thus, multiple candidate subsequences can be determined from the target video sequence based on the sound source localization result.
[0074] More specifically, in order to find users within the neighborhood of the sound source localization results, the sound source localization results can be projected onto the camera image plane using the camera's intrinsic and extrinsic parameters to obtain the image positions of the sound source localization results in each video frame of the target video sequence. Then, based on the image positions, users that meet the requirements and the candidate subsequences in which the users are located can be found.
[0075] Those skilled in the art will understand that voice control commands in vehicles are typically brief, so the target video sequence usually does not need to span a long period of time; that is, the target video sequence is generally a short video sequence. Within a short video sequence, the user typically does not experience significant frame shifts. Therefore, if a user is located within the vicinity of the sound source localization result in a certain video frame, they are likely to remain within the vicinity of the sound source localization result in other video frames as well. Of course, there may be exceptions. Therefore, optionally, for each user within the target video sequence, if the user is located within the vicinity of the sound source localization result in more than one video frame, then the region of interest for that user is extracted to form a candidate subsequence.
[0076] Step S103: Perform cross-correlation calculations between multiple candidate sub-sequences and the audio waveform signal to determine the candidate sub-sequence with the highest cross-correlation degree with the audio waveform signal from multiple candidate sub-sequences, and use it as the target sub-sequence.
[0077] Considering that if a user issues a voice control command, the user's candidate subsequence must have a strong cross-correlation with the audio waveform signal of the voice control command, the candidate subsequence containing the user who is more likely to issue the voice control command can be selected from multiple candidate subsequences, i.e., the target subsequence.
[0078] Based on this, this embodiment can calculate the cross-correlation between each candidate subsequence and the audio waveform signal to obtain the cross-correlation between each candidate subsequence and the audio waveform signal. Then, the target subsequence with the highest cross-correlation with the audio waveform signal can be determined from multiple candidate subsequences. Therefore, the user in the target subsequence is highly likely to be a real user (i.e., a real speaker) initiating the voice control command.
[0079] Step S104: Generate auxiliary verification factors based on the target subsequence, and determine whether the voice control command is issued in real time by an authorized user based on voiceprint similarity and auxiliary verification factors.
[0080] Here, authorized users refer to users who have registered their identity features such as facial features, fingerprints, and voiceprints.
[0081] Since the target subsequence is likely the "video" sequence in which the real user who initiated the voice control command is located, the target subsequence can be used to generate auxiliary verification factors to verify whether the voice control command was issued by an authorized user. Then, based on the voiceprint similarity and the auxiliary verification factors, it can be determined whether the voice control command was issued by an authorized user in real time.
[0082] Since the target subsequence is the subsequence predicted by this application, it is difficult to avoid inaccuracies in the prediction process. Therefore, in a more preferred case, this embodiment can use voiceprint similarity as the main judgment criterion and auxiliary verification factor as the auxiliary judgment criterion to determine whether the voice control command is issued by the authorized user in real time. That is, compared with voiceprint similarity, the auxiliary judgment factor has a smaller weight in the judgment process to improve the accuracy of the judgment result.
[0083] Step S105: If yes, then perform voice control on the vehicle based on the audio waveform signal.
[0084] The vehicle voice control method provided in this application takes into account that if a user issues a voice control command, the user's real location is theoretically consistent with the sound source localization result of the voice control command. At the same time, the user's video sequence must have a very strong cross-correlation with the audio waveform signal of the voice control command. Therefore, in order to determine the real speaker of the voice control command, this application can obtain the audio waveform signal collected by the microphone array on the vehicle for the voice control command, the sound source localization result of the voice control command, and the voiceprint similarity corresponding to the audio waveform signal. Based on the sound source localization result, multiple candidate sub-sequences are determined from the target video sequence. Then, the multiple candidate sub-sequences are cross-correlated with the audio waveform signal respectively to determine the candidate sub-sequence with the highest cross-correlation with the audio waveform signal from the multiple candidate sub-sequences as the target sub-sequence. Then, the real user in the target sub-sequence is highly likely to be the real speaker of the voice control command.
[0085] Based on this, this application can further generate auxiliary verification factors based on the target subsequence, and then determine whether the voice control command is issued by an authorized user in real time based on voiceprint similarity and the auxiliary verification factors. Since the real user in the target subsequence is highly likely to be the real speaker of the voice control command, the auxiliary verification factors generated based on the target subsequence can be used as a basis for judgment to determine whether the voice control command is issued by an authorized user. Therefore, compared with the judgment method that relies solely on voiceprint similarity, this application combines voiceprint similarity and auxiliary verification factors, which can more accurately determine whether the voice control command is issued by an authorized user in real time. If so, the vehicle is controlled by voice based on the audio waveform signal. Since this application not only considers audio-related features but also integrates video-related features under sound source localization, it effectively prevents voice control commands from being exploited by attackers and improves vehicle security.
[0086] In some embodiments of this application, the process of determining the sound source localization result in step S101 above is described in detail.
[0087] In this embodiment, the timestamps of the voice control commands arriving at each microphone in the microphone array can be obtained to obtain the audio arrival timestamps corresponding to the multiple microphones in the microphone array. Then, based on the position coordinates of the multiple microphones, the audio arrival timestamps corresponding to the multiple microphones, and the sound source location, multiple time delay difference equations are established. Finally, the multiple time delay difference equations are solved simultaneously to obtain the solution value of the sound source location, which is used as the sound source localization result.
[0088] Optionally, a time delay difference equation can be established based on the position coordinates of each pair of microphones, the corresponding audio arrival timestamp, and the sound source location, as shown in the following formula:
[0089] Formula (1);
[0090] in, Indicates the location of the sound source. Indicates the first microphone in the array The location coordinates of each microphone Indicates the first microphone in the array The location coordinates of each microphone Indicates the speed of sound. Indicates the first The microphone and the first The theoretical time difference of arrival (TDOA) of a microphone pair consisting of 1 microphone.
[0091] Optionally, a weighted least squares algorithm can be used to solve the system of problems simultaneously to obtain the sound source localization result.
[0092] If the sound source localization result is a two-dimensional position coordinate, then four microphones can be used to solve the problem simultaneously. If the sound source localization result is a three-dimensional position coordinate (which requires the addition of known microphone height values), then five microphones can be used to solve the problem simultaneously.
[0093] Taking two-dimensional position coordinates as an example, we define for To the The distance of one microphone, =1, 2, 3, 4. Solving the simultaneous equations, we get... Initial estimate:
[0094] Formula (2);
[0095] Formula (3);
[0096] Formula (4);
[0097] in, Represents the coefficient matrix. Represents a constant vector. Indicates the first The x-coordinate of each microphone. Indicates the first The y-coordinate of each microphone Represents the x-coordinate of the sound source. This represents the y-coordinate of the sound source. Indicates the first The distance from each microphone to the first microphone is known, and its value is equal to the corresponding TDOA multiplied by the speed of sound.
[0098] Optionally, to improve robustness, RANSAC (Random Sample Consensus) or the M-estimation algorithm can be used to remove outlier TDOAs.
[0099] This embodiment uses a sound source localization method based on time delay difference, which achieves higher localization accuracy.
[0100] In some embodiments of this application, the process of “performing cross-correlation calculations between multiple candidate sub-sequences and audio waveform signals respectively” in step S103 above is described in detail.
[0101] In order to calculate the cross-correlation between the candidate sub-sequences and the audio waveform signal, this embodiment can first perform optical flow calculations on multiple candidate sub-sequences to obtain the motion energy sequences corresponding to each candidate sub-sequence. Here, the motion energy sequences are preferably the motion energy sequences of the mouth and upper body.
[0102] Optionally, PWC-Net (Pyramid, Warping, and Cost Volume Network) or Farneback (Farneback Optical Flow Algorithm) can be used to calculate the optical flow for each candidate subsequence to obtain the motion energy sequence corresponding to each candidate subsequence.
[0103] Of course, other optical flow calculation algorithms can also be used, and this embodiment does not impose specific limitations.
[0104] Furthermore, the audio energy sequence of the audio waveform signal, i.e., the short-time audio energy envelope sequence, can be determined. Then, the cross-correlation between the motion energy sequences corresponding to the multiple candidate sub-sequences and the audio energy sequence is calculated, and these are taken as the cross-correlation between the multiple candidate sub-sequences and the audio waveform signal.
[0105] In one possible implementation, considering that many candidate sub-sequences may be extracted from the target video sequence, in order to improve the calculation efficiency of cross-correlation in this embodiment, a face matching strategy can be used for preliminary screening before performing cross-correlation calculations on multiple candidate sub-sequences with the audio waveform signal.
[0106] That is, in this embodiment, multiple candidate sub-sequences can be matched with the registered face image to obtain the face matching degree corresponding to each candidate sub-sequence. Then, several candidate sub-sequences with a face matching degree higher than a preset matching degree threshold are determined from the multiple candidate sub-sequences, and these several candidate sub-sequences are used as multiple candidate sub-sequences in the cross-correlation calculation step above.
[0107] Optionally, RetinaFace or MTCNN (Multi-task Cascaded Convolutional Networks) can be used to perform face detection on the candidate sub-sequences to obtain face detection results. Then, the face detection results are compared with the registered face images to obtain the face matching degree, which represents the confidence level of the user's identity in the candidate sub-sequences. Based on this, those candidate sub-sequences that are more likely to be authorized users are selected.
[0108] Because face matching has high matching accuracy, it can quickly filter out most of the candidate subsequences that do not meet the matching degree requirements from multiple candidate subsequences, thus reducing the computational cost of cross-correlation calculation.
[0109] In another possible implementation, it is also possible to identify whether the user's mouth or face is not moving. If there is no movement, it means that the user is not making a sound. Based on this, some candidate subsequences that do not meet the requirements can be filtered out.
[0110] Specifically, facial micro-motion detection can be performed on several candidate sub-sequences to determine whether there are candidate sub-sequences in which the user's facial expressions are not moving. If so, the candidate sub-sequences in which the user's facial expressions are not moving are removed from the candidate sub-sequences, and the remaining candidate sub-sequences are used as multiple candidate sub-sequences in the cross-correlation calculation step.
[0111] Optionally, a deep learning model for detecting facial micro-movements can be pre-trained, and the candidate sub-sequences can be input into the deep learning model to obtain the judgment result of whether the user's facial movements are absent in the candidate sub-sequences. The deep learning model is trained using training sub-sequences labeled with the judgment result.
[0112] It should be noted that the above screening mechanism is only an example. In addition, there are other screening methods, such as using human pose estimation algorithms based on OpenPose or HRNet (High-Resolution Network) to determine whether the user is facing away from the vehicle. If so, the user is likely not the speaker of the voice control command, and the candidate subsequence of that user can be screened out.
[0113] This embodiment employs multiple screening mechanisms to eliminate candidate subsequences, reducing the computational load of cross-correlation calculations. This allows for a faster identification of the target subsequence that is more likely to be the actual speaker of the voice control command from multiple candidate subsequences, thus improving efficiency.
[0114] In other embodiments of this application, the process of “generating auxiliary verification factors based on target subsequences and determining whether voice control commands are issued by authorized users in real time based on voiceprint similarity and auxiliary verification factors” in step S104 above will be described in detail.
[0115] Optionally, the auxiliary verification factor can be at least one of the following factors: anti-fraud score, location matching degree, time decay and consistency and face matching degree corresponding to the target subsequence.
[0116] Optionally, anti-fraud detection can be performed based on the audio waveform signal and the target subsequence to obtain an anti-fraud score. Here, the anti-fraud score represents the probability that the voice control command was not issued by a user in the target subsequence.
[0117] More specifically, considering the spectral characteristics of recorded or replicated sounds played by terminal devices, there are significant differences in multiple dimensions between them and the spectral characteristics of sounds emitted by real users in real time. For example, the human voice has a continuous and rich spectrum, especially in the high-frequency harmonics and extremely low-frequency parts, with smooth and natural transitions. However, recorded or replicated sounds may lack extremely low and high frequency components due to the limited frequency response range of terminal devices. Furthermore, recorded or replicated sounds may have been compressed, losing a large amount of high-frequency details and low-frequency resonance. For another example, human voices often have vibrato and slight volume fluctuations, and the beginning and end of each phoneme are very rapid and crisp. However, the fundamental frequency curve of recorded or replicated sounds is often relatively smooth. Even though modern neural speech synthesis technology can mimic the slight volume fluctuations of human voices, its statistical patterns still have subtle differences from those of real people. In addition, terminal devices have difficulty perfectly reproducing rapid transient responses, resulting in the beginning of phonemes not being as natural as human voices. Based on this, optionally, this embodiment can determine the audio spectrum characteristics of the audio waveform signal, and then determine the first confidence probability that the audio waveform signal is a waveform signal played by the terminal device based on the audio spectrum characteristics. The higher the first confidence probability, the greater the possibility that the audio waveform signal is a waveform signal played by the terminal device, and the more likely the voice control command is not the voice of the user in the target subsequence.
[0118] It is also understandable that if the user in the target subsequence makes a sound, then their mouth and face must be able to detect micro-movements. Therefore, this embodiment can perform mouth and face micro-movement detection on the target subsequence to determine the second confidence probability that the user in the target subsequence has no mouth and face movement. The higher the second confidence probability, the greater the probability that the user in the target subsequence has no mouth and face movement, and the more likely that the voice control command is not the sound made by the user in the target subsequence.
[0119] This embodiment can generate an anti-fraud score based on a first confidence probability and a second confidence probability. For example, the anti-fraud score can be obtained by weighted summing of the first confidence probability and the second confidence probability.
[0120] Of course, there are other methods for determining anti-fraud scores. For example, facial micro-motion detection has high accuracy and can relatively accurately detect whether a user in a target subsequence is making facial micro-motions. If not, the anti-fraud score is 1; if so, the anti-fraud score is 0 or close to a minimum value of 0.
[0121] Optionally, the positional matching degree between the user's actual location in the target subsequence and the sound source localization result can be determined. The higher the positional matching degree, the more likely the voice control command is to be the sound emitted by the user in the target subsequence.
[0122] Optionally, this embodiment can also record each user who has successfully authenticated within the most recent preset time period (e.g., the most recent month), including but not limited to: the number of times each user has successfully authenticated, the timestamp of each user's most recent successful authentication, etc. Then, the identity identifiers of users in the target subsequence can be matched with historical records to obtain the time decay memory value of users in the target subsequence who have successfully authenticated within the most recent preset time period.
[0123] Optionally, a preset time decay function can be used, such as To obtain the above time decay memory value, where, This represents the time-decayed memory value at time t (e.g., the current time). This represents the timestamp of the user's most recent successful authentication in the target subsequence. and This indicates the preset adjustment coefficient.
[0124] Considering the possibility that some users in the target subsequence may never have successfully authenticated, optionally, if a user in the target subsequence has never successfully authenticated, the time decay memory value is 0; if a user in the target subsequence has a record of successful authentication within the most recent preset time period, the time decay function described above is used for calculation.
[0125] Of course, there are other ways to calculate the time decay memory value, such as using the exponential moving average algorithm, which calculates the time decay memory value based on the number of times a user in the target subsequence has successfully authenticated within the most recent preset time period. The more times a user has successfully authenticated within the most recent preset time period, the larger the time decay memory value will be.
[0126] Optionally, this embodiment can also perform face matching between the target subsequence and the registered face image to obtain the face matching degree corresponding to the target subsequence. For example, each region of interest in the target subsequence is compared with the feature vector of the registered face image, and then the average of the comparison results corresponding to each region of interest in the target subsequence is calculated as the face matching degree corresponding to the target subsequence; or, for another example, a region of interest is selected from the target subsequence and compared with the feature vector of the registered face image to obtain the face matching degree corresponding to the target subsequence.
[0127] It should be understood that other auxiliary verification factors may also be used, but this application does not specify any limitations.
[0128] In one possible implementation, the process of "determining whether a voice control command is issued in real time by an authorized user based on voiceprint similarity and auxiliary verification factors" can be as follows:
[0129] First, the voiceprint similarity and auxiliary verification factor are concatenated into an input feature vector, which is then input into a pre-trained sound source reliability prediction model to obtain a reliability prediction probability value. This sound source reliability prediction model is trained using a training input feature vector labeled with a reliability prediction probability value.
[0130] It should be noted that the process of concatenating voiceprint similarity and auxiliary verification factors into the input feature vector must be carried out in a predetermined fixed order to prevent the sound source reliability prediction model from making incorrect predictions.
[0131] Optionally, the sound source reliability prediction model can be a Bayesian model, a weighted fusion model, or other neural network models; this application does not impose specific limitations.
[0132] In this embodiment, a higher reliability prediction probability value indicates that the voice control command is more likely to be issued in real time by the authorized user corresponding to the registered face image. Therefore, this embodiment can preset a first probability threshold. If the reliability prediction probability value is greater than or equal to the first probability threshold, it is determined that the voice control command is issued in real time by the authorized user corresponding to the registered face image; otherwise, it is determined that the voice control command is not issued in real time by the authorized user corresponding to the registered face image.
[0133] Optionally, a second probability threshold can be preset. If the second probability threshold is less than the first probability threshold, and the reliability prediction probability value is less than the preset second probability threshold, it is determined that the voice control command is not issued in real time by the authorized user corresponding to the registered face image. If the reliability prediction probability value is less than the first probability threshold and greater than or equal to the second probability threshold, the authorized user is notified to repeat the voice control command in order to initiate secondary verification based on the repeated voice control command.
[0134] The secondary verification process can be the same as described above, or it can be performed only on users in the target subsequence, or other verification methods can be used. This application does not impose any specific limitations.
[0135] In this embodiment, multimodal and multidimensional auxiliary verification factors can be used to help determine whether the voice control command was issued by an authorized user in real time based on voiceprint similarity. This can effectively eliminate distant sound sources outside the vehicle (such as remote control recording) or environmental interference, and obtain more accurate verification results. This effectively improves the target uniqueness and credibility of the voice control command, thereby improving the security of the voice-controlled vehicle and reducing the possibility of the voice control function being exploited by attackers.
[0136] In some other embodiments of this application, the process of “voice control of the vehicle based on the audio waveform signal” in step S105 above will be described in detail.
[0137] Considering that authorized users typically issue voice control commands to unlock or start the vehicle from a distance, this embodiment can optionally divide the area around the vehicle into a multi-level concentric and sector-based security trust zone. For example, by default, the area can be divided into three concentric zones based on the center of the car: Level 1 (0–1m, near-field inside the vehicle), Level 2 (1–2.5m), and Level 3 (>2.5m). Alternatively, different importance levels can be further subdivided into sectors on the horizontal plane according to several azimuth angles.
[0138] Of course, the specific trust zone division can be adaptively expanded or contracted based on environmental noise levels, map constraints, and user history habits; this application does not impose specific limitations. Here, map constraints refer to: using the vehicle's current geographical location and environment type (home garage, public parking lot, roadside, etc.), dynamically adjusting the radius of each level of trust zone and the types of authorized commands based on the map and scene labels.
[0139] Furthermore, different voice control commands can be applied according to different trust zones. For example, in the first-level trust zone, voice control commands requiring high security, such as starting the vehicle, shifting gears, and remotely summoning the car for parking, can be applied. In the second-level trust zone, voice control commands requiring medium security, such as unlocking the doors, opening the trunk, and remotely turning on the air conditioning, can be applied. In the third-level trust zone, voice control commands requiring low security, such as flashing lights and honking the horn to locate the vehicle and checking the vehicle status (remaining battery power / tire pressure), can be applied.
[0140] Based on this, this embodiment can determine the instruction level corresponding to the audio waveform signal collected by the microphone array, and further determine the effective area of the instruction level. For example, if it is a level 1 instruction, the effective area is the level 1 trust zone; if it is a level 2 instruction, the effective area is the level 1 trust zone and the level 2 trust zone; if it is a level 3 instruction, the effective area is the level 1 trust zone, the level 2 trust zone and the level 3 trust zone.
[0141] Therefore, this embodiment can match the sound source localization result with the effective area. If the sound source localization result is located within the effective area, it is determined that the voice control command can be effective. Then, the vehicle can be voice controlled according to the audio waveform signal, that is, the vehicle can be voice controlled according to the voice control command.
[0142] Of course, there can be more refined control strategies. For example, after recognizing the multimodal characteristics of occupants or legally authorized persons in the vehicle (such as the various characteristics in the auxiliary verification factor determination process mentioned above), advanced scenarios such as intelligent partition authorization and child or pet modes can be implemented based on the multimodal characteristics. For example, only the driver-authenticated user is allowed to issue driving-related voice control commands, and the restricted control mode is automatically enabled for the area where children in the back seat or pets are detected, thereby ensuring vehicle safety and precise interaction.
[0143] In some possible implementations, a multi-layered safety strategy can be established. For example, the first level is real-time acoustic event detection, which involves high-decibel detection of ambient sounds during the same time period when the microphone array acquires the audio waveform signal of the voice control command. If a sudden high-decibel sound (such as screaming, impact, glass breaking, etc.) is detected, an emergency alarm is triggered, and appropriate measures are taken if safety conditions are met (such as obstacle detection, vehicle speed below a threshold, or forced parking mode). The second level is physiological audio feature analysis and intention recognition detection of the speaker. When the microphone array acquires the audio waveform signal of the voice control command, physiological audio feature analysis and intention recognition detection are performed on the audio waveform signal to obtain a confidence score representing whether the speaker of the voice control command has a physical abnormality. If the confidence score is higher than the first threshold (such as the driver groaning in pain), an emergency alarm is triggered, and appropriate measures are taken if safety conditions are met. If the confidence score is between the first and second thresholds (such as the driver's breathing disorder), appropriate measures are taken if safety conditions are met.
[0144] Of course, there are other classification strategies, such as weighted confidence fusion of confidence scores with information such as vehicle status (speed, gear, parking brake, door lock, lane / obstacle detection), vehicle location, user location, and number of anomalies detected, to obtain a fused confidence score, and then classifying responses based on the fused confidence score.
[0145] Optionally, the following actions may be taken, provided that safety conditions are met: triggering the parking brake or decelerating, or unlocking the doors within permissible limits.
[0146] Emergency alarms include, but are not limited to: uploading real-time audio and video and location information to the cloud via an encrypted channel, calling pre-set emergency contacts or rescue centers, and local alarms (such as flashing headlights, honking the horn, etc.).
[0147] This embodiment deeply integrates vehicle voice control with safety modules such as in-vehicle network, door lock electronic control, and parking brake to ensure full-process safety control and intelligent assistance for human-vehicle voice interaction, and comprehensively protect the user's life and property safety.
[0148] The above describes a vehicle voice control method provided by the embodiments of this application. The following will describe the apparatus for performing the above vehicle voice control method.
[0149] Please see Figure 2 , Figure 2 This is a schematic diagram of the structure of a vehicle voice control device provided in an embodiment of this application. Figure 2 As shown, the vehicle voice control device may include:
[0150] The data acquisition unit 201 is used to acquire the audio waveform signal collected by the microphone array on the vehicle in response to the voice control command, the sound source localization result of the voice control command, and the voiceprint similarity corresponding to the audio waveform signal;
[0151] The subsequence extraction unit 202 is used to determine multiple candidate subsequences from the target video sequence based on the sound source localization result. The candidate subsequences are composed of the region of interest of the same user in each video frame. The user is located within the preset neighborhood of the sound source localization result. The target video sequence is the in-vehicle and out-of-vehicle video sequence determined by the timestamp of the voice control command arriving at the microphone array.
[0152] The subsequence filtering unit 203 is used to perform cross-correlation calculations between multiple candidate subsequences and the audio waveform signal, so as to determine the candidate subsequence with the highest cross-correlation degree with the audio waveform signal from multiple candidate subsequences, and use it as the target subsequence;
[0153] The command control unit 204 is used to generate auxiliary verification factors based on the target sub-sequence, and determine whether the voice control command is issued by an authorized user in real time based on voiceprint similarity and auxiliary verification factors. If so, the vehicle is controlled by voice based on the audio waveform signal.
[0154] Each module in the aforementioned vehicle voice control device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0155] This application also provides an electronic device, which may include at least one processor and a memory connected to the processor, wherein:
[0156] Memory is used to store computer programs;
[0157] The processor is used to execute computer programs to enable electronic devices to implement any of the vehicle voice control methods provided in the embodiments of this application.
[0158] refer to Figure 3 The diagram illustrates a structural schematic suitable for implementing the electronic device in the embodiments of this application. The electronic device in the embodiments of this application may include, but is not limited to, fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 3 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0159] like Figure 3As shown, the electronic device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. When the electronic device is powered on, the RAM 603 also stores various programs and data required for the operation of the electronic device. The processing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0160] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, memory cards, hard drives, etc.; and communication devices 609. Communication device 609 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0161] This application also provides a computer program product including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the vehicle voice control methods provided in this application.
[0162] This application also provides a computer-readable storage medium carrying one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the vehicle voice control methods provided in this application.
[0163] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0164] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0165] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0166] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
Claims
1. A vehicle voice control method characterized by, include: The system acquires audio waveform signals collected by the microphone array on the vehicle in response to voice control commands, the sound source localization results of the voice control commands, and the voiceprint similarity corresponding to the audio waveform signals. Based on the sound source localization result, multiple candidate subsequences are determined from the target video sequence. The candidate subsequences consist of the region of interest of the same user in each video frame. The user is located within a preset neighborhood of the sound source localization result. The target video sequence is a video sequence inside and outside the vehicle determined by the timestamp of the voice control command arriving at the microphone array. The candidate subsequences are cross-correlated with the audio waveform signal respectively to determine the candidate subsequence with the highest cross-correlation with the audio waveform signal from the candidate subsequences, which is then used as the target subsequence; An auxiliary verification factor is generated based on the target sub-sequence. Based on the voiceprint similarity and the auxiliary verification factor, it is determined whether the voice control command is issued by an authorized user in real time. If so, the vehicle is controlled by voice based on the audio waveform signal.
2. The vehicle voice control method according to claim 1, characterized by, Before performing cross-correlation calculations between the plurality of candidate sub-sequences and the audio waveform signal, the method further includes: The candidate sub-sequences are matched with the registered face images to obtain the face matching degree corresponding to each candidate sub-sequence; From the plurality of candidate sub-sequences, determine a number of candidate sub-sequences whose face matching degree is higher than a preset matching degree threshold; The aforementioned candidate subsequences are used as multiple candidate subsequences in the cross-correlation calculation step.
3. The vehicle voice control method according to claim 2, characterized by, The step of using the plurality of candidate subsequences as multiple candidate subsequences in the cross-correlation calculation step includes: Facial micro-motion detection is performed on the candidate sub-sequences to determine whether there are candidate sub-sequences in which the user's facial expression is not moving. If so, the candidate subsequence with no facial movement of the user is removed from the plurality of candidate subsequences, and the remaining candidate subsequences are used as multiple candidate subsequences in the cross-correlation calculation step.
4. The vehicle voice control method according to any one of claims 1 to 3, characterized in that, The step of performing cross-correlation calculations between the plurality of candidate sub-sequences and the audio waveform signal includes: Optical flow calculations are performed on the multiple candidate sub-sequences to obtain the motion energy sequences corresponding to the multiple candidate sub-sequences; Determine the audio energy sequence of the audio waveform signal; The cross-correlation between the motion energy sequences corresponding to the multiple candidate sub-sequences and the audio energy sequence is calculated, and is used as the cross-correlation between the multiple candidate sub-sequences and the audio waveform signal.
5. The vehicle voice control method according to claim 1, characterized by, The step of generating auxiliary verification factors based on the target subsequence includes: Anti-fraud detection is performed based on the audio waveform signal and the target sub-sequence to obtain an anti-fraud score; Determine the positional matching degree between the user's actual location in the target subsequence and the sound source localization result; Obtain the time decay memory value of the user's successful authentication within the most recent preset time period in the target subsequence; The target subsequence is matched with the registered face image to obtain the face matching degree corresponding to the target subsequence; At least one of the anti-fraud score, the location matching degree, the time decay memory value, and the face matching degree corresponding to the target subsequence is used as the auxiliary verification factor.
6. The vehicle voice control method according to claim 5, characterized by, The anti-fraud detection based on the audio waveform signal and the target sub-sequence, to obtain an anti-fraud score, includes: Determine the audio spectrum characteristics of the audio waveform signal, and determine a first confidence probability that the audio waveform signal is a waveform signal played by the terminal device based on the audio spectrum characteristics; Facial micro-movement detection is performed on the target sub-sequence to determine the second confidence probability that the user's facial movements are not in the target sub-sequence. The anti-fraud score is generated based on the first confidence probability and the second confidence probability.
7. The vehicle voice control method according to claim 1, characterized by, The step of determining whether the voice control command is issued in real time by an authorized user based on the voiceprint similarity and the auxiliary verification factor includes: The voiceprint similarity and the auxiliary verification factor are concatenated into an input feature vector, which is then input into a pre-trained sound source reliability prediction model to obtain a reliability prediction probability value. The sound source reliability prediction model is trained using a training input feature vector labeled with a reliability prediction probability value. If the reliability prediction probability value is greater than or equal to the preset first probability threshold, then it is determined that the voice control command is issued in real time by the authorized user corresponding to the registered face image; If the reliability prediction probability value is less than the preset second probability threshold, it is determined that the voice control command is not issued in real time by the authorized user corresponding to the registered face image, and the second probability threshold is less than the first probability threshold. If the reliability prediction probability value is less than the first probability threshold and greater than or equal to the second probability threshold, the authorized user is notified to repeat the voice control command in order to initiate secondary verification based on the repeated voice control command.
8. The vehicle voice control method according to claim 1, characterized by, The step of performing voice control on the vehicle based on the audio waveform signal includes: Determine the instruction level corresponding to the audio waveform signal, and determine the effective area of the instruction level; If the sound source localization result is located within the effective area, then the vehicle is controlled by voice based on the audio waveform signal.
9. The vehicle voice control method according to claim 1, characterized by, The process of determining the sound source localization result includes: The timestamps of the voice control commands arriving at each microphone in the microphone array are obtained to obtain the audio arrival timestamps corresponding to the multiple microphones in the microphone array. Based on the position coordinates of each of the multiple microphones, the audio arrival timestamps corresponding to the multiple microphones, and the sound source location, multiple time delay difference equations are established. The multiple time delay difference equations are solved simultaneously to obtain the solution value of the sound source location, which is used as the sound source localization result.
10. A vehicle voice control device characterized by comprising: include: The data acquisition unit is used to acquire audio waveform signals collected by the microphone array on the vehicle in response to voice control commands, the sound source localization results of the voice control commands, and the voiceprint similarity corresponding to the audio waveform signals. The subsequence extraction unit is used to determine multiple candidate subsequences from the target video sequence based on the sound source localization result. The candidate subsequences are composed of the region of interest of the same user in each video frame. The user is located within a preset neighborhood of the sound source localization result. The target video sequence is a video sequence inside and outside the vehicle determined by the timestamp of the voice control command arriving at the microphone array. The subsequence filtering unit is used to perform cross-correlation calculations between the plurality of candidate subsequences and the audio waveform signal respectively, so as to determine the candidate subsequence with the highest cross-correlation degree with the audio waveform signal from the plurality of candidate subsequences, and use it as the target subsequence; The instruction control unit is used to generate an auxiliary verification factor based on the target sub-sequence, and determine whether the voice control command is issued by an authorized user in real time based on the voiceprint similarity and the auxiliary verification factor. If so, the unit performs voice control on the vehicle based on the audio waveform signal.