A vehicle control method and device based on multi-modal information, electronic equipment, and storage medium
Patent Information
- Application Number
- CN202511132864.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2045-08-13
AI Technical Summary
[0003]然而现有系统过度依赖预置词库导致方言或自定义语音,且无法支持咳嗽声、口哨等非标准指令,迫使用户被动适应机器语法,对于用户来说显得自主性不强;难以基于实际的车辆运行环境及用户反馈进行智能化决策,导致智能程度及用户驾驶体验受限
[0021] The embodiments of the present invention bring the following beneficial effects: The vehicle control method, device, electronic device, and storage medium based on multimodal information provided in this application are applied to the vehicle terminal; the method integrates the user's biometric data and voiceprint data into a voiceprint key, breaks through the limitation of the preset dictionary, realizes the binding of any voice segment with vehicle commands, and then integrates it with vehicle environment data, current biometric data, and vehicle status data to drive command fission in real time, so as to generate a target control command set containing a first control command that is more in line with user habits and environmental changes, so as to intelligently control the vehicle and improve the driving experience.
Smart Images

Figure CN120998192B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of vehicle control technology, and in particular to a vehicle control method, device, electronic device, and storage medium based on multimodal information. Background Technology
[0002] After more than two decades of development, in-vehicle voice recognition technology has achieved efficient execution of basic functions such as volume adjustment, answering calls, and switching media. With the popularization of software-defined vehicles (SDV) and autonomous driving technology, voice interaction is gradually becoming the core entry point for smart cockpits. Current mainstream solutions are based on pre-built dictionaries and fixed grammatical frameworks, completing command responses through a five-level processing chain of wake-up, recognition, semantic parsing, security verification, and execution.
[0003] However, existing systems rely too heavily on pre-built dictionaries, resulting in dialects or custom voices, and cannot support non-standard commands such as coughs and whistles. This forces users to passively adapt to machine grammar, making them feel less autonomous. Furthermore, the systems struggle to make intelligent decisions based on the actual vehicle operating environment and user feedback, thus limiting their level of intelligence and the user's driving experience. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a vehicle control method, device, electronic device, and storage medium based on multimodal information.
[0005] In a first aspect, embodiments of the present invention provide a vehicle control method based on multimodal information, applied to an in-vehicle infotainment system; the method includes: Feature extraction was performed on the raw voiceprint data of passengers to obtain the acoustic features of Mel frequency cepstral coefficients; Based on the acoustic characteristics of Mel frequency cepstral coefficients and the biometric data of passengers, a dynamic voiceprint key and the initial control command corresponding to the voiceprint key are obtained. By collecting vehicle environment data, vehicle status data, and current biometric data, as well as the voiceprint key, the initial control command is subjected to command fission processing to obtain the target command set; the target command set includes at least one first control command. Control the vehicle based on the first control command.
[0006] Combining the first aspect, the steps for obtaining a dynamic voiceprint key and the initial control command corresponding to the voiceprint key, based on the acoustic characteristics of the Mel frequency cepstral coefficients and the biometric data of the passengers, include: Based on pre-configured weighted information, the acoustic features of Mel frequency cepstral coefficients are fused with biological data to generate intermediate entropy values of a specified number of bits. Based on pre-configured encrypted information, the intermediate entropy value is dynamically encrypted to generate a voiceprint key; Based on the voiceprint key retrieval instruction mapping table, the initial control instruction corresponding to the voiceprint key is obtained.
[0007] In conjunction with the first aspect, the steps of obtaining the target instruction set by performing instruction fission processing on the initial control command through the collected vehicle environmental data, vehicle status data, current biometric data, and voiceprint key include: Vehicle environmental data, vehicle status data, current biometric data, and voiceprint key are input into a preset command generation model to perform command fission, resulting in a fission set of control commands. Collect biofeedback information from passengers when the vehicle is running with the control command set after fission; A biofeedback information-based instruction generation model is used to generate a target instruction set, which includes at least one first control instruction.
[0008] In conjunction with the first aspect, after the step of inputting vehicle environmental data, vehicle status data, current biometric data, and voiceprint key into a preset instruction generation model to perform instruction fission and obtain the fissioned control instruction set, the method further includes: For each control instruction in the fission control instruction set, the target actuator is determined based on the fission control instruction; The control commands generated from the fission are sent to the target actuator to control the target actuator to run according to the fission commands.
[0009] In conjunction with the first aspect, the vehicle-mounted system communicates with the mobile terminal; after inputting vehicle environmental data, vehicle status data, current biometric data, and voiceprint key into a preset instruction generation model to perform instruction fission and obtain the fissioned control instruction set, the system further includes: Based on a predefined conflict rule base, conflict detection is performed on each of the split control instructions in the split control instruction set to obtain the detection results; If the detection result indicates the presence of a conflicting command, record the conflicting command and / or send a feedback message to the mobile terminal.
[0010] In addition to the first aspect, the vehicle-mounted system also communicates and connects with mobile terminals and home devices; By collecting vehicle environment data, vehicle status data, and current biological data, as well as voiceprint keys, the initial control command is processed into a target command set through command fission. At the same time, the initial control command is input into the edge computing node and combined with the status data of mobile terminals and home devices to generate a terminal control command set in real time. The vehicle is controlled based on the first control command, and at the same time, terminal control commands in the terminal control command set are distributed to the corresponding mobile terminal or home device through a quantum encrypted channel.
[0011] In conjunction with the first aspect, the vehicle-mounted system also communicates with mobile terminals and home devices; the initial control commands include at least one vehicle control command and at least one initial terminal control command; the target control command set also includes at least one terminal control command. After obtaining the dynamic voiceprint key and the initial control command corresponding to the voiceprint key based on the acoustic characteristics of the Mel frequency cepstral coefficients and the biometric data of the passengers, the process also includes: By collecting vehicle environment data, vehicle status data, and current biometric data, as well as voiceprint keys, the initial control command is subjected to command fission processing to obtain a target command set; the target command set includes at least one first control command obtained based on vehicle control command fission, and at least one terminal control command obtained based on initial terminal control command fission. The vehicle is controlled based on the first control command, and at the same time, the terminal control command is distributed to the corresponding mobile terminal or home device through a quantum-encrypted channel.
[0012] Secondly, this application provides a vehicle control device based on multimodal information, applied to an in-vehicle infotainment system; the device includes: The feature extraction module is used to extract features from the raw voiceprint data of passengers to obtain the acoustic features of Mel frequency cepstral coefficients. The generation module is used to obtain a dynamic voiceprint key and the initial control command corresponding to the voiceprint key based on the acoustic features of the Mel frequency cepstral coefficients and the biometric data of the passengers. The instruction fission module is used to perform instruction fission processing on the initial control instruction by collecting vehicle environment data, vehicle status data, and current biological data, as well as the voiceprint key, to obtain a target instruction set; the target instruction set includes at least one first control instruction. The control module is used to control the vehicle based on the first control command.
[0013] In conjunction with the second aspect, the generation module includes: The intermediate entropy generation module is used to generate an intermediate entropy value of a specified number of bits based on pre-configured weighted information, fusion of Mel frequency cepstral coefficient acoustic features and biological data; The voiceprint key generation module is used to dynamically encrypt the intermediate entropy value based on pre-configured encryption information to generate a voiceprint key. The initial control command generation module is used to retrieve the command mapping table based on the voiceprint key to obtain the initial control command corresponding to the voiceprint key.
[0014] In conjunction with the second aspect, the instruction fission module includes: The fission module is used to input vehicle environmental data, vehicle status data, current biological data and voiceprint key into a preset instruction generation model to perform instruction fission and obtain a fission set of control instructions. The feedback module is used to collect the biofeedback information of passengers when the vehicle is running with the control command set after fission. An update module is used to update the instruction generation model based on biofeedback information to generate a target instruction set, which includes at least one first control instruction.
[0015] In conjunction with the second aspect, following the fission module, it also includes: The determination module is used to determine the target actuator based on each fission control instruction in the fission control instruction set. The sending module is used to send the split control commands to the target actuator so that the target actuator can run according to the split control commands.
[0016] In conjunction with the second aspect, the vehicle-mounted system communicates with the mobile terminal; following the fission module, it also includes: The detection module is used to perform conflict detection on each of the split control instructions in the split control instruction set based on a predefined conflict rule base, and obtain the detection results. The recording module is used to record the conflicting command and / or send the conflicting command back to the mobile terminal if the detection result indicates that there is a conflicting command.
[0017] In conjunction with the second aspect, the vehicle-mounted system also communicates with mobile terminals and home devices; the device also includes: The computing module is used to process the initial control command by collecting vehicle environment data, vehicle status data, current biological data, and voiceprint key to obtain the target command set. At the same time, it inputs the initial control command into the edge computing node and combines the status data of the mobile terminal and home device to generate the terminal control command set in real time. The encrypted distribution module is used to control the vehicle based on the first control command, and at the same time, distribute the terminal control commands in the terminal control command set to the corresponding mobile terminal or home device through the quantum encrypted channel.
[0018] In conjunction with the second aspect, the vehicle-mounted system also communicates with mobile terminals and home devices; the initial control commands include at least one vehicle control command and at least one initial terminal control command; the target control command set also includes at least one terminal control command. After generating the module, it also includes: The second fission module is used to perform instruction fission processing on the initial control command by collecting vehicle environment data, vehicle status data, and current biological data, as well as voiceprint key, to obtain a target instruction set; the target instruction set includes at least one first control command obtained based on vehicle control command fission, and at least one terminal control command obtained based on initial terminal control command fission. The terminal control command distribution module is used to control the vehicle based on the first control command, and at the same time, distribute the terminal control command to the corresponding mobile terminal or home device through a quantum-encrypted channel.
[0019] Thirdly, this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor runs the computer program to cause the electronic device to perform the above-described method.
[0020] Fourthly, this application provides a storage medium storing computer program instructions, which are read and executed by a processor to perform the above-described method.
[0021] The embodiments of the present invention bring the following beneficial effects: The vehicle control method, device, electronic device, and storage medium based on multimodal information provided in this application are applied to the vehicle terminal; the method integrates the user's biometric data and voiceprint data into a voiceprint key, breaks through the limitation of the preset dictionary, realizes the binding of any voice segment with vehicle commands, and then integrates it with vehicle environment data, current biometric data, and vehicle status data to drive command fission in real time, so as to generate a target control command set containing a first control command that is more in line with user habits and environmental changes, so as to intelligently control the vehicle and improve the driving experience.
[0022] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.
[0023] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0024] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0025] Figure 1 This is a flowchart illustrating the vehicle control method based on multimodal information provided in Embodiment 1 of the present invention. Figure 2 This is a flowchart illustrating the vehicle control method based on multimodal information provided in Embodiment 2 of the present invention. Figure 3A flowchart illustrating the vehicle control method based on multimodal information provided in this embodiment of the invention; Figure 4 A flowchart illustrating the vehicle control method based on multimodal information provided in this embodiment of the invention; Figure 5 A flowchart illustrating the vehicle control method based on multimodal information provided in this embodiment of the invention; Figure 6 A schematic diagram of the structure of a vehicle control device based on multimodal information provided in this embodiment of the invention; Figure 7 This is a schematic diagram of the electronic device structure provided in an embodiment of the present invention.
[0026] Figure label: 10 - Feature extraction module, 20 - Generation module, 30 - Instruction fission module, 40 - Control module; 130 - Processor, 131 - Memory, 132 - Bus, 133 - Communication interface. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] To facilitate understanding of this embodiment, the application scenarios and design concepts of this application embodiment will be briefly introduced below.
[0029] Existing vehicle control systems rely too heavily on pre-built dictionaries, making it difficult to effectively control vehicles using dialects, custom voice commands, and non-standard instructions. Using voiceprint features to control vehicles is also challenging, limiting the level of intelligence and the user's driving experience.
[0030] Based on this, this application provides a vehicle control method, device, electronic device, and storage medium based on multimodal information. This method is applied to the vehicle's infotainment system, which is typically embedded or externally mounted within the vehicle, or communicates with the vehicle system via in-vehicle Ethernet / CAN FD to meet low-latency requirements. This application extracts features from collected voiceprint data and fuses them with biometric data to generate initial control commands, thereby overcoming the limitations of pre-set word libraries and enabling the binding of arbitrary voice segments with vehicle commands. Subsequently, it integrates multi-dimensional data to perform command fission for intelligent vehicle control, enhancing the driving experience.
[0031] Example 1 This application provides a vehicle control method based on multimodal information, combined with Figure 1 As shown, the method includes: S110 extracts features from the raw voiceprint data of passengers to obtain the acoustic features of Mel frequency cepstral coefficients.
[0032] S120, based on the acoustic characteristics of Mel frequency cepstral coefficients and the biometric data of passengers, obtains a dynamic voiceprint key and the initial control command corresponding to the voiceprint key.
[0033] S130: By collecting vehicle environment data, vehicle status data, and current biological data, as well as the voiceprint key, the initial control command is subjected to command fission processing to obtain a target command set, which includes at least one first control command.
[0034] S140 controls the vehicle based on the first control command.
[0035] The method provided in this application extracts features from the raw voiceprint data of passengers, and then converts the extracted Mel-frequency cepstral coefficient (MFCC) acoustic features and biological data into a voiceprint key to determine its association with the initial control command. This breaks through the limitations of the preset dictionary and enables the binding of any voice segment with vehicle commands. Subsequently, the method drives command fission in real time through environmental sensing data, user biological data, and vehicle status data to generate a first control command that is more in line with user habits and environmental changes, thereby improving the driving experience and the intelligence of vehicle driving.
[0036] The step of extracting the Mel-frequency cepstral coefficients (MFCC) acoustic features from the original speaker data in step S110 specifically includes: S111, the original voiceprint data is sequentially processed by noise reduction and pre-emphasis to obtain the speech stream.
[0037] S112 divides the continuous speech stream into multiple short segments.
[0038] S113. For each short time segment, apply a window function to the short time segment and then perform a Fourier transform to obtain a linear spectrum.
[0039] S114 inputs the linear spectrum into the Mel filter bank and outputs the energy value of each filter.
[0040] S115 calculates the logarithm of the energy of each filter and converts it to the cepstral domain to output the Mel frequency cepstral coefficients of a specified number of dimensions, thus obtaining the acoustic characteristics of the Mel frequency cepstral coefficients.
[0041] In this embodiment, the raw voiceprint data (usually a time-series audio waveform sample value) of passengers (including but not limited to the driver, front passenger, rear passenger and temporary visitor) is first collected by a noise-canceling microphone array.
[0042] Then, step S111 is executed to eliminate or reduce unavoidable background noise in the recording environment (such as engine noise, wind noise, air conditioning noise, other passenger noise, etc.) through noise reduction processing, improve the signal-to-noise ratio (SNR) of the speech signal, and then perform pre-emphasis processing to compensate for the attenuation of high-frequency components of the speech signal during propagation (mainly caused by glottal pulse and lip radiation effect), enhance the energy of the high-frequency part, and make the entire spectrum flatter, which is convenient for subsequent processing.
[0043] Common noise reduction methods include spectral subtraction, Wiener filtering, and deep learning-based noise reduction models. Spectral subtraction specifically involves estimating the spectrum of noise (usually the silent segment before the start of speech) and then subtracting the estimated noise spectrum from the spectrum of the noisy speech. Wiener filtering specifically involves estimating the clean speech optimally with the least mean square error based on a statistical model. Deep learning-based noise reduction models specifically involve using a trained neural network to directly map noisy speech to clean speech (usually with better results, but with greater computational cost).
[0044] Pre-emphasis processing specifically uses a first-order high-pass filter, most commonly by applying a simple difference equation:
[0045] in, This is the current sampling point. It is the previous sampling point. It is the pre-emphasis coefficient (usually between 0.95 and 0.97, such as 0.97).
[0046] Understandably, the speech stream obtained after noise reduction and pre-emphasis has less noise (i.e., is "cleaner") and more prominent high-frequency information.
[0047] Among them, the noise-canceling microphone array is used to accurately collect occupant voices in the complex noise background of the vehicle environment. It is the hardware foundation of the voiceprint control system. Voiceprint data is collected by setting it in different positions to cover different seating areas. For example, it is embedded in the steering wheel to collect voiceprint data of the driver's area; it is set in the center of the roof to collect voiceprint data for full vehicle coverage; it is set in the seat headrest to collect voiceprint data of the rear seat area in a directional manner; and it is set in the center console to collect voiceprint data of the front passengers.
[0048] Subsequently, step S112, framing, is performed. Since the speech signal is non-stationary (its statistical characteristics change over time), in order to utilize stationary signal processing techniques (such as Fourier transform), the continuous speech stream needs to be divided into many short time segments (frames). Within each frame (approximately 20-40 milliseconds), the speech signal can be approximated as quasi-stationary (i.e., its statistical characteristics do not change significantly). Specifically, after setting the frame length and frame shift (HopSize), framing is performed, outputting a series of short time segments (i.e., frames).
[0049] Frame shift (or hop size) refers to the interval between the start points of two adjacent frames. It's important to note that the frame shift should be less than the frame length to ensure overlap between frames, maintain information continuity, and avoid spectral distortion caused by truncation. Frame length is typically 20ms, 25ms, or 30ms; 50% (10ms) or 25% (5ms) of the frame length are used as examples here and are not considered definitive.
[0050] Subsequently, step S113 is executed to perform windowing and Fast Fourier Transform (FFT), outputting the linear spectrum (usually amplitude spectrum or power spectrum) of the speech frame.
[0051] Performing an FFT directly on a short-segment signal is equivalent to truncating an infinitely long period signal (rectangular window), which will cause spectral leakage, that is, energy diffuses to the neighboring frequency components, blurring the true spectral peaks. The windowing function is used to reduce this truncation effect.
[0052] Specifically, the short segment is multiplied by a window function. Commonly used window functions include the Hamming window and the Hanning window.
[0053] Subsequently, a Fast Fourier Transform (FFT) is performed to convert the short-time segment signal in the time domain to the frequency domain, obtaining its linear spectrum. A complex array is output to represent the amplitude and phase information of the short-time segment at each discrete frequency point, thus obtaining the linear spectrum (usually the amplitude spectrum or power spectrum) of the short-time segment.
[0054] Subsequently, step S114 is executed to simulate the nonlinear perception characteristics of human ears regarding sound frequencies (human ears are sensitive to low-frequency differences but insensitive to high-frequency differences), and the linear frequency scale is mapped to the Mel scale, which better reflects auditory characteristics. This also serves to reduce dimensionality and decorrelate. Specifically: the Mel scale is defined; a set of triangular bandpass filters is designed on the Mel scale; the frequency range of interest is determined and this range is divided equally on the Mel scale. Then, each Mel band is converted back to a linear frequency. Next, a triangular filter is designed around each center frequency on the linear frequency axis (where the filter response is 1 at the center frequency, linearly decreasing to 0 at the center frequency of the adjacent filter, and adjacent filters overlap). Finally, the linear spectrum (power spectrum) obtained in step S113 is passed through this set of Mel filters. For each filter k, the weighted sum of the spectral value (power value) multiplied by the filter response value within its covered frequency range is calculated to obtain the energy value output by each Mel filter.
[0055] Next, step S115 takes the logarithm of the energy value output by the Mel filter to simulate the nonlinear perception of loudness by the human ear (based on the Weber-Fechner law, where perceived intensity is proportional to the logarithm of physical intensity). Simultaneously, the logarithmic operation compresses the dynamic range of the data, making the features more robust to changes in input volume (amplifying weak signals and compressing strong signals). Subsequently, a Discrete Cosine Transform (DCT) is performed. DCT is an orthogonal transformation that converts this set of correlated energy values into a set of independent (or weakly correlated) coefficients. Typically, only the first L coefficients (e.g., 12 or 13) are needed to well describe the envelope characteristics of the spectrum (primarily reflecting vocal tract characteristics, related to the speaker), while ignoring detail characteristics (primarily reflecting excitation source characteristics, such as fundamental frequency / pitch, related to the speech content). The resulting L-dimensional vector is the Mel frequency cepstral coefficient (MFCC) acoustic feature of this short time segment.
[0056] Therefore, step S110 ensures that effective raw voiceprint data is collected in complex vehicle environments through directional sound pickup and multimodal noise reduction using an anti-noise microphone array. It does not rely on fixed "wake-up words" or "command words," breaking through the traditional constraints of word library limitations. Using any speech segment as the trigger source, after feature extraction, the final output is a low-dimensional, information-dense, and discriminative sequence. This sequence represents the MFCC acoustic characteristics of the speech segment and can more efficiently characterize the essential attributes of the voiceprint. It provides a high-quality input foundation for the subsequent generation of "dynamic voiceprint key generation" and "command fission."
[0057] In conjunction with the first aspect, after step S115, the following also includes: S116, add dynamic difference coefficients to the Mel frequency cepstral coefficients of a specified dimension to form an enhanced feature vector of the first quantitative dimension.
[0058] The dynamic difference coefficients include first-order difference coefficients and second-order difference coefficients; the first quantity is greater than the specified quantity.
[0059] By adding dynamic discrimination coefficients, the shortcomings of static MFCC features in representing temporal dynamic information are addressed. By capturing the changing trend (speed) and acceleration (acceleration) of features between adjacent frames, the voiceprint features more completely reflect the dynamic process of pronunciation (such as tone fluctuations, syllable boundaries, plosive characteristics, etc.). First-order difference coefficients are added to characterize the rate of change of MFCC features between adjacent frames, and second-order discrimination coefficients are further added to characterize the rate of change of the first-order difference coefficients (approximately the second derivative) to reflect the acceleration changes of pronunciation. This expands the feature dimension and improves the robustness of voiceprint recognition under complex vehicle noise conditions, key security, and accuracy of instruction intent parsing.
[0060] Subsequently, step S120 combines the MFCC feature (voiceprint) with real-time biometric data (such as heart rate, fingerprint, facial features, infrared lip movement trajectory, etc.) to generate a dynamic voiceprint key using a specific algorithm (which may include hashing, encryption, feature fusion, etc.). This voiceprint key is an encrypted representation of the user's identity and current physiological state; in this embodiment, the voiceprint key is a 128-bit key. This voiceprint key is mapped to an initial control command; this mapping may be pre-learned (user-personalized configuration) or dynamically generated. The initial command represents the user's core intent (such as "navigate home" or "lower the temperature"). It is worth noting that in this embodiment, the original voiceprint data is destroyed simultaneously with the generation of the dynamic voiceprint key in step S120 to avoid privacy leaks.
[0061] Subsequently, step S130 executes multimodal command fission, using real-time environment, vehicle status, and user status information to intelligently refine and expand the initial command, resulting in a target command set containing at least one first control command. This achieves the dynamic and intelligent transformation (fission) of static, voiceprint-based authentication initial commands into a series of control commands suitable for the specific current situation and capable of safe and effective execution, utilizing rich real-time contextual information (vehicle status data, vehicle environment data, and current biometric data). This represents a significant evolution in vehicle control from "simple command execution" to "intelligent situational understanding and decision-making." Thus, it is no longer limited to executing fixed preset commands but can flexibly adapt to different drivers (identified through voiceprint keys), different driving habits (potentially implicit in biometric data), and different road and weather conditions, providing personalized and optimal control strategies. This makes the vehicle control system more universal, user-friendly, and robust in responding to changing driving environments.
[0062] Finally, the vehicle is controlled based on the first control command obtained in step S130, and the vehicle is regulated to perform corresponding operations (such as adjusting the air conditioning, changing the navigation route, opening and closing the windows, adjusting the seats, etc.) based on the original voiceprint data of the passengers.
[0063] Understandably, after step S110, there is also a voiceprint verification process. If the voiceprint verification fails three times in a row, the local voiceprint key will be destroyed immediately to prevent privacy leakage. At this time, the user needs to re-register the voiceprint account.
[0064] Example 2 This application provides another vehicle control method based on multimodal information, such as... Figure 2 As shown, the method specifically includes: S210 extracts features from the raw voiceprint data of passengers to obtain the acoustic features of Mel frequency cepstral coefficients.
[0065] S220, based on pre-configured weighted information, integrates the acoustic features of Mel frequency cepstral coefficients with biological data to generate an intermediate entropy value with a specified number of bits.
[0066] S230 dynamically encrypts the intermediate entropy value based on pre-configured encryption information to generate a voiceprint key.
[0067] S240: Based on the voiceprint key retrieval instruction mapping table, obtain the initial control instruction corresponding to the voiceprint key.
[0068] S250 uses the collected vehicle environment data, vehicle status data, and current biological data, along with the voiceprint key, to perform instruction fission processing on the initial control command to obtain a target instruction set, which includes at least one first control command.
[0069] S260 controls the vehicle based on the first control command.
[0070] In this embodiment, after extracting MFCC features from the original voiceprint data in step S210, the MFCC features are fused with biological data in step S220. The biological data can refer to heart rate, fingerprint hash, facial features, or other biological data. In this process, it is preferable to first perform dimensional feature alignment, that is, reduce or increase the dimensionality of the biological data to the same dimension as the MFCC features. Then, based on a pre-configured weighting matrix, each element (referring to the biological data and the MFCC features of each dimension) is weighted and summed to generate an intermediate entropy value with a specified number of bits, thereby maximizing information uncertainty and enhancing key unpredictability.
[0071] Subsequently, after initializing dynamic parameters (such as timestamp and vehicle state entropy) and constructing the encryption seed in step S230, a voiceprint key is generated based on the preset encryption information (i.e., dynamic encryption algorithm). It is understandable that since the timestamp and vehicle state entropy are different at each moment, the voiceprint key generated at each moment is different and has uniqueness.
[0072] Understandably, the vehicle's infotainment system pre-stores a mapping table between voiceprint keys and control commands. After obtaining the voiceprint key in step S230, step S240 is executed. By retrieving the mapping relationship in the mapping table, the control command corresponding to the voiceprint key can be obtained, thus obtaining the initial control command.
[0073] The retrieval process is as follows: the user identity is matched through the key hash table, then a personalized instruction set is loaded based on each prefix of the voiceprint key, and then the output is performed according to the preset format to obtain the initial control instructions.
[0074] For example, identity matching is performed based on the voiceprint key hash table "HashTable[K_voice]", which determines "User_007"; then, the mapping relationship is found based on the prefix "A2F3...B0" of the voiceprint key, and the instruction 1 is determined to be: navigation "home". At the same time, the mapping relationship is found based on the prefix "B1D3...C2" of the voiceprint key, and the instruction 2 is determined to be: air conditioner "23℃"... until all the instructions corresponding to the prefixes are found, resulting in n (n≥1) instructions. These n instructions are converted and arranged according to a preset format to obtain the initial control instructions.
[0075] Understandably, the in-vehicle infotainment system also includes a pre-trained command prediction model. After generating a dynamic voiceprint key, it determines the user's identity and constructs a user behavior knowledge graph based on the user's historical operation logs and vehicle state entropy (acceleration mode / steering force / air conditioning setting preferences). Subsequently, the voiceprint key, real-time biometric data (such as heart rate, eye muscle features, etc.), and vehicle environmental data (GPS location, weather, road conditions, traffic information, etc.) are input into the command prediction model to predict commands. For example, if the original voiceprint data is a yawn sound, after feature extraction, it is fused with biometric data (increased heart rate) to generate a dynamic voiceprint key. Based on the command mapping table, the command "navigate to nearby parking lot" is obtained. At the same time, combined with vehicle environmental data (such as the vehicle being in a high-speed driving state), the command prediction model is input to output a predicted command to start a massage.
[0076] Then, steps S250-S260 are executed to perform command splitting and vehicle control in order to perform the corresponding operations.
[0077] In this embodiment, to achieve high-precision binding of voiceprint features with vehicle control commands, a real-time constraint is adopted, meaning the fusion of voiceprint features and biological data must be completed within 20ms. The acquisition of biological data relies on various sensors, such as a PPG optical heart rate sensor, which can monitor the heart rate of passengers in real time. Simultaneously, the PPG heart rate sensor is linked with a steering wheel grip force sensor, using capacitive pressure signals instead on bumpy roads; an infrared lip movement tracking camera captures lip movement trajectories to achieve synchronous verification with voiceprints to prevent recording spoofing; and a capacitive pressure sensor can also be used to detect seat pressure distribution, enabling multi-seat voiceprint binding, but this is only an example and not a limitation. The biological data acquired by the above sensors is integrated with the MFCC feature input into an ASIL-D compliant hardware encryption module to generate a 128-bit voiceprint gene key in real time and perform a destructive process on the original voiceprint data to enhance data security.
[0078] Example 3 This application provides another vehicle control method based on multimodal information, such as... Figure 3 As shown, the method includes: S310 extracts features from the raw voiceprint data of passengers to obtain the acoustic features of Mel frequency cepstral coefficients.
[0079] S320, based on the acoustic characteristics of Mel frequency cepstral coefficients and the biometric data of passengers, obtains a dynamic voiceprint key and the initial control command corresponding to the voiceprint key.
[0080] S330 inputs vehicle environmental data, vehicle status data, current biometric data and voiceprint key into a preset instruction generation model to perform instruction fission, and obtains the fissioned control instruction set.
[0081] S340 collects user biofeedback information when the vehicle is running with the control command set after fission.
[0082] S350 updates the instruction generation model based on user biofeedback information to generate a target instruction set, which includes at least one first control instruction.
[0083] S360 controls the vehicle based on the first control command.
[0084] The method provided in this embodiment differs from that in Embodiment 1.2 in that: after obtaining the initial control command, a nanosecond-level decision engine (TriLoS decision engine equipped with an FPGA accelerator) performs fission processing on the initial control command. The initial control command is used as the "root node," and a local command generation model fuses vehicle environmental data, vehicle status data, current biometric data, and voiceprint keys to assess the current situation and potential user needs. This triggers preset "fission rules," dynamically correcting or expanding the initial control command to generate a control command tree containing multiple sub-control commands (i.e., the "fissioned control command set"). For example, if the initial control command determined based on the voiceprint key is "navigate home," and the collected vehicle environmental data ("rush hour traffic congestion"), vehicle status data ("low battery"), and current biometric data ("driver fatigue") are fused with the voiceprint key, the fissioned control command set is obtained, including multiple fissioned commands: "recommend a faster detour" or "suggest delaying departure," "navigate to the nearest charging station," "activate seat massage," or "suggest parking at the next rest area." For example, if the vehicle's environmental data is "raining", the vehicle's status data is "rear window defroster closed", and the current biological data is "frowning", then the extended control command "turn on rear window defroster" is output. Combining the extended control command and the initial control command, the fission control command set is obtained.
[0085] Subsequently, the vehicle is controlled using the resulting control command set, and the biofeedback information of the passengers is collected. This biofeedback information is then used to update the local command generation model, generating a target command set that better meets the passengers' needs for intelligent vehicle control. Using the example above, after generating the "turn on seat massage" command, the biofeedback information of the passengers when the seat massage is turned on is collected for adjustment. For example, if the collected biofeedback information is "frowning, increased frequency of eye muscle twitching," indicating that the seat massage increases the passengers' discomfort, then the discomfort can be reduced by decreasing the massage intensity, switching to a soothing massage mode, or even turning off the seat massage. This is merely an example and not a limitation.
[0086] This embodiment uses multimodal sensors to collect real-time biological data (such as facial expressions and gestures) and vehicle environmental data. Combined with vehicle status data, a local lightweight federated learning model (i.e., the aforementioned instruction learning model) dynamically corrects instruction logic at a speed of 20ms, forming a closed loop of "perception-decision-execution-evolution". This allows the vehicle to understand subtext. For example, when the user yawns at night, it suggests navigation to a rest stop; when the user frequently clears their throat, it increases the air conditioning humidity. Another example is when it detects that the user frowns three times in a row while playing rock music, it automatically switches the default music in "Adventure Mode" to another style of music. At the same time, it integrates vehicle environmental factors (such as adding the action of turning on fog lights in rainy weather), achieving "online evolution without OTA updates". The accuracy is greatly improved, and all learning data is processed in a closed loop on the vehicle, avoiding the risk of privacy leakage. It truly makes the car a digital twin that understands subtext and gets better and better with use.
[0087] Furthermore, the split control instruction set includes at least one first control instruction for controlling the vehicle.
[0088] After step S330, the following is also included: S331, for each control instruction in the control instruction set after fission, determine the target actuator based on the control instruction after fission.
[0089] S332 sends the split control commands to the target actuator to control the target actuator to run according to the split control commands.
[0090] Understandably, there are multiple actuators in a vehicle, each with different functions. The target actuator is determined based on the control commands after fission. For example, the target actuator for "raising the interior temperature" is "air conditioning", and the target actuator for "air circulation" is "window" or "internal / external circulation device".
[0091] Furthermore, in this embodiment, commands are transmitted in real time via a CAN bus (transmission rate greater than 2Mbps, typically 5Mbps) to drive the target actuator in a direct-pass manner, effectively shortening latency and improving command response speed. Testing and monitoring have shown that this method can further compress the response time from 520ms to less than 182ms, achieving a 93% availability rate in high-speed scenarios. Simultaneously, the CAN bus direct-write engine breaks through the traditional five-layer processing chain (wake-up, voice recognition, semantic parsing, identity verification, and command execution), enhancing security levels.
[0092] In conjunction with the first aspect, the vehicle-mounted system establishes a communication connection with the mobile device. Following step S330, the following steps are also included: S333, based on a predefined conflict rule base, performs conflict detection on each of the split control instructions in the split control instruction set and obtains the detection results.
[0093] S334, if the detection result indicates that there is a conflicting command, the conflicting command is fed back to the vehicle terminal and / or mobile terminal.
[0094] It is understandable that the control commands obtained during the command fission process may conflict with the control commands currently being executed by the vehicle. For example, if the vehicle's interior temperature is higher than the temperature threshold (e.g., 38°C), the generated control command may be "cooling down," but the air conditioning is already at full capacity. In this case, the conflicting command may be recorded on the vehicle's infotainment system or fed back to one or two mobile terminals to prompt passengers to make a decision. This decision may also be recorded to optimize and update the local command generation model.
[0095] Furthermore, while providing feedback to passengers to make decisions, it can also parse conflict control commands based on a preset local command generation model and execute a safety rollback mechanism, such as outputting a "lower window" command.
[0096] Example 4 This application provides another vehicle control method based on multimodal information. In this embodiment, the vehicle's infotainment system is also connected to a mobile terminal and home appliances. Figure 4 As shown, the method includes: S410 extracts features from the raw voiceprint data of passengers to obtain the acoustic features of Mel frequency cepstral coefficients.
[0097] S420, based on the acoustic characteristics of Mel frequency cepstral coefficients and the biometric data of passengers, obtains a dynamic voiceprint key and the initial control command corresponding to the voiceprint key.
[0098] The S430, through the collection of vehicle environmental data, vehicle status data, and current biometric data, as well as voiceprint keys, performs instruction fission processing on the initial control commands to obtain the target instruction set. At the same time, it inputs the initial control commands into the edge computing node and, combined with the status data of the mobile terminal and home devices, generates the terminal control instruction set in real time.
[0099] S440 controls the vehicle based on the first control command, and at the same time, distributes terminal control commands from the terminal control command set to the corresponding mobile terminal or home device through a quantum-encrypted channel.
[0100] The difference between this embodiment and embodiments 1, 2, and 3 is that this method can not only control the vehicle through the first control command, but also control one or more of the mobile terminal and home appliances according to the terminal control command, so as to realize the coordinated control of vehicle-mobile terminal-home appliances.
[0101] Specifically, the vehicle-mounted system also communicates with mobile terminals and home devices. After determining the initial control command based on the voiceprint key, the initial control command is fused with the status data of the mobile terminal and home devices through edge computing nodes to generate a set of terminal control commands. Then, the control commands of each terminal are distributed to the corresponding mobile terminal and home devices to achieve collaborative control of vehicle-mobile terminal-home devices. Specifically, the dehumidification control command corresponding to the voiceprint key is divided into three segments using the Shamir algorithm: vehicle (60%), mobile phone (30%), and home (10%). After quantum encryption transmission, the command is verified and executed by an edge device cluster deployed at the edge nodes. The command is then distributed end-to-end via the 5G network, completing the secure migration of commands from the vehicle to the home device to the mobile terminal within 2.7 seconds (such as the vehicle's "sleep mode" simultaneously turning off home lights and setting the mobile phone to "do not disturb"). This breaks down ecological silos. At the same time, relying on the fragmentation self-destruction mechanism and federated learning architecture, the risk of privacy leakage can be reduced to 0.8%. The entire process achieves cross-ecological secure collaboration from voiceprint to full-domain control (i.e., vehicle-home device-mobile terminal) under the premise of zero privacy leakage, achieving the natural evolution of voiceprint and vehicle interaction, and providing a seamless and reliable ecological neural network for personalized commands.
[0102] Understandably, during the shard migration process, if any shard is missing, the edge node will initiate zero-knowledge proof (i.e., report the missing information to the monitoring node on the vehicle terminal, but does not need to report the specific content of the shard, the complete data status of its own storage, or other details that may expose privacy), and then end the shard migration process.
[0103] In this embodiment, the terminal control command is obtained by fusing and splitting the initial control command with the status data of the edge computing node, the mobile terminal, and the home device.
[0104] Example 5 This application provides another vehicle control method based on multimodal information. In this embodiment, the vehicle's infotainment system is also communicatively connected to a mobile terminal and home appliances. The initial control command includes at least one vehicle control command and at least one initial terminal control command. The target control command set also includes at least one terminal control command. Figure 5 As shown, the method includes: S510 extracts features from the raw voiceprint data of passengers to obtain the acoustic features of Mel frequency cepstral coefficients.
[0105] S520, based on the acoustic characteristics of Mel frequency cepstral coefficients and the biometric data of passengers, obtains a dynamic voiceprint key and an initial control command corresponding to the voiceprint key; the initial control command includes at least one initial terminal control command.
[0106] S530 uses the collected vehicle environment data, vehicle status data, and current biometric data, along with the voiceprint key, to perform instruction fission processing on the initial control command to obtain a target instruction set. The target instruction set includes at least one first control command obtained based on the fission of the vehicle control command, and at least one terminal control command obtained based on the fission of the initial terminal control command.
[0107] S540 controls the vehicle based on the first control command, and at the same time, distributes terminal control commands to the corresponding mobile terminal or home device through a quantum-encrypted channel.
[0108] The difference between this embodiment and embodiments 1, 2, and 3 is that this method can not only control the vehicle through the first control command, but also control one or more of the mobile terminal and home devices according to the terminal control command, so as to realize the coordinated control of vehicle-mobile terminal-home devices; the difference between this embodiment and embodiment 4 is that the initial terminal control command exists in the initial control command in this embodiment, that is, the initial terminal control command for controlling the terminal is obtained in the process of determining the initial control command based on the voiceprint key.
[0109] In this embodiment, the initial terminal control command undergoes multimodal data fusion and command fission following step S530 to generate terminal control commands stored in the target command set. Subsequently, the terminal control commands are distributed to the corresponding mobile terminals or home devices to achieve coordinated control of the vehicle, mobile terminal, and home device.
[0110] In this embodiment, to enhance vehicle control safety, a dynamic arbitration gateway is preferred to intercept illegal commands in real time, eliminating the risk of dangerous driving. In extreme environments, temperatures below 40 degrees Celsius can cause FPGA computation latency to increase exponentially, while temperatures above 85 degrees Celsius can trigger memory bit flips. In such cases, a higher-specification configuration is recommended. Additionally, built-in error correction settings and a nano-hydrophobic coating are added to achieve moisture and condensation prevention. This approach addresses the "hardware-software collaboration" and "cross-layer optimization" problem: reducing biometric errors through noise-resistant sensing and cross-validation to improve accuracy; enhancing response speed through FPGA hardware acceleration and CAN bus direct-write protocol to ensure real-time performance; and reducing leakage risks through quantum-classical hybrid encryption and zero-knowledge proofs to ensure security. Ultimately, under automotive-grade BOM cost constraints, this achieves a generational leap from "fixed response" to "autonomous evolution" in voiceprint control.
[0111] Example 6 A second aspect of this application also provides a vehicle control device based on multimodal information, combined with... Figure 6 As shown, the device includes: a feature extraction module 10, a generation module 20, an instruction fission module 30, and a control module 40.
[0112] The feature extraction module 10 is used to extract features from the raw voiceprint data of the passengers to obtain the acoustic features of the Mel frequency cepstral coefficients.
[0113] The generation module 20 is used to obtain a dynamic voiceprint key and the initial control command corresponding to the voiceprint key based on the acoustic characteristics of the Mel frequency cepstral coefficients and the biometric data of the passengers.
[0114] The instruction fission module 30 is used to perform instruction fission processing on the initial control instruction by collecting vehicle environment data, vehicle status data and current biological data, as well as voiceprint key, to obtain a target instruction set; the target instruction set includes at least one first control instruction.
[0115] In conjunction with the second aspect, the generation module 20 includes: an intermediate entropy value generation module, a voiceprint key generation module, and an initial control command generation module.
[0116] The intermediate entropy generation module is used to generate an intermediate entropy value of a specified number of bits based on pre-configured weighted information, fusion of Mel frequency cepstral coefficient acoustic features and biological data.
[0117] The voiceprint key generation module is used to dynamically encrypt the intermediate entropy value based on pre-configured encryption information to generate a voiceprint key; The initial control command generation module is used to retrieve the command mapping table based on the voiceprint key to obtain the initial control command corresponding to the voiceprint key.
[0118] In conjunction with the second aspect, the instruction fission module 30 includes: a fission module, a feedback module, and an update module.
[0119] The fission module is used to input vehicle environmental data, vehicle status data, current biological data and voiceprint key into a preset instruction generation model to perform instruction fission and obtain the fission control instruction set.
[0120] The feedback module is used to collect the biofeedback information of passengers when the vehicle is running with the control command set after fission.
[0121] The update module is used to update the instruction generation model based on biofeedback information to generate a target instruction set, which includes at least one first control instruction.
[0122] In conjunction with the second aspect, after the fission module 30, there are also: a determination module and a sending module.
[0123] The determination module is used to determine the target actuator based on each post-fission control instruction in the post-fission control instruction set.
[0124] The sending module is used to send the split control commands to the target actuator so that the target actuator can run according to the split control commands.
[0125] In conjunction with the second aspect, the vehicle-mounted terminal communicates with the mobile terminal; after the fission module 30, it also includes: a detection module and a recording module.
[0126] The detection module is used to perform conflict detection on each of the split control instructions in the split control instruction set based on a predefined conflict rule base, and obtain the detection results; The recording module is used to record and / or feed back conflicting instructions to the mobile terminal if the detection result indicates that there are conflicting instructions.
[0127] In conjunction with the second aspect, the vehicle-mounted system also communicates with mobile terminals and home devices; the device also includes a computing module and an encrypted distribution module.
[0128] The computing module is used to process the initial control command by collecting vehicle environment data, vehicle status data, current biological data, and voiceprint key to obtain the target command set. At the same time, it inputs the initial control command into the edge computing node and combines the status data of mobile terminal and home device to generate the terminal control command set in real time. The encrypted distribution module is used to control the vehicle based on the first control command, and at the same time, distributes the terminal control commands in the terminal control command set to the corresponding mobile terminal or home device through the quantum encrypted channel.
[0129] In conjunction with the second aspect, the vehicle-mounted system also communicates with mobile terminals and home devices; the initial control commands include at least one vehicle control command and at least one initial terminal control command; the target control command set also includes at least one terminal control command; after the generation module 20, it also includes: The second fission module is used to perform instruction fission processing on the initial control command by collecting vehicle environment data, vehicle status data, and current biological data, as well as voiceprint key, to obtain a target instruction set; the target instruction set includes at least one first control command obtained based on vehicle control command fission, and at least one terminal control command obtained based on initial terminal control command fission. The terminal control command distribution module is used to control the vehicle based on the first control command, and simultaneously distributes the terminal control command to the corresponding mobile terminal or home device through a quantum-encrypted channel. Thirdly, embodiments of this application provide an electronic device, combined with... Figure 7 As shown, the electronic device includes a memory 131 and a processor 130. The memory 131 stores a computer program, and the processor 130 runs the computer program to make the electronic device perform the above-described method.
[0130] Furthermore, combined Figure 7 The electronic device shown also includes a bus 132 and a communication interface 133, with the processor 130, the communication interface 133 and the memory 131 connected via the bus 132.
[0131] The memory 131 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 133 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc. The bus 132 may be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 7 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0132] Processor 130 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 130 or by instructions in software form. Processor 130 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 131, and processor 130 reads the information in memory 131 and, in conjunction with its hardware, completes the steps of the method described in the foregoing embodiments.
[0133] Fourthly, embodiments of this application provide a storage medium storing computer program instructions, which are read and executed by a processor to perform the above-described method.
[0134] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system and apparatus described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0135] Furthermore, in the description of the embodiments of the present invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.
[0136] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0137] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "second," "third," and "secondary" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0138] Finally, it should be noted that the above embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A vehicle control method based on multimodal information, characterized in that, Applied to in-vehicle infotainment systems; the method includes: Feature extraction was performed on the raw voiceprint data of passengers to obtain the acoustic features of Mel frequency cepstral coefficients; Based on the acoustic characteristics of the Mel frequency cepstral coefficients and the biometric data of the passengers, a dynamic voiceprint key and an initial control command corresponding to the voiceprint key are obtained. By collecting vehicle environment data, vehicle status data, and current biometric data, as well as the voiceprint key, the initial control command is subjected to command fission processing to obtain a target command set; the target command set includes at least one first control command; the current biometric data is the biometric data of the current passenger. Based on the first control command, control the vehicle; The step of performing instruction fission processing on the initial control command to obtain the target instruction set by collecting vehicle environmental data, vehicle status data, and current biometric data, as well as the voiceprint key, includes: The vehicle environment data, vehicle status data, current biometric data, and voiceprint key are input into a preset instruction generation model to perform instruction fission, resulting in a fission control instruction set. For each fission control instruction in the fission control instruction set, a target actuator is determined based on the fission control instruction. The fission control instruction is sent to the target actuator to control the target actuator to operate according to the fission control instruction. Biometric feedback information of the occupants is collected when the vehicle is running with the fission control instruction set. The instruction generation model is updated based on the biometric feedback information to generate the target instruction set.
2. The method according to claim 1, characterized in that, The steps of obtaining a dynamic voiceprint key and an initial control command corresponding to the voiceprint key based on the acoustic characteristics of the Mel frequency cepstral coefficients and the biometric data of the passenger include: Based on pre-configured weighted information, the acoustic features of the Mel frequency cepstral coefficients are fused with the biological data to generate an intermediate entropy value with a specified number of bits. Based on pre-configured encryption information, the intermediate entropy value is dynamically encrypted to generate the voiceprint key; Based on the voiceprint key retrieval instruction mapping table, the initial control instruction corresponding to the voiceprint key is obtained.
3. The method according to claim 1, characterized in that, After the step of inputting the vehicle environment data, the vehicle status data, the current biometric data, and the voiceprint key into a preset instruction generation model to perform instruction fission and obtain the fissioned control instruction set, the method further includes: Based on a predefined conflict rule base, conflict detection is performed on each of the split control instructions in the split control instruction set to obtain the detection results; If the detection result indicates the presence of a conflicting command, the conflicting command is recorded and / or fed back to the mobile terminal.
4. The method according to claim 1, characterized in that, While executing the step of processing the initial control command to obtain the target command set by collecting vehicle environment data, vehicle status data, and current biological data, as well as the voiceprint key, the initial control command is input into the edge computing node and combined with the status data of the mobile terminal and home device to generate the terminal control command set in real time. The vehicle is controlled based on the first control command, and at the same time, terminal control commands in the terminal control command set are distributed to the corresponding mobile terminal or home device through a quantum encrypted channel.
5. The method according to claim 1, characterized in that, The initial control command includes at least one vehicle control command and at least one initial terminal control command; the target command set also includes at least one terminal control command. By collecting vehicle environment data, vehicle status data, and current biometric data, as well as the voiceprint key, the initial control command is subjected to command fission processing to obtain the target command set. The vehicle is then controlled based on the first control command. At the same time, the terminal control command is distributed to the corresponding mobile terminal or home device through a quantum encryption channel.
6. A vehicle control device based on multimodal information, characterized in that, Applied to in-vehicle infotainment systems; the device includes: The feature extraction module is used to extract features from the raw voiceprint data of passengers to obtain the acoustic features of Mel frequency cepstral coefficients. The generation module is used to obtain a dynamic voiceprint key and an initial control command corresponding to the voiceprint key based on the acoustic features of the Mel frequency cepstral coefficients and the biometric data of the passenger. The instruction fission module is used to perform instruction fission processing on the initial control instruction using the collected vehicle environment data, vehicle status data, and current biometric data, as well as the voiceprint key, to obtain a target instruction set; the target instruction set includes at least one first control instruction; the current biometric data is the biometric data of the current passenger. A control module is used to control the vehicle based on the first control command; The instruction fission module is used to input the vehicle environment data, vehicle status data, current biometric data, and voiceprint key into a preset instruction generation model to perform instruction fission, thereby obtaining a fissioned control instruction set; for each fissioned control instruction in the fissioned control instruction set, a target actuator is determined based on the fissioned control instruction; the fissioned control instruction is sent to the target actuator to control the target actuator to run according to the fissioned control instruction; biometric feedback information of the passengers is collected when the vehicle runs according to the fissioned control instruction set; and the instruction generation model is updated based on the biometric feedback information to generate the target instruction set.
7. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program and the processor running the computer program to cause the electronic device to perform the method of any one of claims 1 to 5.
8. A storage medium, characterized in that, The storage medium stores computer program instructions, which, when read and executed by a processor, perform the method described in any one of claims 1 to 5.
Citation Information
Patent Citations
Vehicle lamp external voice turn-on method and system
CN117373449A
Implementation method and system of vehicle voice assistant, optimization method of mobile terminal voice assistant and storage medium
CN118230736A