Vehicle control method and apparatus, device and medium

By integrating voice and video data into a command generation model, and combining it with the user's multi-dimensional emotional characteristics, the problem of inaccurate control in in-vehicle voice recognition systems has been solved, achieving higher vehicle control accuracy and user experience.

WO2026149019A1PCT designated stage Publication Date: 2026-07-16CHONGQING CHANGAN AUTOMOBILE CO LTD

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
CHONGQING CHANGAN AUTOMOBILE CO LTD
Filing Date
2025-11-13
Publication Date
2026-07-16

AI Technical Summary

Technical Problem

In existing technologies, in-vehicle voice recognition systems often fail to provide accurate voice recognition results due to overly simplistic consideration of factors, resulting in imprecise vehicle control and a poor user experience.

Method used

By acquiring voice and video data from the vehicle, and using an instruction generation model to fuse voice and video features, target instruction text is generated. By combining multi-dimensional emotional features such as the user's tone, pitch, volume, and facial expressions, the accuracy of intent recognition is improved.

Benefits of technology

It improves the accuracy of vehicle control and user experience, and can set function parameters according to user emotions, providing more intelligent services and feedback, and reducing the risks caused by distracted operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025134806_16072026_PF_FP_ABST
    Figure CN2025134806_16072026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a vehicle control method and apparatus, a device and a medium. The method comprises: acquiring voice data and video data of a target user in a vehicle, and inputting the voice data and the video data into an instruction generation model to acquire target instruction text output by the instruction generation model; and then, by means of the target instruction text, controlling the vehicle, the instruction generation model being configured to fuse an audio feature corresponding to the voice data and an image feature corresponding to the video data, so as to generate the target instruction text, the audio feature comprising a first text sub-feature and a first emotion sub-feature, the image feature comprising a second text sub-feature and a second emotion sub-feature, the first emotion sub-feature being used to represent at least one of tone, intonation and volume of the target user's speech, and the second emotion sub-feature being used to represent a facial expression and / or a lip movement of the target user. The present invention improves the accuracy of vehicle control and improves user's driving and riding experience.
Need to check novelty before this filing date? Find Prior Art

Description

Vehicle control methods, devices, equipment and media

[0001] This application claims priority to Chinese Patent Application No. 202510044082.8, filed on January 10, 2025, entitled “Vehicle Control Method, Apparatus, Device and Medium”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This invention relates to the field of automotive control technology, specifically to a vehicle control method, device, equipment, and medium. Background Technology

[0003] With the upgrading of automobile consumption, the demand for personalization, and the development of intelligent connected vehicles, in-vehicle voice recognition systems have become one of the important factors for many people when purchasing a car. In-vehicle voice recognition systems can process users' voice commands in real time and control the vehicle to respond quickly, greatly improving the user's driving experience and reducing the risk of accidents caused by distracted operation of vehicle functions.

[0004] Currently, the main approach is to use an Automatic Speech Recognition (ASR) model to recognize the user's voice commands, obtain the voice recognition results, and then control the vehicle to execute the corresponding operations based on the voice recognition results.

[0005] However, existing technology relies on voice recognition to identify the user's intention to control the vehicle, and then controls the vehicle based on the identified intention. But because it considers too few factors, the identified intention is inaccurate, resulting in imprecise vehicle control and a poor user experience. Summary of the Invention

[0006] The purpose of this invention is to provide a vehicle control method, device, equipment, and medium to solve the problems of insufficient precision in vehicle control and poor user experience in the prior art.

[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0008] Firstly, a vehicle control method includes:

[0009] Acquire voice and video data of the target user in the vehicle;

[0010] The voice data and video data are input into the instruction generation model to obtain the target instruction text output by the instruction generation model. The instruction generation model is used to fuse the audio features corresponding to the voice data and the image features corresponding to the video data to generate the target instruction text. The audio features include a first text sub-feature and a first emotion sub-feature, and the image features include a second text sub-feature and a second emotion sub-feature. The first emotion sub-feature is used to represent at least one of the target user's tone, intonation, and volume, and the second emotion sub-feature is used to represent the target user's facial expression and / or lip changes.

[0011] The vehicle is controlled by the target instruction text.

[0012] Based on the aforementioned technical methods, the command generation model integrates the target user's voice and video data, combined with the target user's current emotions, to comprehensively determine the user's intent, thereby outputting more accurate target command text. For the user, the entire process is seamless, but the accuracy of vehicle control based on the target command text and the target user's driving experience are greatly improved, enabling the vehicle to provide more intelligent services and feedback to the target user.

[0013] Furthermore, in the instruction generation model, the weight of the first emotion sub-feature is greater than the weight of the second emotion sub-feature.

[0014] Based on the aforementioned technical means, since the rhythm, tone, volume, and other elements contained in the first emotion sub-feature can directly reflect the user's current emotion, the first emotion sub-feature is given a higher weight than the second emotion sub-feature. This allows the model to pay more attention to the first emotion sub-feature during the inference process, thereby ensuring that the determined user emotion and user intent are more accurate and improving the accuracy of the target instruction text output by the subsequent instruction generation model.

[0015] Furthermore, the acquisition of the target user's voice and video data in the vehicle includes:

[0016] The voice data is collected through the microphone in the vehicle, and the initial video data is collected through the camera in the vehicle;

[0017] Based on the voice data, determine the target location of the target user corresponding to the voice data in the vehicle;

[0018] The initial video data is processed based on the target location to obtain the video data corresponding to the target user.

[0019] Based on the aforementioned technical means, by obtaining the corresponding target user's video data from the voice data, it is possible to associate the voice data and video data of the same user when there are multiple users in the vehicle. This effectively manages and coordinates the needs of different users, avoids confusion and misunderstanding, and further improves the accuracy of the generated target instruction text.

[0020] Furthermore, the acquisition of the voice data via a microphone in the vehicle includes:

[0021] Initial voice data is collected using a microphone in the vehicle;

[0022] The initial speech data is subjected to noise reduction processing to generate the speech data.

[0023] Based on the above technical means, noise reduction processing of the initial speech data can significantly reduce the background noise of the initial speech data, making the human voice in the obtained speech data clearer and easier to recognize and understand.

[0024] Furthermore, the step of performing noise reduction processing on the initial speech data to generate the speech data includes:

[0025] The filter coefficients of the adaptive filter are determined based on the environmental noise measurement or echo signal of the initial speech data.

[0026] Based on the filter coefficients, the initial speech data is denoised using the adaptive filter to generate the speech data.

[0027] Based on the above technical means, by adjusting the filter coefficients, the noise component in the output signal (speech data) is minimized, the echo in the initial speech data is canceled, the recognizability of human voice is improved, and the interference of noise on the human voice extraction process is reduced.

[0028] Furthermore, the instruction generation model includes:

[0029] The first modal projection multilayer sensing layer, the second modal projection multilayer sensing layer, and the Transformer network;

[0030] The input of the Transformer network is connected to the output of the first modal projection and the output of the second modal projection, respectively.

[0031] The first modal projection multilayer perception layer is used to project the initial audio features corresponding to the speech data onto a preset feature space to generate audio features. The initial audio features include a first initial text sub-feature and a first initial emotion sub-feature.

[0032] The second modal projection multilayer perception layer is used to project the initial image features corresponding to the video data onto the preset feature space to generate image features. The initial image features include a second initial text sub-feature and a second initial emotion sub-feature.

[0033] The Transformer network is used to output the target instruction text based on the encoded word embedding sequence, which is obtained by word segmentation and position encoding of the audio features and the image features.

[0034] Based on the aforementioned technical means, data from different modalities (i.e., video data and audio data) are projected onto the same or compatible feature space through a first modality projection multilayer perception layer and a second modality projection multilayer perception layer, allowing them to be aligned and compared within a shared representation space. Simultaneously, both audio and image features contain emotion sub-features representing the target user's emotions. By segmenting and positionally encoding the audio and image features, encoded word embedding sequences are obtained and then input into a Transformer network for processing. The encoders of the Transformer network in the instruction generation model can share the same network structure and parameters. Compared to constructing separate complex encoders for video and audio data, this technique significantly reduces the total number of parameters in the instruction generation model without affecting processing accuracy.

[0035] Furthermore, the instruction generation model also includes an embedding layer, the input of which is connected to the output of the first modal projection multilayer perception layer and the output of the second modal projection multilayer perception layer, respectively, and the output of which is connected to the input of the Transformer network.

[0036] The embedding layer is used to map each word to a corresponding word embedding sequence;

[0037] The positional encoding corresponding to each word is added to the word embedding sequence corresponding to the word to generate the encoded word embedding sequence.

[0038] Based on the aforementioned techniques, the encoded word embedding sequence captures the semantic information of words, mapping each word to its corresponding encoded word embedding sequence. This allows the Transformer network to directly utilize this semantic information for processing. Furthermore, since the Transformer network is built on a self-attention mechanism, the semantic information of the encoded word embedding sequence can be effectively integrated with the self-attention mechanism. The self-attention mechanism can better focus on the semantic relationships between words at different positions, thereby generating more accurate target instruction text.

[0039] Furthermore, the instruction generation model also includes:

[0040] Image-to-audio layer binding;

[0041] The output of the image-bound audio layer is connected to the input of the first modal projection multilayer perception layer. The image-bound audio layer is used to adjust the feature format of the initial audio features to the feature format of the initial image features.

[0042] Based on the above technical means, by binding the image to the audio layer, the data of the two modalities of image and audio are bound or associated, so as to realize cross-modal data fusion and understanding in the future.

[0043] Furthermore, before acquiring the voice and video data of the target user in the vehicle, the method further includes:

[0044] The instruction generation model is preloaded into the vehicle's memory.

[0045] Based on the above technical means, by preloading the instruction generation model into the vehicle's memory, the loading time of the instruction generation model can be effectively reduced, and the inference efficiency of the target instruction model can be improved.

[0046] Secondly, a vehicle control method, prior to acquiring the voice data and video data of a target user in the vehicle, the method further includes:

[0047] The initial instruction generation model is obtained and stored by training the sample voice data, sample video data and tags of the sample users.

[0048] The initial instruction generation model is converted to the inference architecture supported by the vehicle to generate the instruction generation model;

[0049] The instruction generation model is deployed to the vehicle.

[0050] Based on the aforementioned technical methods, and leveraging the powerful computing capabilities of the cloud platform, a high-accuracy initial command generation model is pre-trained using sample voice data, sample video data, and labels from sample users to generate and store this model. Subsequently, the initial command generation model is converted into a vehicle-supported inference architecture to generate a new command generation model, which is then deployed to the vehicle. In this way, the vehicle does not need to perform complex model training locally, thus saving computing resources and time. Furthermore, the cloud platform can utilize its platform advantages to generate higher-quality command generation models, enabling them to have better generalization ability and accuracy. During model inference, after the vehicle collects voice and video data, it can directly process the voice and video data on-vehicle using the command generation model to control the vehicle, eliminating the need for interaction with the cloud platform and improving processing efficiency.

[0051] Furthermore, the method also includes:

[0052] Reduce the representation accuracy of the weights and / or the output of the activation function of the initial instruction generation model.

[0053] Based on the above technical means, the memory space required for storing weights, the amount of data in the model calculation process, and the burden on the vehicle to run the initial instruction to generate the model are effectively reduced, thereby improving the model running speed.

[0054] Furthermore, the method also includes:

[0055] Delete redundant connections and / or redundant neurons in the initial instruction generation model.

[0056] Based on the above technical means, the model structure of the initial instruction generation model is effectively simplified, and the memory space required for storing weights and the amount of data in the model operation process are further reduced.

[0057] Thirdly, a vehicle control device includes:

[0058] The acquisition module is used to acquire voice and video data of the target user in the vehicle.

[0059] An input module is used to input the voice data and the video data into an instruction generation model, and obtain the target instruction text output by the instruction generation model. The instruction generation model is used to fuse the audio features corresponding to the voice data and the image features corresponding to the video data to generate the target instruction text. The audio features include a first text sub-feature and a first emotion sub-feature. The image features include a second text sub-feature and a second emotion sub-feature. The first emotion sub-feature is used to represent at least one of the target user's tone, intonation, and volume. The second emotion sub-feature is used to represent the target user's facial expression and / or lip changes.

[0060] The control module is used to control the vehicle via the target instruction text.

[0061] Furthermore, in the instruction generation model, the weight of the first emotion sub-feature is greater than the weight of the second emotion sub-feature.

[0062] Furthermore, the acquisition module is specifically used for:

[0063] The voice data is collected through the microphone in the vehicle, and the initial video data is collected through the camera in the vehicle;

[0064] Based on the voice data, determine the target location of the target user corresponding to the voice data in the vehicle;

[0065] The initial video data is processed based on the target location to obtain the video data corresponding to the target user.

[0066] Furthermore, the acquisition module is specifically used for:

[0067] Initial voice data is collected using a microphone in the vehicle;

[0068] The initial speech data is subjected to noise reduction processing to generate the speech data.

[0069] Furthermore, the acquisition module is specifically used for:

[0070] The filter coefficients of the adaptive filter are determined based on the environmental noise measurement or echo signal of the initial speech data.

[0071] Based on the filter coefficients, the initial speech data is denoised using the adaptive filter to generate the speech data.

[0072] Furthermore, the instruction generation model includes:

[0073] The first modal projection multilayer sensing layer, the second modal projection multilayer sensing layer, and the Transformer network;

[0074] The input of the Transformer network is connected to the output of the first modal projection and the output of the second modal projection, respectively.

[0075] The first modal projection multilayer perception layer is used to project the initial audio features corresponding to the speech data onto a preset feature space to generate audio features. The initial audio features include a first initial text sub-feature and a first initial emotion sub-feature.

[0076] The second modal projection multilayer perception layer is used to project the initial image features corresponding to the video data onto the preset feature space to generate image features. The initial image features include a second initial text sub-feature and a second initial emotion sub-feature.

[0077] The Transformer network is used to output the target instruction text based on the encoded word embedding sequence, which is obtained by word segmentation and position encoding of the audio features and the image features.

[0078] Furthermore, the instruction generation model also includes an embedding layer, the input of which is connected to the output of the first modal projection multilayer perception layer and the output of the second modal projection multilayer perception layer, respectively, and the output of which is connected to the input of the Transformer network.

[0079] The embedding layer is used to map each word to a corresponding word embedding sequence;

[0080] The positional encoding corresponding to each word is added to the word embedding sequence corresponding to the word to generate the encoded word embedding sequence.

[0081] Furthermore, the instruction generation model also includes:

[0082] Image-to-audio layer binding;

[0083] The output of the image-bound audio layer is connected to the input of the first modal projection multilayer perception layer. The image-bound audio layer is used to adjust the feature format of the initial audio features to the feature format of the initial image features.

[0084] Furthermore, before acquiring the voice and video data of the target user in the vehicle, the vehicle control device also includes a preloading module for preloading the instruction generation model into the vehicle's memory.

[0085] Fourthly, a vehicle control device, prior to acquiring voice and video data of a target user in the vehicle, further includes a deployment module for:

[0086] The initial instruction generation model is obtained and stored by training the sample voice data, sample video data and tags of the sample users.

[0087] The initial instruction generation model is converted to the inference architecture supported by the vehicle to generate the instruction generation model;

[0088] The instruction generation model is deployed to the vehicle.

[0089] Furthermore, the deployment module is also used for:

[0090] Reduce the representation accuracy of the weights and / or the output of the activation function of the initial instruction generation model;

[0091] Furthermore, the deployment module is also used for:

[0092] Delete redundant connections and / or redundant neurons in the initial instruction generation model;

[0093] Fifthly, a vehicle includes: a processor, a memory, and computer-executable instructions stored in the memory and executable on the processor, wherein the processor executes the computer-executable instructions to implement the vehicle control method described in the first aspect.

[0094] A sixth aspect is a cloud platform, comprising: a processor, a memory, and computer-executable instructions stored in the memory and executable on the processor, wherein the processor executes the computer-executable instructions to implement the vehicle control methods shown in the first and second aspects.

[0095] A seventh aspect is a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, are used to implement the vehicle control methods shown in the first and second aspects.

[0096] Eighth aspect, a computer program product comprising a computer program, which, when executed by a processor, is used to implement the vehicle control methods shown in the first and second aspects.

[0097] In a ninth aspect, embodiments of the present invention provide a chip, the chip including a memory and a processor, the memory storing code and data, the memory being coupled to the processor, and the processor running a program in the memory causing the chip to perform the vehicle control methods shown in the first and second aspects above.

[0098] In a tenth aspect, embodiments of the present invention provide a computer program, which, when executed by a processor, performs the vehicle control methods described in the first and second aspects above.

[0099] The beneficial effects of this invention are:

[0100] (1) This technical solution integrates the target user's voice data and video data through an instruction generation model, and combines the target user's current emotions to comprehensively determine the user's intent, thereby outputting more accurate target instruction text. The target user's tone, intonation, and volume, as reflected in the voice data, are the primary indicators, while facial expressions and / or lip changes, as reflected in the video data, are secondary indicators. This multi-dimensional approach effectively ensures the accuracy of the user's emotional state determination. Simultaneously, the user's intent is comprehensively determined based on the text reflected in the voice and video data, as well as the aforementioned determined user emotions, thus generating the target instruction text. When controlling the vehicle based on the target instruction text, not only can the vehicle's functions be controlled, but the parameters within those functions can also be set according to the user's current emotions, improving the accuracy of vehicle control, making the in-vehicle environment more aligned with the user's current needs, enhancing the user's driving experience, and providing more intelligent services and feedback.

[0101] (2) The instruction generation model includes a first modal projection multilayer perceptron, a second modal projection multilayer perceptron, and a Transformer network. The first and second modal projection multilayer perceptrons project the initial audio and image features into the same preset feature space, obtaining audio and image features. These features are then segmented and positionally encoded before being input into the Transformer network for further processing to obtain the target instruction text. Compared to constructing separate complex encoders for video and audio data, the instruction generation model proposed in this invention requires only one encoder to achieve the fusion processing of audio and image features, significantly reducing the total number of model parameters without affecting processing accuracy.

[0102] (3) This invention utilizes a cloud platform to train the model in the cloud and sends the trained instruction generation model to the vehicle. The vehicle does not need to perform complex model training locally, thus saving computing resources and time. Furthermore, the cloud platform can leverage its advantages to generate higher-quality instruction generation models, enabling them to have better generalization ability and accuracy. During model inference, after the vehicle collects voice and video data, it can directly run the instruction generation model on the vehicle to process the voice and video data, saving the interaction process between the vehicle and the cloud platform and improving processing efficiency. Attached Figure Description

[0103] Figure 1 is a schematic diagram of the vehicle structure provided in an embodiment of the present invention;

[0104] Figure 2 is a schematic flowchart of the vehicle control method provided in an embodiment of the present invention;

[0105] Figure 3 is a schematic diagram of a scenario for the vehicle control method provided in an embodiment of the present invention;

[0106] Figure 4 is a schematic diagram of the instruction generation model provided in an embodiment of the present invention;

[0107] Figure 5 is a schematic flowchart of the vehicle control method provided in an embodiment of the present invention.

[0108] Figure 6 is a system schematic diagram of the vehicle software system provided in an embodiment of the present invention;

[0109] Figure 7 is a second scenario diagram of the vehicle control method provided in an embodiment of the present invention;

[0110] Figure 8 is a schematic diagram of the vehicle control device provided in an embodiment of the present invention;

[0111] Figure 9 is a schematic diagram of the vehicle control device provided in an embodiment of the present invention.

[0112] Figure 10 is a second structural schematic diagram of the vehicle provided in an embodiment of the present invention;

[0113] Figure 11 is a schematic diagram of the structure of the cloud platform provided in an embodiment of the present invention. Detailed Implementation

[0114] The embodiments of the present invention will be described below with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention and not for limiting the scope of protection of the present invention.

[0115] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this invention are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0116] Before introducing the present invention, the application background of the present invention will be explained first.

[0117] With the upgrading of automobile consumption, personalized demands, and the development of intelligent connected vehicles, in-vehicle voice recognition systems have become an important consideration for many car buyers. In-vehicle voice recognition systems can interact with users, analyze their speech, identify their intentions and needs, and then control the vehicle. Controlling the vehicle through an in-vehicle voice recognition system allows the driver to keep both hands on the steering wheel and eyes on the road, thus reducing the risk of accidents caused by distracted operation of vehicle functions. Therefore, with in-vehicle voice recognition systems, drivers can control vehicle functions more conveniently without sacrificing driving comfort.

[0118] In practical applications, the system can capture user voice input via microphone and then convert it into text using Automatic Speech Recognition (ASR). The system then identifies the user's intent based on this text and controls the vehicle to perform the corresponding operation based on the voice recognition result (user intent). In other words, the in-vehicle voice recognition system directly generates corresponding commands based on the user's voice input and then controls the vehicle according to those commands.

[0119] In existing technologies, since most of a user's attention is focused on driving the vehicle, the commands they issue when they want to actively control the vehicle are relatively simple, such as "turn on the air conditioning." Even when the user doesn't actively want to control the vehicle, existing technologies can still control it through voice recordings, such as turning on the air conditioning or opening the windows to cool it down when the user says "I'm so hot." However, in both of these cases, specific control parameters are lacking. Taking "I'm so hot" as an example, the in-vehicle voice recognition system would recognize the user's intention to "cool down" based on this statement, and then control the vehicle to open windows or turn on the air conditioning to lower the interior temperature. For turning on the air conditioning, the control parameters such as temperature and fan speed are usually determined based on default values, common values, and ambient temperature. Similarly, for opening windows, the control parameters such as the window opening degree and the number of windows to be opened can also be determined based on default values, common values, and ambient temperature.

[0120] However, a user's intention to control the vehicle is not solely reflected in their words. Controlling the vehicle solely through words can lead to lower accuracy in vehicle control, thus affecting the user's driving experience.

[0121] In summary, existing technologies lack a clear and comprehensive understanding of user intentions and needs, resulting in imprecise vehicle control and a poor user experience.

[0122] Based on the aforementioned technical issues, and considering that the speaker's emotions also determine the user's control intentions regarding the vehicle, and that these emotions are generally expressed through tone of voice, intonation, and facial expressions—parameters that can be determined based on the user's voice and video containing the user's face—a multimodal large-scale model based on video and voice can be used to fuse the speaker's video and voice data. By comprehensively determining the speaker's control intentions regarding the vehicle based on both the primary language (what the speaker says) and the secondary language (the speaker's emotions), a more accurate target instruction text can be generated, allowing for vehicle control based on this target instruction text. This significantly improves the accuracy of vehicle control and enhances the user's driving experience.

[0123] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0124] Figure 1 is a schematic diagram of the vehicle structure provided in an embodiment of the present invention; as shown in Figure 1, the intelligent connected vehicle cockpit of the vehicle is equipped with a microphone (MIC) and a camera.

[0125] The microphone can be a regular microphone, a condenser microphone, or a micro-electro-mechanical systems (MEMS) microphone. The type of microphone can be determined according to the actual situation, and the embodiments of the present invention do not impose specific limitations on it.

[0126] In practical applications, the sound received by the MIC includes mixed and overlapping voice signals from users in the intelligent cockpit of intelligent connected vehicles, music signals from speakers, noise signals from the vehicle's powertrain system, road-tire noise, wind noise, and other noise source signals.

[0127] The cameras in the intelligent cockpit of intelligent connected vehicles can be infrared cameras (IC), red-green-blue cameras (RGB cameras), stereo cameras (SC), infrared and RGB combination cameras (IR+RGB Combo Cameras), thermal cameras (TC), fisheye cameras (FC), etc. The type of camera can be determined according to the actual situation, and the embodiments of the present invention do not impose specific limitations on this.

[0128] Optionally, the pixel resolution of the camera is not less than 640*480 or other preset pixel resolutions, which can be determined according to the actual situation. This embodiment of the invention does not impose specific limitations on this.

[0129] In practical applications, cameras are used to capture facial videos of each user in the vehicle.

[0130] Optionally, the MIC and camera can be installed in the center of the car dashboard, or in other locations in the vehicle, as long as the MIC can capture the user's voice and the camera can capture facial videos of each user in the smart cockpit of the intelligent connected vehicle. This embodiment of the invention does not impose specific restrictions on the installation location of the MIC and camera.

[0131] In practical applications, the microphone (MIC) converts the acquired audio signals into digital signals via an analog-to-digital converter (ADC) / A2B (Automotive Audio Bus). The camera converts the acquired image signals into digital signals via an image signal processor (ISP). Both the microphone and camera then transmit the converted digital signals to the intelligent cockpit system's host computer for processing. The host computer contains a command generation model that generates target command text based on the audio acquired by the microphone and the video captured by the camera, and then controls the vehicle based on this target command text.

[0132] Figure 2 is a schematic flowchart of a vehicle control method provided in an embodiment of the present invention. As shown in Figure 2, the vehicle control method is applied to a vehicle or a cloud platform, and the vehicle control method may include the following steps:

[0133] S21. Obtain the voice and video data of the target user in the vehicle.

[0134] The voice data is acquired through a microphone installed inside the vehicle, and the video data is acquired through a camera installed inside the vehicle. The target user is the user who is speaking.

[0135] Understandably, when a user speaks, the system uses microphones and cameras installed inside the vehicle to acquire the user's voice and video data, respectively, in order to provide a data analysis basis for the subsequent instruction generation model.

[0136] Referring to Figure 1, voice data can be converted from sound signals into corresponding digital signals, and video data can be converted from image signals into corresponding digital signals.

[0137] S22. Input the voice data and video data into the instruction generation model, and obtain the target instruction text output by the instruction generation model.

[0138] The instruction generation model is pre-trained based on sample voice data, sample video data, and sample control instruction text from sample users. It is used to fuse the audio features corresponding to the voice data and the image features corresponding to the video data to generate the target instruction text.

[0139] Furthermore, the audio features include a first text sub-feature and a first emotion sub-feature, and the image features include a second text sub-feature and a second emotion sub-feature. The first emotion sub-feature is used to represent at least one of the target user's tone of voice, intonation, and volume, and the second emotion sub-feature is used to represent the target user's facial expressions and / or lip changes.

[0140] For example, lip changes may include the degree of lip opening and closing, lip shape, and the rate of lip shape change, which can be determined according to the actual situation. This embodiment of the invention does not impose specific limitations on these aspects.

[0141] In one possible implementation, the weight of the first emotion sub-feature in the instruction generation model is greater than the weight of the second emotion sub-feature.

[0142] In the above methods, tone, pitch, and volume are instinctive and direct ways of expressing emotions. In human communication, elements such as tone, volume, and rhythm of speech can quickly convey emotional information. While facial expressions and / or lip movements are also direct manifestations of emotion, they are easily affected by external factors such as makeup and lighting. Therefore, assigning a higher weight to the first emotion sub-feature than the second emotion sub-feature allows the model to focus more on the first emotion sub-feature during inference, thereby ensuring a more accurate determination of the user's emotion and intent, and improving the accuracy of the target instruction text output by the subsequent instruction generation model.

[0143] It should be understood that the instruction generation model can be a large multimodal model based on video and audio. This instruction generation model can process multiple types of data such as video and audio simultaneously, integrate information from different sources, reduce information loss and latency, improve the accuracy of understanding the intent and needs of the target user, and thus output more accurate target instruction text.

[0144] Understandably, voice and video data are input into the command generation model. The command generation model then merges and processes the voice and video data to capture the target user's primary and secondary languages, further understanding the target user's intentions and needs, and thus outputting target command text to improve the accuracy of subsequent vehicle control based on the target command text.

[0145] For example, suppose a driver says "It's too hot" while driving. The instruction generation model can determine the vehicle function that needs to be controlled, such as turning on the air conditioning, based on the driver's voice and video data. Furthermore, it can determine the driver's emotion based on the voice and video data, and then determine the corresponding control parameters for the appropriate function. For instance, when the driver is angry, the air conditioning temperature is set lower to quickly lower the interior temperature; when the driver is calm, the air conditioning temperature is set higher to maintain a comfortable temperature for the human body. Compared to traditional ASR models, this invention, through its instruction generation model, can more accurately determine the user's current needs and identify more precise user intentions, thereby further improving the accuracy of vehicle control and the user's driving experience.

[0146] For example, emotions can include happiness, sadness, anger, surprise, etc.

[0147] It should be understood that the structure of the instruction generation model and the process by which the instruction generation model processes data will be explained in detail in subsequent embodiments, and will not be repeated here.

[0148] S23. Control the vehicle through the target instruction text.

[0149] In one possible implementation, a corresponding target control command can be generated based on the target command text, and then the vehicle can be controlled to execute the target control command in order to control the vehicle.

[0150] Optionally, after controlling the vehicle, the control result can be fed back to the user via voice and / or text. For example, the user can be informed via voice that "the window is open," and the text "the window is open" can also be displayed as a pop-up on the vehicle's display screen. Providing feedback to the user through multiple methods ensures that the control result is effectively communicated to the target user.

[0151] The vehicle control method provided in this invention acquires voice and video data of a target user in the vehicle, inputs the voice and video data into an instruction generation model, and obtains the target instruction text output by the instruction generation model. Then, the vehicle is controlled using the target instruction text. The instruction generation model fuses audio features corresponding to the voice data and image features corresponding to the video data to generate the target instruction text. The audio features include a first text sub-feature and a first emotion sub-feature, and the image features include a second text sub-feature and a second emotion sub-feature. The first emotion sub-feature represents at least one of the target user's tone, intonation, and volume, and the second emotion sub-feature represents the target user's facial expression and / or lip changes. In this technical solution, the tone, intonation, and volume of the target user's speech, as reflected in the voice data, are the primary factors, while the facial expression and / or lip changes of the target user in the video data are secondary factors. This multi-dimensional approach effectively ensures the accuracy of the user's emotion determination. Simultaneously, the user's intention is comprehensively determined based on the text reflected in the voice and video data, as well as the aforementioned determined user emotion, to generate the target instruction text. When controlling the vehicle based on target command text, it not only controls the vehicle's functions but also adjusts the parameters of those functions according to the user's current mood, improving the accuracy of vehicle control and making the in-car environment more aligned with the user's current needs. The driver can keep both hands on the steering wheel and eyes on the road, allowing for more precise control of vehicle functions without sacrificing driving comfort. This reduces the risk of accidents caused by distracted operation, enhances the user's driving experience, and provides more intelligent services and feedback.

[0152] In some embodiments, S21 can be implemented through the following process:

[0153] First, voice data is collected through the microphone in the vehicle, and initial video data is collected through the camera in the vehicle. Then, based on the voice data, the target location of the target user in the vehicle is determined. Finally, the initial video data is processed based on the target location to obtain the video data corresponding to the target user.

[0154] It should be noted that, referring to Figure 1, when there are multiple users in the vehicle, the initial video data collected by the vehicle's cameras may include video data from all users. In order to enable the instruction generation model to more accurately understand the target user's emotions and intentions, the initial video data also needs to be processed. That is, the target user's target location is determined based on the target user's voice data, and then the initial video data is cropped or the target user is marked in the initial video based on the target location, so as to finally obtain the video data corresponding to the target user.

[0155] For example, suppose there are two passengers in the car: the driver and the front passenger. The vehicle's microphone captures the front passenger saying, "I want to listen to some music." Simultaneously, the camera also captures initial video data of the vehicle, which includes the faces of both the driver and the front passenger. First, the front passenger's voice data is analyzed to determine that the sound originates from the front passenger's location. Then, the front passenger's image is cropped from the initial video data, or the front passenger is labeled within the initial video data, thus generating corresponding video data for the front passenger.

[0156] Optionally, in some scenarios, if each location of the vehicle is equipped with a corresponding camera, then after acquiring the target user's voice data, the target user's target location is determined based on the target user's voice data, and the initial video data collected by the camera corresponding to the target user's location is used to determine the target video data.

[0157] It should be understood that in this way, when there are multiple users in the vehicle, the voice data and video data of the same user can be associated, thereby effectively managing and coordinating the needs of different users, avoiding confusion and misunderstanding, and further improving the accuracy of the generated target instruction text.

[0158] Furthermore, in one feasible approach, collecting voice data via a microphone in the vehicle can be achieved through the following process:

[0159] Initial voice data is collected through a microphone in the vehicle, and then noise reduction is performed on the initial voice data to generate voice data.

[0160] Understandably, a vehicle's microphone captures not only human voices but also ambient noise, such as engine noise, traffic noise, and navigation sounds. Background noise interferes with the clarity of the speech signal, making it difficult for subsequent command generation models to accurately recognize human voices, leading to misunderstandings or incorrect outputs. Therefore, noise reduction processing can be applied to the initial speech data to reduce background noise, making the human voice in the resulting speech data clearer and easier for subsequent recognition and understanding of the user's words.

[0161] Furthermore, the specific process of noise reduction for the initial speech data is as follows:

[0162] First, the filter coefficients of the adaptive filter are determined based on the environmental noise measurement or echo signal of the initial speech data. Then, based on the filter coefficients, the initial speech data is denoised using the adaptive filter to generate speech data.

[0163] In one possible implementation, the adaptive filter can use ambient noise measurements or echo signals as reference signals, calculate filter coefficients through the adaptive filter mathematical expression, and continuously iterate to adjust and optimize the filter coefficients to minimize their error. Then, based on the filter coefficients, the initial speech data is filtered to minimize the noise components in the output signal (speech data) or cancel the echo in the speech signal.

[0164] The specific formula for calculating the filter coefficients in an adaptive filter is as follows:

[0165] Where x(n) is the initial speech data; d(n) is the reference signal; w k (n) represents the k-th filter coefficient; μ is the step size parameter, controlling the learning rate; n is the ordinal number, i.e., (1…N); k is the ordinal number, i.e., (1…K); P x is the power of the initial speech data; M is the filter order.

[0166] It is understandable that by adjusting the filter coefficients, the noise component in the output signal (speech data) can be minimized, the echo in the initial speech data can be canceled, the recognizability of human voice can be improved, and the interference of noise on the human voice extraction process can be reduced.

[0167] Figure 3 is a schematic diagram of a scenario of the vehicle control method provided in this embodiment of the invention. As shown in Figure 3, the scenario includes: acquiring initial voice data of the target user through a microphone. It should be noted that the initial voice data includes human voice and ambient noise. Then, an active filter amplifier or a programmable amplifier is used to amplify the acquired initial voice data. It should be understood that this method can increase the intensity of the initial voice data, thereby improving the subsequent processing effect. Further, a low-pass filter is used to perform high-frequency filtering on the amplified initial voice data. Specifically, a low-pass filter can be used to remove high-frequency noise exceeding 20kHz, retaining the main components of the amplified initial voice data to generate filtered initial voice data. It is understood that by removing unnecessary high-frequency noise and retaining the main voice frequency range, the clarity of the human voice is significantly improved, making it easier for the subsequent instruction generation model to extract key information, thereby improving the accuracy of the generated target instruction text.

[0168] Furthermore, the filtered initial speech data is converted from analog to digital using an ADC / A2B converter. It should be understood that since current audio digital signal processors (Audio DSPs) primarily process digital signals, which have advantages such as insensitivity to noise and interference and ease of storage, copying, and processing, analog-to-digital conversion of the filtered initial speech data is necessary.

[0169] Furthermore, the converted initial speech data is input to an audio digital signal processor for processing. The audio digital signal processor uses an adaptive filter to further process the converted initial speech data, including noise reduction or echo cancellation, to generate speech data. The specific process has been described in the above embodiments and will not be repeated here.

[0170] Furthermore, the processed voice data is input to the vehicle's main controller (System on Chip, SOC). It should be noted that the Audio DSP can send the generated voice data to the SOC's hardware abstraction layer via the integrated circuit's built-in audio (Inter-IC Sound, IIS) bus, Pulse Code Modulation (PMC) bus, Time Division Multiplexing (TDM) bus, and Pulse Density Modulation (PDM) bus. The SOC pre-deploys an instruction generation model to generate target instruction text based on the voice and video data.

[0171] Optionally, a System-on-Chip (SOC) typically includes Audio Video Navigation (AVN), a Head Unit (HU), and an In-Vehicle Infotainment (IVI).

[0172] Simultaneously, video data of the target user is captured via a camera. This video data is then processed using a Field-Programmable Gate Array (FPGA) before being input to the vehicle's main controller. The FPGA performs filtering, enhancement, and edge detection on the video data to improve its quality.

[0173] Finally, the SOC acquires the target user's voice and video data, processes the voice and video data using the instruction generation model, and outputs the target instruction text.

[0174] It should be understood that the vehicle is equipped with a communication module that can obtain instructions from the cloud platform to generate models by utilizing fourth-generation mobile communication technology (4G), fifth-generation mobile communication technology (5G), and wireless Fidelity (WIFI).

[0175] Understandably, by inputting the target user's voice and video data into the instruction generation model in this way, the instruction generation model can effectively extract key features and information, identify more accurate user intentions, and generate more accurate target instruction text.

[0176] Next, the structure of the instruction generation model and the process by which the instruction generation model processes data will be explained.

[0177] In some embodiments, the instruction generation model includes a first modal projection multilayer sensing layer, a second modal projection multilayer sensing layer, an embedding layer, and a Transformer network.

[0178] The input of the embedding layer is connected to the output of the first modal projection multilayer sensing layer and the output of the second modal projection multilayer sensing layer, and the output of the embedding layer is connected to the input of the Transformer network.

[0179] The first modal projection multilayer perception layer is used to project the initial audio features corresponding to the speech data onto a preset feature space to generate audio features. The initial audio features include a first initial text sub-feature and a first initial emotion sub-feature. The audio features include a first text sub-feature and a first emotion sub-feature.

[0180] Specifically, the instruction generation model can extract features from the speech data to obtain initial audio features, which include a first initial text sub-feature and a first initial emotion sub-feature. The first initial text sub-feature represents the text used for speech recognition of the speech data, and the first initial emotion sub-feature represents at least one of the tone, intonation, and volume of the speech data.

[0181] For example, tone of voice includes gentleness, sarcasm, etc.; intonation refers to the pitch variation of a user's voice, for example, a high-pitched tone usually conveys excitement, while a low-pitched tone may indicate frustration; volume refers to the loudness of a user's voice, a louder volume may indicate excitement or anger, while a softer volume may indicate shyness or unease.

[0182] The second modal projection multilayer perception layer is used to project the initial image features corresponding to the video data onto a preset feature space to generate image features. The initial image features include the second initial text sub-feature and the second initial emotion sub-feature, and the image features include the second text sub-feature and the second emotion sub-feature.

[0183] Specifically, the instruction generation model can extract features from video data to obtain initial image features, which include a second initial text sub-feature and a second initial emotion sub-feature. The second initial text sub-feature represents text identified based on the target user's lip movements, and the second initial emotion sub-feature represents facial expressions and / or lip changes.

[0184] The embedding layer maps each word to a corresponding word embedding sequence. Then, the positional encoding of each word is added to the word embedding sequence to generate an encoded word embedding sequence. In other words, the words are obtained by segmenting the first text sub-feature, the first emotion sub-feature, the second text sub-feature, and the second emotion sub-feature respectively.

[0185] The Transformer network consists of multiple Transformer layers. It is used to output target instruction text based on the encoded word embedding sequence. The encoded word embedding sequence is obtained by word segmentation and position encoding of audio and image features.

[0186] In this embodiment, data from different modalities (i.e., video data and audio data) are projected onto the same or compatible feature space through a first modality projection multilayer perception layer and a second modality projection multilayer perception layer, allowing them to be aligned and compared within a shared representation space. Simultaneously, both audio and image features include emotion sub-features representing the target user's emotions. This enables the Transformer network to more accurately identify user intent based on the emotion sub-features contained in both audio and image features, resulting in more accurate target instruction text output. The encoder of the Transformer network in this instruction generation model is a shared encoder, allowing video and audio data to share the same network structure and parameters. In other words, compared to constructing separate complex encoders for video and audio data, this technique significantly reduces the total number of parameters in the instruction generation model without affecting processing accuracy.

[0187] Furthermore, the instruction generation model also includes an image-to-audio layer.

[0188] The output of the image-bound audio layer is connected to the input of the first modal projection multilayer perception layer. The image-bound audio layer is used to adjust the feature format of the initial audio features to the feature format of the initial image features.

[0189] In the above embodiments, by binding the image to the audio layer, the data of the two modalities of image and audio are bound or associated, so as to realize cross-modal data fusion and understanding in the future.

[0190] Next, we will explain the structure of the instruction generation model through a specific example.

[0191] Figure 4 is a schematic diagram of the instruction generation model provided in an embodiment of the present invention. As shown in Figure 4, the instruction generation model includes: a contrastive language image layer, a second modal projection multilayer perceptron layer, an audio encoder, a convolutional downsampling layer, an image-bound audio layer, a first modal projection multilayer perceptron layer, an embedding layer, and a Transformer network.

[0192] The Transformer network consists of L stacked Transformer layers.

[0193] The Transformer network consists of an encoder and a decoder. The encoder includes a multi-head attention layer, a feedforward network, and a residual connection + layer normalization. The decoder includes a multi-head attention layer, a feedforward network, a residual connection + layer normalization, and a fully connected layer.

[0194] Among them, the contrastive language image layer is used to extract the initial image features of the video data; the positional encoding is used to represent the positional information of words in the features; the multi-head attention layer is used to achieve parallel computation through multiple attention heads and capture contextual information at different levels; the feedforward neural network is used to enhance the representational ability of the model; the residual connection is used to connect the inputs and outputs at different levels to improve training stability; and the layer normalization is used to normalize the outputs of each layer to improve training efficiency.

[0195] Using the instruction generation model shown in Figure 4, the process of processing data based on the instruction generation model will be explained.

[0196] Initial image features corresponding to video data are acquired through image-to-audio binding. These initial image features are then projected into a preset feature space using a second modal projection multilayer perceptron layer to generate image features. Initial audio features corresponding to speech data are acquired through an audio encoder. These initial audio features are then convolved using a convolutional downsampling layer to reduce data dimensionality, generating convolved initial audio features. Next, an image-to-audio binding layer adjusts the feature format of the convolved initial audio features to match the feature format of the initial image features. Finally, a first modal projection multilayer perceptron layer projects the formatted initial audio features into the preset feature space to generate audio features.

[0197] Furthermore, the audio and image features are segmented into words, specifically the first text sub-feature, the first emotion sub-feature, the second text sub-feature, and the second emotion sub-feature are segmented into words respectively, resulting in multiple words. Then, an embedding layer maps each word to its corresponding word embedding sequence.

[0198] Furthermore, positional encoding is added to each word embedding sequence to generate an encoded word embedding sequence. Then, all encoded word embedding sequences are divided into multiple parts and fed into different attention heads in the encoder's multi-head attention layer. Each attention head calculates self-attention or cross-attention to generate a context representation. The context representations are then concatenated to generate the used context. Next, the used context is concatenated with the corresponding word embedding sequence to generate the target input sequence. The target input sequence is encoded through multiple Transformer layers to generate the target sequence, and finally, the target sequence is decoded by a decoder to generate the target instruction text.

[0199] Specifically, word segmentation is performed on audio and image features, and the resulting segmentation can be represented by the input sequence. For example, if someone says "turn on the air conditioner," the text sub-feature corresponding to the collected speech data is "turn on the air conditioner." The input sequence obtained by segmenting this text sub-feature is {turn on, air conditioner}. In other words, the input sequence can be represented as: X = {x1, x...} 2, , ......, x n,}

[0200] Where X is an input sequence, and x1, x2, ..., xn are words obtained through word segmentation.

[0201] Furthermore, in order for the instruction generation model to better understand the input sequence, it is also necessary to perform an embedding vector operation on the input sequence, that is, to convert the words in the input sequence into vector representations and generate a word embedding sequence, so that the words are transformed from discrete symbolic representations (such as text) whose semantic information is not directly manifested into continuous numerical vectors containing semantic information.

[0202] The word embedding sequence of the input sequence X can be represented as: E = {e1, e2} 2, , ......, e n,}

[0203] Where E is the word embedding sequence of the input sequence X; e n It is the marker x n Word embedding sequence.

[0204] Then, to help the instruction generation model understand the sequence and contextual relationships, the word embedding sequence needs to be positionally encoded, that is, positional information is added to the embedding vector, enabling the instruction generation model to understand the semantics of each token and its position in the input sequence. Finally, the modality projection multilayer perceptron projects data from different modalities (i.e., video data and audio data) onto the same or compatible feature space, allowing them to be aligned and compared in a shared representation space, generating audio and image features, which are then input into the Transformer network. The Transformer network allows each word to pay attention to other words in the input sequence when generating its representation. Through the calculation of attention weights, the instruction generation model can capture the dependencies between words and further process and nonlinearly transform the context vector.

[0205] The decoder in the Transformer layer typically employs a masking mechanism during decoding to ensure that each position only focuses on previous positions, thus achieving sequential generation. It references the hidden representation generated by the encoder to integrate information from the input sequence. Further, the encoded and decoded hidden states are passed through a linear layer to generate a vector the same size as the vocabulary. Each element of these vectors corresponds to a word in the vocabulary, representing the probability of that word's generation. By applying a function (softmax function), the output of the linear layer is converted into a probability distribution, representing the likelihood of each word being selected in the current step. During generation, multiple most probable sequences are maintained, and the sequence with the highest overall probability is ultimately selected.

[0206] For the Transformer layer, the input of the l-th layer is: H (l-1) The input to the first layer of the (word embedding sequence) is: E pos The output formula of the Transformer layer can be expressed as:

[0207] Among them, H enc The encoder output is represented by ReLU, the linear rectified function, and softmax, the normalized exponential function. W... Q W K W VAll are learnable weight matrices; W1, W2, b1, b2 are learnable parameters; d k The dimension of the key is used to scale the dot product.

[0208] The output formula for a modal projection multilayer sensing decoder can be expressed as:

[0209] Among them, H (dec) The output of the decoder; FFN stands for Feed-Forward Network; CrossAttention is the cross-attention mechanism; SelfAttention is the self-attention mechanism. For the target sequence Y shifted-right Input.

[0210] The probability formula for mapping a word to a word in the vocabulary through a linear layer can be expressed as: P(y t |y <x,X)=softmax(H dec W out )

[0211] Where P is the probability of the linear layer mapping to each word in the vocabulary; t is the time step; y is the target word; y t For the target word corresponding to t; W out This is the weight matrix at the output.

[0212] Optionally, in some embodiments, the model can be trained through a cloud platform to obtain an instruction generation model, and then the instruction generation model can be deployed to the vehicle so that the vehicle can generate target instruction text locally through the instruction generation model.

[0213] Specifically, Figure 5 is a schematic flowchart of the vehicle control method provided in this embodiment of the invention. As shown in Figure 5, this vehicle control method is applied to a cloud platform and includes:

[0214] S51. Train the initial instruction generation model by using sample voice data, sample video data and labels from sample users.

[0215] During the model training phase, sample voice data and sample video data can be collected in real time from sample users driving sample vehicles via microphone and camera, respectively. Then, based on the sample users' control intentions while driving the vehicle, labels can be generated. These labels represent sample control command text for controlling the vehicle, including the vehicle's specific functions and parameters.

[0216] For example, while a sample user is driving the sample vehicle, if the user feels hot, they can say "It's too hot." At this point, the user's voice data is collected in real-time via microphone, and video data is collected in real-time via camera. The user can then activate relevant functions and adjust their parameters based on their current feeling until they feel comfortable. For example, the user can turn on the air conditioning and adjust the temperature (e.g., 24°C) until they no longer feel hot. Finally, a label is assigned based on the user's adjustments. For example, the label could be "Air conditioning on, adjusted to 24°C."

[0217] Then, training, testing, and validation sets can be established based on sample speech data, sample video data, and labels. The model can then be trained using these sets to obtain initial instructions for model generation.

[0218] Specifically, the initial model is trained based on the training set, and the model parameters of the initial model are continuously optimized until the training cutoff condition is met. The current initial model is then determined as the initial instruction generation model.

[0219] Before model training, the shared encoder of the Transformer network in the initial model is pre-trained on a large-scale corpus. During pre-training, the shared encoder learns general knowledge of language syntax and semantics. Thus, after the shared encoder is pre-trained, the parameters of this shared encoding part can be frozen or fine-tuned. In actual training, only minor adjustments to the initial model based on the training set are needed to achieve the training objective, greatly reducing the training time and improving training efficiency.

[0220] It should be noted that the instruction generation model also has other functions, such as speech translation, which translates the recognized speech content from one language to another; intent classification, which identifies the user's intent by analyzing the user's speech input; and slot filling, which extracts key information (such as time, location, people, etc.) from the user's speech and fills it into predefined slots.

[0221] During the training of the instruction generation model, it can also be trained for each specific task. For example, during training, some words in the input text are randomly masked, and the instruction generation model needs to predict these masked words. The goal is to minimize the cross-entropy loss between the predicted sequence and the target sequence. The specific formula for the cross-entropy loss is:

[0222] Where L represents the cross-entropy loss between the predicted and target sequences; P represents the probability of the linear layer mapping to each word in the vocabulary; t is the time step; y represents the target word; y tThe target word is the word corresponding to t.

[0223] S52. Reduce the weights of the initial instruction generation model and / or the representation accuracy of the output of the activation function.

[0224] In one possible implementation, the data format of the weights in the initial instruction generation model is reduced from 32-bit floating-point numbers to 16-bit; and / or, the data format of the output of the activation function is reduced from 32-bit floating-point numbers to 16-bit, thereby reducing the size and computational burden of the initial instruction generation model.

[0225] Reducing the representation precision of the weights in the initial instruction generation model significantly reduces the memory space required to store the weights. Similarly, reducing the representation precision of the activation function's output can also reduce the amount of data involved in the computation, thereby reducing the computational resources required for model execution and accelerating computation.

[0226] S53. Delete redundant connections and / or redundant neurons in the initial instruction generation model.

[0227] By removing redundant connections and / or redundant neurons, inference time can be reduced, thus reducing computational cost while maintaining model performance.

[0228] Redundant connections and neurons participate in computation during both forward and backward propagation of the model. Removing them significantly reduces the computational cost when calculating the model output during forward propagation. However, connections and neurons in the model require memory to store related parameters (such as connection weights). The presence of redundant connections and neurons increases the model's memory footprint. Therefore, removing redundant connections and / or redundant neurons generated from the initial instructions simplifies the model structure and reduces its memory usage.

[0229] S54. Transform the initial instruction generation model to the inference architecture supported by the vehicle to generate an instruction generation model.

[0230] Because the inference architecture supported by the cloud platform is different from that supported by the vehicle, the model trained on the cloud platform cannot be directly deployed to the vehicle and needs to be converted first.

[0231] Optionally, the inference architecture supported by the vehicle can be the Open Neural Network Exchange (ONNX) format.

[0232] Optionally, before conversion, the knowledge generated by the initial instruction model can be compressed into a smaller model, maintaining accuracy while being more efficient.

[0233] S55. Deploy the instruction generation model to the vehicle.

[0234] The instruction generation model is deployed to the vehicle and loaded by the application to generate the target instruction text. During inference, the multi-core architecture of the vehicle's SoC chip can be used to process the input data in parallel.

[0235] In the above embodiments, leveraging the powerful computing capabilities of the cloud platform, a high-accuracy initial command generation model is pre-trained using sample voice data, sample video data, and labels from sample users to generate and store the model. Subsequently, various operations are performed on the initial command generation model to reduce its complexity, and it is then converted into a vehicle-supported inference architecture to generate a new command generation model, which is then deployed to the vehicle. In this way, the vehicle does not need to perform complex model training locally, thus saving computing resources and time. Furthermore, the cloud platform can utilize its platform advantages to generate a higher-quality command generation model, enabling it to have better generalization ability and accuracy. During model inference, after the vehicle collects voice and video data, it can directly process the voice and video data on the vehicle side using the command generation model to control the vehicle, saving the interaction process with the cloud platform and improving processing efficiency.

[0236] Optionally, after the instruction generation model is deployed to the vehicle on the cloud platform, the vehicle can also preload the instruction generation model into the vehicle's memory to reduce the loading time of the instruction generation model and improve the inference efficiency of the target instruction model.

[0237] In some scenarios, vehicles can send model retrieval requests to the cloud platform. These requests are used to request instructions from the cloud platform to generate a model. Upon receiving the model retrieval request, the cloud platform will generate the model from the instructions and deploy it to the vehicle.

[0238] Furthermore, the cloud platform can continuously iterate and update the instruction generation model to ensure the accuracy of instruction generation. Therefore, the instruction generation model deployed on the vehicle also needs to be updated through the cloud platform. Specifically, after iteratively updating the instruction generation model, the cloud platform can send model update information to the vehicle. This information indicates that the version of the instruction generation model stored in the cloud platform is higher than the version deployed in the vehicle. Upon receiving the model update information, the vehicle sends a model update request to the cloud platform to retrieve the higher version of the instruction generation model stored in the cloud platform. After receiving the model update request, the cloud platform deploys the higher version of the instruction generation model to the vehicle.

[0239] In one possible implementation, the cloud platform can send the latest version (i.e., the aforementioned higher version) of the model parameters to the vehicle via Over-The-Air (OTA) technology, thereby enabling the deployment of the higher version instruction-generated model into the vehicle.

[0240] Understandably, this approach leverages the continuous update capabilities of the cloud platform to ensure that the vehicle always uses the latest instruction generation model, thereby guaranteeing the accuracy of the target instructions generated by the model. Furthermore, the vehicle does not require frequent local updates and maintenance, reducing operating and management costs and significantly improving the performance of intelligent vehicles while enhancing the user experience.

[0241] Based on the specific implementation of the vehicle control method in the above embodiments, and to facilitate a clearer understanding of the entire software system, the software system will be described in detail below.

[0242] Figure 6 is a schematic diagram of the vehicle software system provided in an embodiment of the present invention; as shown in Figure 6, the software system includes a cloud platform side and a vehicle side.

[0243] The cloud platform has functions for generating instruction models for training and downloading models.

[0244] The vehicle includes an audio application layer, an audio service layer, an audio framework layer, a hardware abstraction layer (HAL), an audio DSP, a low-pass filter, and a signal processing layer.

[0245] In the signal processing layer, initial voice data of the target user in the vehicle is first acquired. Then, the acquired audio signal is amplified using an active filter amplifier or a programmable amplifier. After amplification, noise processing is performed, employing a low-pass filter to remove high-frequency noise exceeding 20kHz. Next, the Audio DSP uses an adaptive filter to denoise the initial voice data, generating voice data, which is then input to the HAL via a bus. Correspondingly, initial video data of the target user in the vehicle is acquired, generated through image processing, and then input to the Hardware Abstraction Layer (HAL).

[0246] In the hardware abstraction layer, video and audio data are processed by an instruction generation model to generate target control text, which is then used to control the vehicle. Optionally, the instruction generation model can also perform speech translation, intent classification, slot filling, and emotion recognition, and the output is fed into the audio application layer for further processing through the audio framework layer and audio service layer.

[0247] Figure 7 is a schematic diagram of a second scenario of the vehicle control method provided in this embodiment of the invention. As shown in Figure 7, this scenario includes a cloud platform and a vehicle. The cloud platform needs to pre-deploy the instruction generation model to the vehicle. In practical applications, the vehicle acquires the target user's voice data and video data, and then amplifies the voice data using an active filter amplifier or a programmable amplifier. After signal amplification, noise processing is performed, using a low-pass filter to remove high-frequency noise exceeding 20kHz while retaining the main components of the voice signal. Further, an adaptive filter is used to denoise the initial voice data to generate voice data. Then, the generated voice data and video data are input into the instruction generation model to obtain the target instruction text output by the instruction generation model. Finally, the vehicle is controlled using the target instruction text.

[0248] It should be noted that the specific implementation process is illustrated in the above embodiments, and the embodiments of the present invention will not be described in detail here.

[0249] Figure 8 is a schematic diagram of the vehicle control device provided in an embodiment of the present invention; as shown in Figure 8, the vehicle control device 80 includes:

[0250] The acquisition module 81 is used to acquire the voice data and video data of the target user in the vehicle;

[0251] The input module 82 is used to input voice data and video data into the instruction generation model and obtain the target instruction text output by the instruction generation model. The instruction generation model is used to fuse the audio features corresponding to the voice data and the image features corresponding to the video data to generate the target instruction text. The audio features include a first text sub-feature and a first emotion sub-feature, and the image features include a second text sub-feature and a second emotion sub-feature. The first emotion sub-feature is used to represent at least one of the target user's tone, pitch, and volume, and the second emotion sub-feature is used to represent the target user's facial expression and / or lip changes.

[0252] Control module 83 is used to control the vehicle via target command text.

[0253] Furthermore, in the instruction generation model, the weight of the first emotion sub-feature is greater than the weight of the second emotion sub-feature.

[0254] Furthermore, module 81 is specifically used for:

[0255] Voice data is collected via microphones in the vehicle, and initial video data is collected via cameras in the vehicle.

[0256] Based on the voice data, determine the target location of the target user in the vehicle corresponding to the voice data;

[0257] The initial video data is processed based on the target location to obtain the video data corresponding to the target user.

[0258] Furthermore, module 81 is specifically used for:

[0259] Initial voice data was collected using a microphone in the vehicle;

[0260] The initial speech data is denoised to generate speech data.

[0261] Furthermore, module 81 is specifically used for:

[0262] The filter coefficients of the adaptive filter are determined based on the environmental noise measurement or echo signal of the initial speech data.

[0263] Based on the filter coefficients, an adaptive filter is used to denoise the initial speech data to generate speech data.

[0264] Furthermore, the instruction generation model includes:

[0265] The first modal projection multilayer sensing layer, the second modal projection multilayer sensing layer, and the Transformer network;

[0266] The input of the Transformer network is connected to the output of the first modal projection and the output of the second modal projection, respectively.

[0267] The first modal projection multilayer perception layer is used to project the initial audio features corresponding to the speech data onto a preset feature space to generate audio features. The initial audio features include a first initial text sub-feature and a first initial emotion sub-feature.

[0268] The second modal projection multilayer perception layer is used to project the initial image features corresponding to the video data onto a preset feature space to generate image features. The initial image features include the second initial text sub-feature and the second initial emotion sub-feature.

[0269] The Transformer network is used to output target instruction text based on the encoded word embedding sequence, which is obtained by word segmentation and position encoding of audio and image features.

[0270] Furthermore, the instruction generation model also includes an embedding layer, the input of which is connected to the output of the first modal projection multilayer perception layer and the output of the second modal projection multilayer perception layer, respectively, and the output of the embedding layer is connected to the input of the Transformer network.

[0271] The embedding layer is used to map each word to a corresponding word embedding sequence;

[0272] The positional encoding corresponding to each word is added to the word embedding sequence corresponding to the word to generate the encoded word embedding sequence.

[0273] Furthermore, the instruction generation model also includes:

[0274] Image-to-audio layer binding;

[0275] The output of the image-bound audio layer is connected to the input of the first modal projection multilayer perception layer. The image-bound audio layer is used to adjust the feature format of the initial audio features to the feature format of the initial image features.

[0276] Furthermore, before acquiring the voice and video data of the target user in the vehicle, the vehicle control device 80 also includes a preloading module for preloading the instruction generation model into the vehicle's memory.

[0277] The vehicle control method apparatus provided in this embodiment of the invention can be used to execute the vehicle control method on the vehicle side in any of the above embodiments. Its implementation principle and technical effect are similar, and will not be described again here.

[0278] Figure 9 is a second structural schematic diagram of the vehicle control device provided in an embodiment of the present invention; as shown in Figure 9, the vehicle control device 80 further includes a deployment module 84, used for:

[0279] The model is trained by using sample voice data, sample video data, and tags from sample users, and the initial instructions are obtained and stored.

[0280] The initial instruction generation model is converted to a vehicle-supported inference architecture to generate an instruction generation model;

[0281] Deploy the instruction generation model to the vehicle.

[0282] Furthermore, the deployment module 84 is also used to reduce the weights of the initial instruction generation model and / or the representation accuracy of the output of the activation function.

[0283] Furthermore, deployment module 84 is also used to remove redundant connections and / or redundant neurons from the initial instruction generation model.

[0284] The vehicle control method apparatus provided in this embodiment of the invention can be used to execute the vehicle control method on the cloud platform side in any of the above embodiments. Its implementation principle and technical effect are similar, and will not be described again here.

[0285] It should be noted that the division of the various modules in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, these modules can be implemented entirely in software through processing element calls, or entirely in hardware. Alternatively, some modules can be implemented through processing element calls in software, while others are implemented in hardware. Moreover, these modules can be fully or partially integrated together, or implemented independently. The processing element here can be an integrated circuit with signal processing capabilities. During implementation, each step of the above method or each of the above modules can be completed through the integrated logic circuits in the hardware of the processor element or through software instructions.

[0286] Figure 10 is a second schematic diagram of the structure of a vehicle provided in an embodiment of the present invention. As shown in Figure 10, the vehicle 100 may include: a processor 101, a memory 102, and computer execution instructions stored in the memory 102 and executable on the processor 101. When the processor 101 executes the computer execution instructions, it implements the vehicle control method on the vehicle side provided in any of the foregoing embodiments.

[0287] Optionally, the various devices in the vehicle 100 can be connected via a system bus.

[0288] The memory 102 can be a separate memory unit or a memory unit integrated into the processor. The number of processors can be one or more.

[0289] Optionally, vehicle 100 may also include a communication interface for interacting with other devices.

[0290] The vehicle provided in this embodiment of the invention can be used to execute the vehicle control method on the vehicle side provided in any of the above method embodiments. Its implementation principle and technical effect are similar, and will not be repeated here.

[0291] Figure 11 is a schematic diagram of the structure of the cloud platform provided in an embodiment of the present invention. As shown in Figure 11, the cloud platform 110 may include: a processor 111, a memory 112, and computer execution instructions stored in the memory 112 and executable on the processor 111. When the processor 111 executes the computer execution instructions, it implements the vehicle control method on the cloud platform side provided in any of the foregoing embodiments.

[0292] Optionally, the various devices in the cloud platform 110 can be connected via a system bus.

[0293] The memory 112 can be a separate memory unit or a memory unit integrated into the processor. The number of processors can be one or more.

[0294] Optionally, the cloud platform 110 may also include a communication interface for interacting with other devices.

[0295] The cloud platform provided in this embodiment of the invention can be used to execute the vehicle control method on the cloud platform side provided in any of the above method embodiments. Its implementation principle and technical effect are similar, and will not be described again here.

[0296] It should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0297] The system bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The system bus can be divided into address bus, data bus, control bus, etc. For ease of representation, only one thick line is used in the diagram, but this does not indicate that there is only one bus or one type of bus. Memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0298] All or part of the steps in the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a readable memory. When the program is executed, it performs the steps of the above-described method embodiments. The aforementioned memory (storage medium) includes: read-only memory (ROM), RAM, flash memory, hard disk, solid-state drive, magnetic tape, floppy disk, optical disk, and any combination thereof.

[0299] This invention provides a computer-readable storage medium storing computer-executable instructions that, when executed on a computer, cause the computer to perform the aforementioned vehicle control method.

[0300] The aforementioned computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory, electrically erasable programmable read-only memory, erasable programmable read-only memory, programmable read-only memory, read-only memory, magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0301] Optionally, a readable storage medium can be coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Alternatively, the readable storage medium can be an integral part of the processor. Both the processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components within the device.

[0302] This invention also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium, and when the at least one processor executes the computer program, it can implement the above-described vehicle control method.

[0303] This embodiment also provides a chip, which includes a memory and a processor. The memory stores code and data, and the memory is coupled to the processor. The processor runs the program in the memory so that the chip can be used to execute the vehicle control methods provided in the above-described embodiments.

[0304] This embodiment also provides a computer program, which, when executed by a processor, is used to perform the vehicle control methods provided in the foregoing embodiments.

[0305] It should be understood that the present invention is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. A vehicle control method, characterized in that, include: Acquire voice and video data of the target user in the vehicle; The voice data and video data are input into the instruction generation model to obtain the target instruction text output by the instruction generation model. The instruction generation model is used to fuse the audio features corresponding to the voice data and the image features corresponding to the video data to generate the target instruction text. The audio features include a first text sub-feature and a first emotion sub-feature, and the image features include a second text sub-feature and a second emotion sub-feature. The first emotion sub-feature is used to represent at least one of the target user's tone, intonation, and volume, and the second emotion sub-feature is used to represent the target user's facial expression and / or lip changes. The vehicle is controlled by the target instruction text.

2. The method according to claim 1, characterized in that, In the instruction generation model, the weight of the first emotion sub-feature is greater than the weight of the second emotion sub-feature.

3. The method according to claim 1 or 2, characterized in that, The acquisition of voice and video data of the target user in the vehicle includes: The voice data is collected through the microphone in the vehicle, and the initial video data is collected through the camera in the vehicle; Based on the voice data, determine the target location of the target user corresponding to the voice data in the vehicle; The initial video data is processed based on the target location to obtain the video data corresponding to the target user.

4. The method according to claim 3, characterized in that, The process of collecting the voice data via a microphone in the vehicle includes: Initial voice data is collected using a microphone in the vehicle; The initial speech data is subjected to noise reduction processing to generate the speech data.

5. The method according to claim 4, characterized in that, The step of performing noise reduction processing on the initial speech data to generate the speech data includes: The filter coefficients of the adaptive filter are determined based on the environmental noise measurement or echo signal of the initial speech data. Based on the filter coefficients, the initial speech data is denoised using the adaptive filter to generate the speech data.

6. The method according to any one of claims 1-5, characterized in that, The instruction generation model includes: The first modal projection multilayer sensing layer, the second modal projection multilayer sensing layer, and the Transformer network; The input of the Transformer network is connected to the output of the first modal projection and the output of the second modal projection, respectively. The first modal projection multilayer perception layer is used to project the initial audio features corresponding to the speech data onto a preset feature space to generate audio features. The initial audio features include a first initial text sub-feature and a first initial emotion sub-feature. The second modal projection multilayer perception layer is used to project the initial image features corresponding to the video data onto the preset feature space to generate image features. The initial image features include a second initial text sub-feature and a second initial emotion sub-feature. The Transformer network is used to output the target instruction text based on the encoded word embedding sequence, which is obtained by word segmentation and position encoding of the audio features and the image features.

7. The method according to claim 6, characterized in that, The instruction generation model further includes an embedding layer, the input of which is connected to the output of the first modal projection multilayer sensing layer and the output of the second modal projection multilayer sensing layer, respectively, and the output of which is connected to the input of the Transformer network. The embedding layer is used to map each word to a corresponding word embedding sequence; The positional encoding corresponding to each word is added to the word embedding sequence corresponding to the word to generate the encoded word embedding sequence.

8. The method according to claim 6 or 7, characterized in that, The instruction generation model also includes an image-bound audio layer; The output of the image-bound audio layer is connected to the input of the first modal projection multilayer perception layer. The image-bound audio layer is used to adjust the feature format of the initial audio features to the feature format of the initial image features.

9. The method according to any one of claims 1-8, characterized in that, Before acquiring the voice and video data of the target user in the vehicle, the method further includes: The instruction generation model is preloaded into the vehicle's memory.

10. The method according to any one of claims 1-8, characterized in that, Before acquiring the voice and video data of the target user in the vehicle, the method further includes: The initial instruction generation model is obtained and stored by training the sample voice data, sample video data and tags of the sample users. The initial instruction generation model is converted to the inference architecture supported by the vehicle to generate the instruction generation model; The instruction generation model is deployed to the vehicle.

11. The method according to claim 10, characterized in that, The method further includes: Reduce the representation accuracy of the weights and / or the output of the activation function of the initial instruction generation model.

12. The method according to claim 10 or 11, characterized in that, The method further includes: Delete redundant connections and / or redundant neurons in the initial instruction generation model.

13. A vehicle control device, characterized in that, include: The acquisition module is used to acquire voice and video data of the target user in the vehicle. An input module is used to input the voice data and the video data into an instruction generation model, and obtain the target instruction text output by the instruction generation model. The instruction generation model is used to fuse the audio features corresponding to the voice data and the image features corresponding to the video data to generate the target instruction text. The audio features include a first text sub-feature and a first emotion sub-feature, and the image features include a second text sub-feature and a second emotion sub-feature. The first emotion sub-feature is used to represent at least one of the target user's tone, intonation, and volume, and the second emotion sub-feature is used to represent the target user's facial expression and / or lip changes. The control module is used to control the vehicle via the target instruction text.

14. A vehicle comprising: A processor, a memory, and computer-executable instructions stored in the memory and executable on the processor, characterized in that the processor executes the computer-executable instructions to implement the method as claimed in any one of claims 1 to 9.

15. A cloud platform, comprising: A processor, a memory, and computer-executable instructions stored in the memory and executable on the processor, characterized in that the processor executes the computer-executable instructions to implement the method as claimed in any one of claims 1 to 12.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1 to 12.

17. A chip, characterized in that, The chip includes a memory and a processor. The memory stores code and data and is coupled to the processor. The processor runs a program in the memory that causes the chip to perform the method described in any one of claims 1 to 12.

18. A program product, characterized in that, include: A computer program that, when the program product is run on a computer, causes the computer to perform the method described in any one of claims 1 to 12.

19. A computer program, characterized in that, When the computer program is executed by a processor, it is used to perform the method described in any one of claims 1 to 12.