Vehicle control method, device and equipment based on emotion recognition and storage medium
By monitoring in-vehicle environmental information in real time and dynamically adjusting the weights of facial motion units, environmental interference is eliminated, and verification is performed using voice and physiological signals. This solves the problem of inaccurate emotion recognition in vehicles, improving user experience and interaction reliability.
Patent Information
- Application Number
- CN202511607315.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-04
- Publication Date
- 2026-02-27
AI Technical Summary
In existing technologies, relying solely on the biosignals exhibited by the user cannot effectively distinguish whether facial movements are physiological reactions caused by environmental discomfort or facial expressions driven by genuine emotions, resulting in inaccurate interactive feedback and affecting user experience.
By monitoring in-vehicle environment information in real time to form an environment vector, the weight parameters of facial action units in the emotion recognition model are dynamically adjusted to remove environmental interference, obtain the emotional state of the driver and passengers after removing in-vehicle environmental interference, and combine voice and physiological signs for auxiliary verification to trigger corresponding vehicle control operations.
It significantly improves the accuracy of emotion recognition and the precision of interactive feedback, thereby enhancing the reliability of vehicle human-machine interaction and user experience.
Smart Images

Figure CN121572993A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle technology, and more specifically to a vehicle control method, device, equipment, and storage medium based on emotion recognition. Background Technology
[0002] As the level of automotive intelligence continues to improve, the human-machine interaction systems of modern vehicles are evolving from simple function execution to a higher level of understanding user intentions and proactively providing services. Based on this, emotion recognition technology, as the core of realizing emotional interaction in intelligent cockpits, is receiving increasing attention. Emotion recognition technology refers to the technology that, by sensing and interpreting the emotional state of the driver or passengers in real time, allows vehicles to automatically adjust the environment, recommend content, or provide comfort, thereby significantly improving driving comfort and safety.
[0003] In related technologies, in-vehicle cameras capture facial images of users, and then pre-trained computer vision algorithms are used to extract and classify key feature points in facial expressions (such as changes in eyebrows and corners of the mouth), which are then mapped to specific emotion categories, such as happiness, calmness, or anger. The core basis for its analysis and judgment comes directly from the biosignals exhibited by the user.
[0004] However, because the interior of a car is a dynamically changing physical environment, temperature, air quality, and even odors can directly trigger unconscious physiological reactions in passengers, such as frowning or covering their noses. Therefore, relying solely on the user's own biosignals cannot effectively distinguish whether these facial movements are physiological reactions caused by environmental discomfort or facial expressions driven by genuine emotions. This may lead to inaccurate or even offensive interactive feedback, severely hindering the improvement of the user experience.
[0005] It should be noted that the information disclosed in the background section of this application is intended only to enhance the understanding of the general background of this application, and should not be construed as an admission or in any way implying that the information constitutes prior art known to those skilled in the art. Summary of the Invention
[0006] In view of this, this application provides a vehicle control method, device, equipment and storage medium based on emotion recognition, which helps to solve the problem that the existing technology cannot effectively distinguish whether facial movements are physiological reactions caused by environmental discomfort or facial expressions driven by real emotions by relying solely on the biosignals exhibited by the user, which may trigger inaccurate or even offensive interactive feedback, seriously restricting the improvement of user experience.
[0007] In a first aspect, embodiments of this application provide a vehicle control method based on emotion recognition, comprising: Determine an environmental vector based on the acquired in-vehicle environmental information, where the environmental vector is a vector representing the comprehensive state of the in-vehicle environment; Dynamically adjust the weight parameters of each facial action unit in the emotion recognition model according to the environmental vector; Input the facial image features of the driver and passengers at the current moment into the adjusted emotion recognition model to obtain the emotion state of the driver and passengers after removing the in-vehicle environment interference; Trigger a vehicle control operation corresponding to the emotion state according to the emotion state of the driver and passengers after removing the in-vehicle environment interference.
[0008] In the embodiment of the present application, by real-time monitoring the in-vehicle environmental information to form an environmental vector, the comprehensive state of the in-vehicle environment can be quantitatively characterized; based on this environmental vector, dynamically adjusting the weight parameters of the facial action units can effectively reduce the weight of the physiological reactions caused by environmental discomfort in emotion recognition; this targeted parameter correction makes the finally obtained emotion state a more real emotional reflection after removing environmental interference, significantly improving the accuracy of emotion recognition. On this basis, the system triggers corresponding vehicle control operations according to this purified emotion state, ensuring the accuracy of the interaction feedback and improving the reliability and user experience of vehicle human-computer interaction.
[0009] In a possible implementation manner, the in-vehicle environmental information includes: in-vehicle temperature information, air quality information, and in-vehicle odor information; the determining an environmental vector according to the acquired in-vehicle environmental information includes: Obtain the in-vehicle temperature information, air quality information, and in-vehicle odor information; Perform data preprocessing on the in-vehicle temperature information, the air quality information, and the in-vehicle odor information respectively to generate corresponding standardized in-vehicle temperature values, standardized air quality values, and standardized in-vehicle odor values; Combine the standardized in-vehicle temperature values, the standardized air quality values, and the standardized in-vehicle odor values to form an environmental vector.
[0010] In the embodiment of the present application, by uniformly preprocessing and standardizing the three key environmental information of in-vehicle temperature, air quality, and odor, the differences in the magnitudes of different sensor data are effectively eliminated. At the same time, combining the standardized in-vehicle temperature values, standardized air quality values, and standardized in-vehicle odor values to form an environmental vector provides clear and comprehensive environmental information for the emotion recognition model, laying a reliable data foundation for accurately adjusting the weight of facial action units subsequently, and finally significantly improving the accuracy of emotion state recognition in complex environments.
[0011] In one possible implementation, the data preprocessing includes: data filtering, data fusion, and data standardization.
[0012] In this embodiment, the quality and reliability of the environmental vector are systematically improved through the collaborative processing of filtering, fusion, and standardization of in-vehicle temperature, air quality, and odor information. It can be understood that data filtering effectively eliminates instantaneous noise and abnormal fluctuations during sensor acquisition, ensuring data stability; data fusion integrates the complementarity of multi-source heterogeneous environmental information, generating more representative comprehensive environmental indicators; and data standardization unifies the scale and range of different physical quantity data, making it a standardized input that the emotion recognition model can directly process. These preprocessing operations collectively provide a high-quality, consistent environmental information foundation for subsequent model parameter adjustments.
[0013] In one possible implementation, dynamically adjusting the weight parameters of each facial action unit in the emotion recognition model based on the environment vector includes: According to the formula: The weight parameters of each facial action unit in the emotion recognition model are determined and dynamically adjusted. in, To adjust the weight parameters of the i-th facial motion unit, To adjust the weight parameters of the i-th facial motion unit, , , These are the environment vectors of the i-th facial action unit. , , Standardized values of interior temperature in CRRC Standardized air quality values Standardized values of in-vehicle odor The sensitivity coefficient.
[0014] In this embodiment, a mathematical formula based on a sensitivity coefficient is introduced to accurately map the quantitative information of the environmental vector to the parameter adjustment process of the emotion recognition model. This formula establishes a quantitative influence relationship between each facial action unit and different environmental factors, enabling the system to specifically adjust the weights of each facial action unit according to the specific environmental state, significantly improving the accuracy of the emotion recognition model.
[0015] In one possible implementation, the step of inputting the facial image features of the driver and passengers at the current moment into the adjusted emotion recognition model to obtain the emotional state of the driver and passengers after removing interference from the in-vehicle environment includes: Based on the facial image features of the driver and passengers at the current moment, the intensity values corresponding to multiple facial action units are obtained; The intensity values corresponding to the facial action units are weighted and calculated based on the adjusted weight parameters to obtain the emotion score. Based on the emotional score, the emotional state of the driver and passengers is obtained after removing interference from the in-vehicle environment.
[0016] In this embodiment, by weighting the collected facial motion unit intensity value with the weight parameters adjusted by the environment vector, the emotion score can accurately reflect the true emotional tendency of the driver and passengers after removing environmental interference, thereby ensuring that the final obtained emotional state is a reliable result based on the purified emotional signal, and effectively improving the reliability of the emotional state output.
[0017] In one possible implementation, triggering a vehicle control operation corresponding to the emotional state of the driver and passengers, after removing interference from the in-vehicle environment, includes: Based on the emotional state of the driver and passengers after removing interference from the in-vehicle environment, and combined with the in-vehicle environment information, a fusion decision is performed; Based on the results of the fusion decision operation, corresponding control commands are generated to trigger one or more vehicle interaction devices to work together.
[0018] In this embodiment of the application, by fusing emotional state with real-time in-vehicle environmental information after removing environmental interference, the vehicle control operation can simultaneously respond to the real emotional needs of the driver and passengers and the specific environmental conditions, avoiding the one-sidedness that may result from relying on a single information source, thereby generating more comprehensive and accurate control commands and providing users with a better interactive experience.
[0019] In one possible implementation, after obtaining the emotional state of the driver and passengers after removing in-vehicle environmental disturbances, the method further includes: Acquire the voice signals and / or physiological signs of the driver and passengers; Based on the voice signal and / or physiological sign signal, the emotional state after removing interference from the in-vehicle environment is verified.
[0020] In this embodiment, by introducing voice signals and physiological signs as independent information sources, the emotional state derived from facial expression analysis is reconfirmed and calibrated, thereby effectively correcting the recognition bias caused by individual expression habits and ultimately significantly improving the reliability of the emotional state judgment results.
[0021] Secondly, this application provides a vehicle control device based on emotion recognition, comprising: The environment vector determination module is used to determine the environment vector based on the acquired in-vehicle environment information, wherein the environment vector is a vector characterizing the overall state of the in-vehicle environment; The weight parameter adjustment module is used to dynamically adjust the weight parameters of each facial action unit in the emotion recognition model according to the environment vector. The emotion state acquisition module is used to input the facial image features of the driver and passengers at the current moment into the adjusted emotion recognition model to obtain the emotion state of the driver and passengers after removing interference from the in-vehicle environment. The vehicle control operation triggering module is used to trigger vehicle control operations corresponding to the emotional state of the driver and passengers after removing interference from the in-vehicle environment.
[0022] Thirdly, this application provides an electronic device, comprising: processor; Memory; And a computer program, wherein the computer program is stored in the memory, the computer program including instructions that, when executed by the processor, cause the electronic device to perform the method described in any one of the first aspects.
[0023] Fourthly, this application provides a computer-readable storage medium including a stored program, wherein, when the program is executed, it controls the device on which the computer-readable storage medium is located to perform the method described in any one of the first aspects.
[0024] Understandably, the emotion-based vehicle control device provided in the second aspect, the electronic device provided in the third aspect, and the computer-readable storage medium provided in the fourth aspect are all used to perform some or all of the methods provided in this application. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods, and will not be repeated here. Attached Figure Description
[0025] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is a schematic diagram of an application scenario provided by an embodiment of this application.
[0027] Figure 2 This is a flowchart illustrating a vehicle control method based on emotion recognition, provided as an embodiment of this application.
[0028] Figure 3 This is a schematic diagram of the structure of a vehicle control device based on emotion recognition, provided as an embodiment of this application.
[0029] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0030] To better understand the technical solution of this application, the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0031] It should be understood that the described embodiments are merely some, not all, of the embodiments in this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.
[0032] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. The singular forms “a,” “the,” and “the” used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0033] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0034] Emotion recognition, a cutting-edge field at the intersection of artificial intelligence and human psychology, aims to infer an individual's inner emotional state by analyzing their physiological characteristics and behavioral patterns. In the specific application scenario of the smart cockpit, emotion recognition technology primarily relies on computer vision technology. It captures facial expression images of occupants (especially the driver) through in-vehicle cameras, extracting feature data from key facial movement units such as raised corners of the mouth, furrowed eyebrows, and widened eyes. These subtle facial muscle movement patterns are highly correlated with several basic human emotions (such as happiness, surprise, anger, and disgust). The system analyzes and classifies these features using a pre-trained emotion recognition model, ultimately outputting a judgment of the current emotional state of the occupants. For ease of understanding, a specific application scenario is first illustrated below.
[0035] See Figure 1 This is a schematic diagram illustrating an application scenario provided by an embodiment of this application. For example... Figure 1As shown, when the emotion recognition system identifies features such as "frowning" and "nose wrinkling" from a driver's facial image, it can determine that the driver is in an unhappy state based on the mapping between key feature points and emotion categories in the emotion recognition system. At this time, the emotion recognition system automatically executes vehicle control operations related to "unhappiness." For example, Figure 1 The vehicle's center console screen displays "The car senses you're a little unhappy. Listen to a beautiful melody to relax," and controls the playback of related music.
[0036] However, because the interior of a car is a dynamically changing physical environment, temperature, air quality, and even odors can directly trigger unconscious physiological reactions in passengers, such as frowning or covering their noses. Therefore, relying solely on the user's own biosignals cannot effectively distinguish whether these facial movements are physiological reactions caused by environmental discomfort or facial expressions driven by genuine emotions. This may lead to inaccurate or even offensive interactive feedback, severely hindering the improvement of the user experience.
[0037] For example, when a vehicle emits an odor, the driver may also exhibit facial expressions such as "frowning" and "wrinkling the nose." In this case, if the emotion recognition system identifies the driver as unhappy based on these facial expressions and automatically performs vehicle control operations related to "unhappiness," it may result in unpleasant interactive feedback that negatively impacts the user experience.
[0038] To address the aforementioned issues, this embodiment of the application uses real-time monitoring of the in-vehicle environment to form an environmental vector, thereby quantifying the overall state of the in-vehicle environment. Based on this environmental vector, the weight parameters of the facial motion unit are dynamically adjusted, effectively reducing the weight of physiological reactions caused by environmental discomfort in emotion recognition. This targeted parameter correction ensures that the final emotional state is a more authentic emotional reflection stripped of environmental interference, significantly improving the accuracy of emotion recognition. Furthermore, the system triggers corresponding vehicle control operations based on this purified emotional state, ensuring the accuracy of interactive feedback and improving the reliability of vehicle human-machine interaction and user experience. Specifically, detailed descriptions are provided below in conjunction with the accompanying drawings and specific embodiments.
[0039] See Figure 2 This is a flowchart illustrating a vehicle control method based on emotion recognition provided in an embodiment of this application. This method can be applied to the application scenarios shown above, such as... Figure 2 As shown, it mainly includes the following steps.
[0040] Step S201: Determine the environment vector based on the acquired in-vehicle environment information.
[0041] In this embodiment, an environmental monitoring module deployed within the vehicle cabin, such as a temperature sensor, humidity sensor, light sensor, air quality sensor, and odor sensor, continuously collects in-vehicle environmental information reflecting the in-vehicle environment. After acquiring the in-vehicle environmental information, the system performs a series of standardized data processing steps to construct an environmental vector that can be used for model calculations.
[0042] The environment vector is a vector representing the overall state of the in-vehicle environment. As a holistic indicator of environmental state, the core function of the environment vector is to provide precise and quantifiable input conditions for the subsequent adaptive adjustment of parameters in the emotion recognition model.
[0043] For example, when the air composition inside the car changes due to external exhaust fumes, this change is captured and reflected in the corresponding dimension of the environmental vector, thus providing key evidence for subsequent steps to identify and distinguish between the resulting frown (environmental cause) and the frown caused by road rage (emotional cause).
[0044] In one possible implementation, the in-vehicle environment information includes: in-vehicle temperature information, air quality information, and in-vehicle odor information. It is understood that in-vehicle temperature information can characterize the thermal comfort level inside the vehicle, thus affecting the user's facial movements; air quality information can characterize the cleanliness of the air inside the vehicle, thus affecting the user's facial movements; and in-vehicle odor information can characterize abnormal gases inside the vehicle, thus affecting the user's facial movements.
[0045] In practical applications, temperature sensor arrays, air quality sensor arrays, and odor sensor arrays are installed inside the vehicle to simultaneously collect the above three types of physical parameters, ensuring a comprehensive perception of the environmental conditions.
[0046] It should be noted that this application does not impose specific restrictions on the in-vehicle environment information. In some other implementations, the in-vehicle environment information may also include other factors that affect the comfort of drivers and passengers, such as ambient humidity and ambient light intensity. Those skilled in the art can make corresponding adjustments according to actual needs.
[0047] After acquiring in-vehicle temperature, air quality, and odor information, data preprocessing is required to generate corresponding standardized values for in-vehicle temperature, air quality, and odor. This preprocessing aims to eliminate differences caused by varying physical dimensions and sensor ranges, mapping the raw readings to a unified numerical range.
[0048] Finally, the system combines these standardized values in a predetermined order to form the environmental vector defined in this scheme. Specifically, the standardized values of in-vehicle temperature, air quality, and in-vehicle odor are combined to form the environmental vector. This environmental vector, as a whole, can comprehensively reflect the instantaneous state of the vehicle cabin environment.
[0049] In this embodiment, by uniformly preprocessing and standardizing the three key environmental information of in-vehicle temperature, air quality, and odor, the differences in magnitude between different sensor data are effectively eliminated. At the same time, the standardized values of in-vehicle temperature, air quality, and odor are combined to form an environmental vector, providing clear and comprehensive environmental information for the emotion recognition model. This lays a reliable data foundation for the subsequent accurate adjustment of the weights of facial action units, and ultimately significantly improves the accuracy of emotion state recognition in complex environments.
[0050] In one possible implementation, the aforementioned data preprocessing can be further decomposed into three key technical steps: data filtering, data fusion, and data standardization. These three steps work sequentially to ensure that the final generated environment vector has high reliability and usability.
[0051] Specifically, data filtering is the first step for processing raw sensor signals. Because of the complex operating environment of sensors, their readings are susceptible to transient interference, resulting in random fluctuations or spike noise. The purpose of filtering is to smooth out these abnormal data points and preserve the true trend of environmental parameters.
[0052] For example, when a passenger briefly opens a car window, causing a sudden influx of outside air, the odor sensor reading may show a brief, sharp peak. Through filtering algorithms, such non-persistent interference can be effectively suppressed, preventing it from having an excessive impact on the construction of the environmental vector.
[0053] After initial filtering, the system performs data fusion processing on multiple readings from similar sensors. Multiple sensors of the same type are typically deployed in different locations within a vehicle to obtain comprehensive environmental information. The fusion processing aims to integrate these distributed readings into a more representative, comprehensive indicator.
[0054] For example, the system can fuse and calculate multiple temperature sensor readings from the driver's seat, front passenger seat, and rear seats to generate a single temperature index that accurately reflects the thermal comfort of the entire cabin, avoiding the one-sidedness that may result from relying on readings from only a single location.
[0055] Finally, data standardization maps the filtered and fused environmental parameters to a unified numerical range. Because parameters such as temperature, air quality, and odor have different physical meanings and dimensions, their original numerical ranges vary greatly. Standardization eliminates these dimensional differences, converting all environmental parameters into dimensionless standardized values, thus enabling subsequent vector operations and model processing within the same mathematical space. This step makes weighted comparisons and comprehensive calculations between different environmental parameters possible.
[0056] In this embodiment, the quality and reliability of the environmental vector are systematically improved by collaboratively processing the in-vehicle temperature information, air quality information, and in-vehicle odor information through filtering, fusion, and standardization.
[0057] It should also be noted that, according to actual needs, those skilled in the art can set the above filtering process to different algorithms such as moving average filtering and Kalman filtering; the fusion process to strategies such as weighted average and confidence-based fusion; and the standardization process to various methods such as min-max standardization and Z-score standardization. This application does not impose specific restrictions on these methods.
[0058] Step S202: Dynamically adjust the weight parameters of each facial action unit in the emotion recognition model based on the environment vector.
[0059] In this embodiment of the application, the system takes the environment vector as input and calculates the adjustment amount of the weight parameters of each facial action unit in the model based on the preset environment-weight mapping relationship.
[0060] In the initial training phase, emotion recognition models learn the statistical association between facial movements and basic emotions under standard, ideal conditions. In one specific implementation, this emotion recognition model can be expressed by a formula. For example: .
[0061] in, Score based on emotion. This indicates the intensity of each facial movement unit. The weight parameters represent the weights of each facial motion unit. This is a bias term.
[0062] However, during actual driving, the in-car environment can directly trigger specific, non-emotional facial movements. In this application, the weight parameters of each facial action unit in the emotion recognition model are dynamically adjusted by analyzing the state of the current environment vector. This consequently reduces the importance of the in-vehicle environment in emotional decision-making.
[0063] For example, when the environment vector indicates that there is strong light shining inside the vehicle, the system will recognize the prevalence of "squinting" as a visual physiological response in this environment, and thus automatically reduce the weight coefficient of "squinting" when recognizing emotions such as "disgust" or "anger".
[0064] In one possible implementation, according to the formula: The weight parameters of each facial action unit in the emotion recognition model are determined and dynamically adjusted. To adjust the weight parameters of the i-th facial motion unit, To adjust the weight parameters of the i-th facial motion unit, , , These are the environment vectors of the i-th facial action unit. , , Standardized values of interior temperature in CRRC Standardized air quality values Standardized values of in-vehicle odor The sensitivity coefficient. It can be understood that the adjusted weight is equal to the original weight minus an adjustment amount, which is obtained by multiplying each component of the environment vector by its corresponding sensitivity coefficient and then summing the results.
[0065] The aforementioned sensitivity coefficients are the core of this model's ability to achieve precise adjustments. Each facial movement unit corresponds to a unique set of sensitivity coefficients, which define the degree to which the movement unit is affected by different environmental factors. For example, some facial movements that mainly involve the nose area may have a high sensitivity coefficient to changes in odor, while movements that mainly involve the forehead area may be more sensitive to changes in temperature.
[0066] In this embodiment, a mathematical formula based on a sensitivity coefficient is introduced to accurately map the quantitative information of the environmental vector to the parameter adjustment process of the emotion recognition model. This formula establishes a quantitative influence relationship between each facial action unit and different environmental factors, enabling the system to specifically adjust the weights of each facial action unit according to the specific environmental state, significantly improving the accuracy of the emotion recognition model.
[0067] Step S203: Input the facial image features of the driver and passengers at the current moment into the adjusted emotion recognition model to obtain the emotional state of the driver and passengers after removing interference from the in-vehicle environment.
[0068] In practical applications, the system first captures real-time facial images of drivers and passengers using in-vehicle vision sensors and extracts facial image features that reflect muscle movement patterns. Then, the current facial image features of the drivers and passengers are input into the adjusted emotion recognition model. When processing these facial image features, the emotion recognition model uses newly updated internal parameters for evaluation, and its decision-making logic incorporates the recognition and compensation for environmental interference factors. Finally, through the internal calculations and inferences of the emotion recognition model, the system outputs a judgment result on the emotional state of the drivers and passengers.
[0069] Understandably, this judgment has largely eliminated facial movement interference caused by in-vehicle environmental factors.
[0070] In one possible implementation, intensity values corresponding to multiple facial action units are obtained based on the facial image features of the driver and passengers at the current moment; the intensity values corresponding to the facial action units are weighted and calculated based on the adjusted weight parameters to obtain an emotion score; and based on the emotion score, the emotional state of the driver and passengers after removing interference from the in-vehicle environment is obtained.
[0071] Specifically, the system first needs to quantify and extract the activation intensity of facial action units from the acquired facial images. By analyzing facial key points and muscle movement areas, the system calculates the specific intensity value of each facial action unit at the current moment. These intensity values objectively reflect the significance of micro-facial movements such as drooping eyebrows, upturned corners of the mouth, and nasal flaring, constituting the original feature set for emotion computing.
[0072] The system then performs a weighted summation calculation based on the facial motion unit intensity values obtained above and the weight parameters adjusted by the environment vector, thereby obtaining the emotion score. Specifically, the emotion score... .in, This represents the emotional score after optimization based on the current environmental state. This represents the weight parameters of the facial motion unit after optimization for the current environmental state. This indicates the intensity of each facial movement unit. This is a bias term.
[0073] Finally, based on the weighted calculated emotion score, the system performs the final emotion state determination. Specifically, the system compares the emotion score with a preset determination threshold, or uses a classifier for pattern recognition, and finally outputs an emotion state classification result that removes interference from the in-vehicle environment.
[0074] It should be noted that the above implementation demonstrates a scheme based on weighted summation. In other implementations, the emotion score can be calculated using other mathematical models, such as probabilistic statistical models or deep learning neural network models; and the determination of emotion state can also employ multi-class decision functions or other pattern recognition methods. This application does not impose specific limitations in this regard.
[0075] In this embodiment, by weighting the collected facial motion unit intensity value with the weight parameters adjusted by the environment vector, the emotion score can accurately reflect the true emotional tendency of the driver and passengers after removing environmental interference, thereby ensuring that the final obtained emotional state is a reliable result based on the purified emotional signal, and effectively improving the reliability of the emotional state output.
[0076] In one possible implementation, after obtaining the emotional state of the driver and passengers after removing interference from the in-vehicle environment, the results of the aforementioned emotional state determination after removing interference from the environment are verified by integrating other independent information sources, thereby constructing a more robust emotion perception system.
[0077] Specifically, the system can simultaneously collect the voice signals and / or physiological signs of drivers and passengers; then, based on the voice signals and / or physiological signs, it can assist in verifying the emotional state after removing interference from the in-vehicle environment.
[0078] Understandably, voice signals are acquired through an in-vehicle microphone array, and their acoustic characteristics, such as tone, speed, and rhythm, carry rich emotional information. Physiological signals are monitored through contact or non-contact sensors (such as a hand heart rate detection module on the steering wheel or a body motion sensor built into the seat), reflecting physiological responses triggered by emotional changes, such as subtle changes in body state like heart rate variability and skin conductance. These signals are independent of facial expressions, providing a completely new dimension of observation for emotion recognition.
[0079] For example, when the system judges that the driver is in a "calm" state based on facial expressions, but at the same time detects that the driver's voice signal has a rapid and high-pitched tone, or that the physiological signal shows a significant increase in heart rate, the system will recognize this inconsistency, thereby lowering the confidence level of the final "calm" state, and tending to output more detailed judgments such as "potential anxiety" or "emotional conflict", and may even initiate further interaction to find out the user's true state.
[0080] In this embodiment, by introducing voice signals and physiological signs as independent information sources, the emotional state derived from facial expression analysis is reconfirmed and calibrated, thereby effectively correcting the recognition bias caused by individual expression habits and ultimately significantly improving the reliability of the emotional state judgment results.
[0081] Step S204: Based on the emotional state of the driver and passengers after removing interference from the in-vehicle environment, trigger the vehicle control operation corresponding to the emotional state.
[0082] In this embodiment, the system has a pre-stored database of correspondences between emotional states and vehicle control strategies. This database defines the types of interactive feedback to be triggered under different emotional tags. For example, when the system determines that the driver or passenger is in a state of "fatigue" or "distraction" due to long-term driving, the triggered control strategy will revolve around "refreshing" and "safety warning"; while when the system recognizes that the driver or passenger is in a state of "relaxation" or "pleasure", the strategy may focus on "maintaining comfort" and "enhancing the experience".
[0083] In practical applications, a more precise and user-friendly control strategy can be generated by comprehensively considering the core emotional needs of drivers and passengers and the actual conditions of the in-vehicle environment. One possible implementation involves performing a fusion decision based on the emotional state of the drivers and passengers after removing in-vehicle environmental interference, combined with in-vehicle environmental information; based on the result of the fusion decision, corresponding control commands are generated to trigger the collaborative operation of one or more vehicle interaction devices.
[0084] Specifically, the system does not make decisions solely based on the emotional state after removing environmental interference. Instead, it collaboratively analyzes this emotional state with real-time in-vehicle environmental information to execute a fusion decision. This decision-making process treats the emotional state as a signal of the user's intrinsic needs and environmental information as an external constraint. Through predefined decision rules or models, it outputs a comprehensive interaction solution. For example, when the system identifies that the occupants are in a "frustrated" state and simultaneously detects that the in-vehicle temperature is too high, the fusion decision logic will determine that "high temperature" is one of the possible triggers for "frustration," and thus formulate a combined instruction to "lower the air conditioning temperature" and "play soothing music," rather than simply executing the latter.
[0085] It should be noted that the specific implementation methods of integrated decision-making are diverse. Its decision-making logic can be based on conditional judgments according to preset rules, optimization calculations based on cost functions or utility models, or classification or regression models trained through machine learning. Furthermore, the combination of collaborative devices is not limited to air conditioning and audio systems, but can be extended to ambient lighting, seat ventilation and massage, fragrance generators, and other in-vehicle devices. This application does not impose specific limitations in this regard.
[0086] Understandably, by integrating emotional states free from environmental interference with real-time in-vehicle environmental information for decision-making, vehicle control operations can simultaneously respond to the real emotional needs of drivers and passengers as well as specific environmental conditions. This avoids the one-sidedness that may result from relying on a single information source, thereby generating more comprehensive and accurate control commands and providing users with a better interactive experience.
[0087] In this embodiment, an environmental vector is formed by real-time monitoring of the in-vehicle environment, thereby quantifying the overall state of the in-vehicle environment. Based on this environmental vector, the weight parameters of the facial motion unit are dynamically adjusted, effectively reducing the weight of physiological reactions caused by environmental discomfort in emotion recognition. This targeted parameter correction results in a more authentic emotional state obtained after removing environmental interference, significantly improving the accuracy of emotion recognition. On this basis, the system triggers corresponding vehicle control operations based on this purified emotional state, ensuring the accuracy of interactive feedback and improving the reliability of vehicle human-machine interaction and user experience.
[0088] Corresponding to the above-described method embodiments, this application also provides a vehicle control device based on emotion recognition. Specifically, see [link to relevant documentation]. Figure 3 This is a schematic diagram of a vehicle control device based on emotion recognition, provided in an embodiment of this application. Specifically, the figure shows a vehicle control device 300 based on emotion recognition. The vehicle control device 300 includes: an environment vector determination module 301, a weight parameter adjustment module 302, an emotion state acquisition module 303, and a vehicle control operation triggering module 304. Specifically, the environment vector determination module 301 determines an environment vector based on acquired in-vehicle environment information; the weight parameter adjustment module 302 dynamically adjusts the weight parameters of each facial action unit in the emotion recognition model based on the environment vector; the emotion state acquisition module 303 inputs the facial image features of the driver / passenger at the current moment into the adjusted emotion recognition model to obtain the driver / passenger's emotion state after removing in-vehicle environment interference; and the vehicle control operation triggering module 304 triggers a vehicle control operation corresponding to the emotion state based on the driver / passenger's emotion state after removing in-vehicle environment interference.
[0089] For details, please refer to the embodiments described above. For the sake of brevity, this application will not repeat them here.
[0090] Corresponding to the above method embodiments, this application also provides a schematic diagram of the structure of an electronic device. See also Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 400 may include a processor 401, a memory 402, and a communication unit 403. These components communicate through one or more buses. Those skilled in the art will understand that the structure of the electronic device shown in the figure does not constitute a limitation on the embodiments of the present invention. It may be a bus-shaped structure or a star-shaped structure, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0091] The communication unit 403 is used to establish a communication channel, enabling the electronic device to communicate with other devices. It receives user data from other devices or sends user data to other devices.
[0092] The processor 401 serves as the control center of the electronic device, connecting various parts of the device via various interfaces and lines. It executes software programs, instructions, and / or modules stored in the memory 402, and calls data stored in the memory to perform various functions and / or process data. The processor may be composed of integrated circuits (ICs), such as a single packaged IC or multiple packaged ICs with the same or different functions connected together. For example, the processor 401 may consist only of a central processing unit (CPU). In this embodiment, the CPU may have a single processing core or include multiple processing cores.
[0093] The memory 402 is used to store the execution instructions of the processor 401. The memory 402 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0094] When the execution instructions in memory 402 are executed by processor 401, the electronic device 400 is able to perform operations. Figure 2 Some or all of the steps in the illustrated embodiments.
[0095] In a specific implementation, this application also provides a computer storage medium, wherein the computer storage medium may store a program, and when the program is executed, it may include some or all of the steps in the various embodiments of the simulation scene generation method provided by this invention. The storage medium may be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0096] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent the existence of A alone, the simultaneous existence of A and B, or the existence of B alone. A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects have an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0097] Those skilled in the art will recognize that the units and algorithm steps described in the embodiments disclosed herein can be implemented using electronic hardware, computer software, or a combination of electronic hardware and software. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0098] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0099] In the several embodiments provided in this application, any function, if implemented as a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0100] The same or similar parts between the various embodiments in this specification can be referred to mutually. In particular, the device embodiments and terminal embodiments are basically similar to the method embodiments, so the description is relatively simple, and the relevant parts can be referred to the description in the method embodiments.
Claims
1. A vehicle control method based on emotion recognition, characterized by, The method comprises the following steps: According to the obtained in-vehicle environment information, an environment vector is determined, wherein the environment vector is a vector representing the comprehensive state of the in-vehicle environment; According to the environment vector, the weight parameters of each facial action unit in the emotion recognition model are dynamically adjusted; The facial image features of the driver or passenger at the current time are input into the adjusted emotion recognition model to obtain the emotional state of the driver or passenger after removing the in-vehicle environment interference; According to the emotional state of the driver or passenger after removing the in-vehicle environment interference, a vehicle control operation corresponding to the emotional state is triggered.
2. The method of claim 1, wherein, The in-vehicle environment information includes in-vehicle temperature information, air quality information, and in-vehicle odor information; and the environment vector is determined according to the obtained in-vehicle environment information, which comprises the following steps: Obtain the in-vehicle temperature information, air quality information, and in-vehicle odor information; The in-vehicle temperature information, air quality information, and in-vehicle odor information are respectively pre-processed to generate corresponding in-vehicle temperature standardized values, air quality standardized values, and in-vehicle odor standardized values; The in-vehicle temperature standardized values, air quality standardized values, and in-vehicle odor standardized values are combined to form an environment vector.
3. The method of claim 2, wherein, The data preprocessing includes data filtering processing, data fusion processing, and data standardization processing.
4. The method according to claim 1 or 2, characterized in that, According to the environment vector, the weight parameters of each facial action unit in the emotion recognition model are dynamically adjusted, which comprises the following steps: According to the formula: , the weight parameters of each facial action unit in the dynamic adjustment emotion recognition model are determined; in, To adjust the weight parameters of the i-th facial motion unit, To adjust the weight parameters of the i-th facial motion unit, , , These are the environment vectors of the i-th facial action unit. , , Standardized values of interior temperature in CRRC Standardized air quality values Standardized values of in-vehicle odor The sensitivity coefficient.
5. The method of claim 1, wherein, The facial image features of the driver or passenger at the current time are input into the adjusted emotion recognition model to obtain the emotional state of the driver or passenger after removing the in-vehicle environment interference, which comprises the following steps: Based on the facial image features of the driver or passenger at the current time, intensity values corresponding to a plurality of facial action units are obtained; Based on the adjusted weight parameters, the intensity values corresponding to the facial action units are weighted and calculated to obtain an emotional score; Based on the emotional score, the emotional state of the driver or passenger after removing the in-vehicle environment interference is obtained.
6. The method of claim 1, wherein, According to the emotional state of the driver or passenger after removing the in-vehicle environment interference, a vehicle control operation corresponding to the emotional state is triggered, which comprises the following steps: According to the emotional state of the driver or passenger after removing the in-vehicle environment interference, a fusion decision is made in combination with the in-vehicle environment information; Based on the result of the fusion decision operation, a corresponding control instruction is generated to trigger one or more vehicle interaction devices to work cooperatively.
7. The method of claim 1, wherein, After obtaining the emotional state of the driver or passenger after removing the in-vehicle environment interference, the following steps are further included: Obtain the speech signal and / or physiological sign signal of the driver or passenger; Based on the speech signal and / or physiological sign signal, the emotional state after removing the in-vehicle environment interference is verified.
8. A vehicle control device based on emotion recognition, characterized by, The method comprises the following steps: An environment vector determination module is configured to determine an environment vector according to obtained in-vehicle environment information, wherein the environment vector is a vector representing the comprehensive state of the in-vehicle environment; A weight parameter adjustment module is configured to dynamically adjust the weight parameters of each facial action unit in the emotion recognition model according to the environment vector; The emotional state acquisition module is configured to input a facial image feature of the driver or passenger at a current time into the adjusted emotional recognition model to acquire an emotional state of the driver or passenger after removing the interference of the in-vehicle environment. The vehicle control operation triggering module is configured to trigger a vehicle control operation corresponding to the emotional state of the driver or passenger after removing the interference of the in-vehicle environment according to the emotional state of the driver or passenger.
9. An electronic device, comprising: The electronic device comprises: a processor; a memory; and a computer program, wherein the computer program is stored in the memory, and the computer program comprises instructions which, when executed by the processor, cause the electronic device to perform the method of any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium comprises a stored program, wherein the program controls a device where the computer-readable storage medium is located to perform the method of any one of claims 1 to 8 when the program is running. The computer-readable storage medium comprises a stored program, wherein the program controls a device where the computer-readable storage medium is located to perform the method of any one of claims 1 to 8 when the program is running.