Vehicle atmosphere adjusting method, device and equipment and storage medium
By collecting and analyzing multimodal data from drivers and passengers, the system generates emotional state recognition results and adjusts the in-vehicle environment, solving the problem that the emotions of drivers and passengers affect driving safety and comfort, and achieving intelligent environmental optimization.
Patent Information
- Application Number
- CN202511245141.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-11-07
Smart Images

Figure CN120902750A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of vehicles, and in particular to a vehicle atmosphere adjusting method, device, equipment and storage medium. BACKGROUND
[0002] In the current era of rapid development of the automotive industry, automobiles have become an indispensable means of transportation in people's daily life. With the continuous improvement of consumers' requirements for automobile quality, the optimization of driving experience has become the focus of major automobile manufacturers and related enterprises.
[0003] In the related art, the emotional state of the driver and passengers during driving has a key influence on driving safety, comfort and overall driving experience. Positive emotions help to improve the reaction speed and decision-making ability of the driver, while negative emotions can lead to safety hazards such as inattention, so how to effectively focus on and improve the emotional state of the driver and passengers has become a problem to be solved.
[0004] In summary, the problems in the related art need to be solved. SUMMARY
[0005] The present application aims to at least partially solve one of the technical problems in the related art.
[0006] To this end, an object of an embodiment of the present application is to provide a vehicle atmosphere adjusting method, device, equipment and storage medium.
[0007] In order to achieve the above technical purpose, the technical solution adopted by the embodiments of the present application comprises:
[0008] On the one hand, the embodiments of the present application provide a vehicle atmosphere adjusting method, which comprises:
[0009] Collecting facial expression image data, speech data and body movement video data of the driver and passengers of a target vehicle;
[0010] Determining a first emotional evaluation score corresponding to the driver and passengers and a first confidence corresponding to the first emotional evaluation score according to the facial expression image data, determining a second emotional evaluation score corresponding to the driver and passengers and a second confidence corresponding to the second emotional evaluation score according to the speech data, and determining a third emotional evaluation score corresponding to the driver and passengers and a third confidence corresponding to the third emotional evaluation score according to the body movement video data;
[0011] According to the first confidence, the second confidence and the third confidence, the first emotional evaluation score, the second emotional evaluation score and the third emotional evaluation score are analyzed by fusion, and a sentiment state recognition result corresponding to the driver and passengers is determined.
[0012] According to the emotional state recognition result, a vehicle atmosphere adjustment scheme of the target vehicle is determined, and an operation state of an atmosphere component in the target vehicle is controlled according to the vehicle atmosphere adjustment scheme; wherein the atmosphere component includes an atmosphere lamp, a sound and a fragrance.
[0013] In addition, the vehicle atmosphere adjustment method according to the above-mentioned embodiment of the present application can have the following additional technical features:
[0014] Further, in an embodiment of the present application, the collection of the facial expression image data, the voice data and the body movement video data of the driver and passenger of the target vehicle includes:
[0015] The physiological signal data of the driver and passenger is continuously detected;
[0016] When the change rate of the physiological signal data exceeds a preset threshold, the facial expression image data, the voice data and the body movement video data of the driver and passenger of the target vehicle are collected.
[0017] Further, in an embodiment of the present application, the collection of the facial expression image data, the voice data and the body movement video data of the driver and passenger of the target vehicle includes:
[0018] The driving parameters of the target vehicle are collected;
[0019] According to the driving parameters, the driving state of the driver of the target vehicle is determined;
[0020] When the driving state meets a preset target state, the facial expression image data, the voice data and the body movement video data of the driver and passenger of the target vehicle are collected.
[0021] Further, in an embodiment of the present application, the determination of the first emotional evaluation score corresponding to the driver and passenger and the first confidence corresponding to the first emotional evaluation score according to the facial expression image data includes:
[0022] An image emotion recognition model established in advance is obtained;
[0023] The facial expression image data is input into the image emotion recognition model, the emotional tendency of the driver and passenger is predicted through the image emotion recognition model, and the first emotional evaluation score corresponding to the driver and passenger and the first confidence corresponding to the first emotional evaluation score are obtained.
[0024] Further, in an embodiment of the present application, the image emotion recognition model established in advance is obtained, including:
[0025] Obtaining a batch of training data, wherein the training data comprises sample expression image data of a sample person and a sentiment label corresponding to the sample expression image data, the sentiment label being used to represent a real sentiment tendency of the sample person;
[0026] Inputting the sample expression image data into an initialized image sentiment recognition model, predicting the sentiment tendency of the sample person through the image sentiment recognition model, and obtaining a predicted sentiment evaluation score corresponding to the sample person;
[0027] Determining a training loss value according to the sentiment label and the predicted sentiment evaluation score;
[0028] Updating parameters of the image sentiment recognition model according to the loss value, and obtaining an established image sentiment recognition model.
[0029] Further, in an embodiment of the present application, the determining of the training loss value according to the sentiment label and the predicted sentiment evaluation score comprises:
[0030] Determining the training loss value according to the sentiment label and the predicted sentiment evaluation score through a cross-entropy loss function.
[0031] Further, in an embodiment of the present application, the fusion analysis of the first sentiment evaluation score, the second sentiment evaluation score and the third sentiment evaluation score according to the first confidence, the second confidence and the third confidence to determine the sentiment state recognition result corresponding to the driver or passenger comprises:
[0032] Determining a weighted weight corresponding to the first sentiment evaluation score, the second sentiment evaluation score and the third sentiment evaluation score according to the first confidence, the second confidence and the third confidence;
[0033] Weighted fusion of the first sentiment evaluation score, the second sentiment evaluation score and the third sentiment evaluation score according to the weighted weight to obtain a fusion sentiment evaluation score corresponding to the driver or passenger;
[0034] Determining the sentiment state recognition result corresponding to the driver or passenger according to the fusion sentiment evaluation score.
[0035] On the other hand, an embodiment of the present application provides a vehicle atmosphere adjusting device, the device comprising:
[0036] A collection unit configured to collect facial expression image data, voice data and body movement video data of a driver or passenger of a target vehicle;
[0037] an evaluation unit configured to determine a first emotional evaluation score corresponding to the driver or passenger and a first confidence corresponding to the first emotional evaluation score according to the facial expression image data, determine a second emotional evaluation score corresponding to the driver or passenger and a second confidence corresponding to the second emotional evaluation score according to the voice data, and determine a third emotional evaluation score corresponding to the driver or passenger and a third confidence corresponding to the third emotional evaluation score according to the body movement video data;
[0038] an analysis unit configured to perform fusion analysis on the first emotional evaluation score, the second emotional evaluation score and the third emotional evaluation score according to the first confidence, the second confidence and the third confidence, and determine an emotional state recognition result corresponding to the driver or passenger;
[0039] an execution unit configured to determine a vehicle atmosphere adjustment scheme of the target vehicle according to the emotional state recognition result, and control an operation state of an atmosphere component in the target vehicle according to the vehicle atmosphere adjustment scheme, wherein the atmosphere component includes an atmosphere lamp, a sound and a fragrance.
[0040] In another aspect, an electronic device is provided, including:
[0041] at least one processor;
[0042] at least one memory configured to store at least one program;
[0043] when the at least one program is executed by the at least one processor, the at least one processor is caused to implement the vehicle atmosphere adjustment method.
[0044] In another aspect, an electronic device is provided, including:
[0045] In another aspect, an electronic device is provided, including:
[0046] The advantages and beneficial effects of the present application will be partially given in the following description, partially will become obvious from the following description, or will be learned by the practice of the present application:
[0047] This application discloses a vehicle atmosphere adjustment method, apparatus, device, and storage medium. Addressing the problem in related technologies where the emotional state of drivers and passengers affects driving safety and comfort, but effective adjustment methods are lacking, this application simultaneously collects facial expression images, voice data, and body movement videos of in-vehicle occupants. Subsequently, it calculates corresponding emotion assessment scores and their corresponding confidence levels based on the aforementioned data. Next, it fuses the confidence levels and emotion scores of multimodal data to generate accurate emotion state recognition results. Based on the recognition results, it automatically generates and executes an in-vehicle atmosphere adjustment scheme, dynamically optimizing the in-vehicle environment by controlling the operating status of components such as ambient lighting, audio, and fragrance, thereby improving driving safety and passenger comfort. This application can accurately determine the emotional state of drivers and passengers and intelligently and automatically adjust the vehicle's environmental atmosphere. Attached Figure Description
[0048] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following description is provided with accompanying drawings of the relevant technical solutions in the embodiments of this application or the prior art. It should be understood that the accompanying drawings described below are only for the purpose of clearly illustrating some embodiments of the technical solutions in this application. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0049] Figure 1 This is a schematic diagram illustrating the implementation environment of a vehicle atmosphere adjustment method provided in this application embodiment;
[0050] Figure 2 This is a flowchart illustrating a vehicle atmosphere adjustment method provided in an embodiment of this application;
[0051] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0052] The present application will be further described below with reference to the accompanying drawings and specific embodiments. The described embodiments should not be considered as limitations on the present application, and all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of the present application.
[0053] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description herein is for describing particular embodiments only and is not intended to be limiting of the application.
[0055] In the current era of rapid development of the automotive industry, cars have become an indispensable means of transportation in people's daily life. With the continuous improvement of consumers' requirements for car quality, the optimization of driving experience has become the focus of major car manufacturers and related enterprises.
[0056] In the related art, the emotional state of the driver and passengers during driving has a key influence on driving safety, comfort and overall driving experience. Positive emotions help to improve the reaction speed and decision-making ability of the driver, while negative emotions can lead to safety hazards such as inattention, so how to effectively focus on and improve the emotional state of the driver and passengers has become a problem to be solved.
[0057] Therefore, in the embodiments of the present application, a vehicle atmosphere adjustment method, device, equipment and storage medium are provided. In view of the problem in the related art that the emotional state of the driver and passengers affects driving safety and comfort but lacks effective adjustment means, the present application synchronously collects facial expression images, voice data and body movement videos of the driver and passengers in the vehicle; then, the corresponding emotional evaluation scores and their corresponding confidence levels are calculated based on the above data respectively; next, the confidence levels and emotional scores of the multi-modal data are fused to generate an accurate emotional state recognition result; according to the recognition result, a vehicle atmosphere adjustment scheme is automatically generated and executed, the running state of components such as atmosphere lights, sound, fragrance, etc. is controlled, the in-vehicle environment is dynamically optimized, and thus the driving safety and riding comfort are improved. The present application can accurately determine the emotional state of the driver and passengers and intelligently and automatically adjust the environmental atmosphere of the vehicle.
[0058] Please refer to Figure 1 , Figure 1 An implementation environment schematic diagram of a vehicle atmosphere adjustment method provided in the embodiments of the present application is shown. In this implementation environment, the main hardware and software subjects involved include a terminal device 110 and a background server 120. The terminal device 110 and the background server 120 are in communication connection.
[0059] Specifically, the vehicle atmosphere adjusting method provided in the embodiments of the present application can be executed on the terminal device 110 side alone or based on data interaction between the terminal device 110 and the background server 120. The terminal device 110 can be a vehicle-mounted terminal, for example, a central control unit of a vehicle; the background server 120 can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, CDN (Content Delivery Network), and big data and artificial intelligence platform.
[0060] The terminal device 110 and the background server 120 can establish a communication connection through a wireless network or a wired network. The wireless network or the wired network uses standard communication technology and / or protocol, and the network can be set as the Internet or any other network, for example, any combination of LAN (Local Area Network), MAN (Metropolitan Area Network), WAN (Wide Area Network), mobile, wired or wireless network, private network or virtual private network.
[0061] Of course, it can be understood that the implementation environment in Figure 1 is only some optional application scenarios of the vehicle atmosphere adjusting method provided in the embodiments of the present application, and the actual application is not fixed to the software and hardware environment shown in Figure 1 .
[0062] Next, in combination with the foregoing introduction of the implementation environment, a vehicle atmosphere adjusting method provided in the embodiments of the present application is introduced and described.
[0063] Please refer to Figure 2 , Figure 2 is a schematic diagram of a vehicle atmosphere adjusting method provided in the embodiments of the present application, which includes but is not limited to:
[0064] Step 210, collecting facial expression image data, voice data and body movement video data of the driver and passenger of the target vehicle;
[0065] Step 220, determining a first emotional evaluation score corresponding to the driver or passenger and a first confidence corresponding to the first emotional evaluation score according to the facial expression image data, determining a second emotional evaluation score corresponding to the driver or passenger and a second confidence corresponding to the second emotional evaluation score according to the voice data, and determining a third emotional evaluation score corresponding to the driver or passenger and a third confidence corresponding to the third emotional evaluation score according to the body movement video data;
[0066] Step 230, performing fusion analysis on the first emotional evaluation score, the second emotional evaluation score and the third emotional evaluation score according to the first confidence, the second confidence and the third confidence, and determining a recognition result of the emotional state corresponding to the driver or passenger;
[0067] Step 240, determining a vehicle atmosphere adjustment scheme of the target vehicle according to the recognition result of the emotional state, and controlling the operation state of an atmosphere component in the target vehicle according to the vehicle atmosphere adjustment scheme; wherein the atmosphere component includes atmosphere lights, sound and fragrance.
[0068] In the embodiments of the present application, a vehicle atmosphere adjustment method is provided, which aims to improve driving safety and comfort through multi-modal emotion recognition and intelligent environment control. The implementation process of the method is described in detail below.
[0069] In the embodiments of the present application, the vehicle performing the vehicle atmosphere adjustment method is recorded as a target vehicle. The target vehicle can be an intelligent connected vehicle or an intelligent cockpit vehicle which integrates a perception system, an intelligent computing unit and a multi-modal atmosphere adjustment component. The specific type of the target vehicle is not limited in the present application.
[0070] For the target vehicle, first, step 210 is performed, that is, the facial expression image data, voice data and body movement video data of the driver and passengers of the target vehicle are collected. In the target vehicle, a variety of high-performance data collection devices can be integrated. Exemplarily, in terms of visual data collection, a number of high-definition cameras can be embedded in key positions such as the rearview mirror, A-pillar and roof console of the target vehicle. These cameras have wide-angle vision and infrared night vision functions, ensuring that the driver's and each passenger's facial expressions can be clearly captured without dead angles, generating high-quality facial expression image data. At the same time, to accurately capture body language (such as hand gestures, posture changes, shoulder shrugs, etc.) that can reflect emotional state, 3D cameras or millimeter wave radars that can perceive depth of field can be installed on both sides of the vehicle roof or B-pillar to record body movement video data containing spatial depth information, thereby effectively distinguishing different passengers and understanding the meaning of their actions. In addition, in terms of auditory data collection, the target vehicle can hide an array of multiple high signal-to-noise ratio microphones in the vehicle interior. This array can not only record all voice data in the vehicle with high fidelity, including conversation content, tone, sighs, etc., but also can locate and separate sound sources through beamforming technology, effectively filtering environmental interference such as air conditioner noise, engine sound, wind noise, etc., thereby accurately locking and extracting voice data.
[0071] It should be noted that in the embodiments of the present application, the collection of facial expression image data, voice data and body movement video data of the driver and passengers needs to be performed under user authorization, and these data are only used for emotional analysis, and are usually anonymized or discarded after processing, thereby fully protecting the data security and privacy rights and interests of the user.
[0072] In step 220, the driver and passengers can be emotionally recognized based on the facial expression image data, voice data and body movement video data, thereby obtaining a quantitative emotional evaluation score and its corresponding confidence.
[0073] Specifically, for facial expression image data, a pre-trained image emotion recognition model can be invoked for analysis. The model can be based on a deep convolutional neural network (CNN) architecture, such as VGG or ResNet, and trained on a large-scale facial expression database, such as FER-2013 or AffectNet, capable of recognizing basic emotions such as anger, disgust, fear, happiness, sadness, surprise, and neutral. Upon receiving facial expression image data, the image emotion recognition model can first perform face detection and alignment, then extract texture and geometric features of key facial regions (such as eyebrows, eyes, and mouth), and finally classify and output a first sentiment assessment score representing the primary emotion category (e.g., happiness can be quantified as a high score, and sadness as a low score). At the same time, the model calculates a first confidence of this classification, which is usually determined by the highest value of the classification probability vector or the calibrated probability of the model, reflecting the influence of the current image quality, lighting conditions, and expression clarity on the credibility of the judgment result. If severe occlusion, excessive backlight, or facial deviation from the lens is detected, the confidence will be significantly reduced.
[0074] For speech data, the analysis process can be completed by acoustic feature analysis and semantic understanding in coordination. The original audio stream of speech data is first preprocessed by noise reduction and voice activity detection (VAD) to segment the valid speech segment. Then, the acoustic analysis module extracts low-level features of the speech segment, such as pitch, energy, speech rate, and spectrogram, and uses a classifier such as support vector machine (SVM) or recurrent neural network (RNN) to calculate a second sentiment assessment score based on paralanguage. In parallel, the semantic analysis module converts speech to text through automatic speech recognition (ASR), and then uses sentiment analysis tools in natural language processing (NLP) (such as BERT-based models) to analyze the sentiment polarity of the text content, and the result is combined with the acoustic analysis result to form the final second sentiment assessment score. The second confidence of this modality is a comprehensive consideration of the signal-to-noise ratio of the acoustic signal, the clarity of the speech, and the explicitness of the sentiment keywords in the semantic analysis.
[0075] For the limb movement video data, computer vision based behavior analysis techniques can be adopted. For example, the coordinates of human body key points (e.g., the positions of shoulders, elbows, hands, and head) can be extracted from the video frames by pose estimation algorithms such as OpenPose or MediaPipe, and a skeletal model can be constructed. Subsequently, the motion trajectories, amplitudes, velocities, and specific gestures (e.g., raising the forehead, waving the arms, and clasping the chest) of these key points are analyzed, and a time series classification model (e.g., 3D CNN or LSTM) is used to determine the corresponding emotional state, outputting a third emotional assessment score. For example, rapid and large amplitude movements can be related to excitement or anger, while slow movements with lowered head and drooping shoulders can be related to depression. The third confidence of this modality is heavily dependent on the quality of the video shooting, whether the limbs are blocked (e.g., by a steering wheel), and the typicality of the movement pattern. Blurred images or severely cropped limbs will result in a decreased confidence.
[0076] Of course, it can be understood that the above is only used to exemplarily illustrate the determination process of the emotional assessment score and the confidence, and does not mean to limit the specific implementation process.
[0077] In step 230, after obtaining the emotional assessment scores and confidences in each dimension, the multi-modal information fusion stage can be entered. In the embodiments of the present application, the three confidences (the first confidence, the second confidence, and the third confidence) can be standardized and compared to evaluate the data quality and reliability of each modality in the current environment. If the confidence of a certain modality is significantly higher than the other two (e.g., exceeds a pre-set threshold range relative to the confidences of the other two), the system tends to adopt the emotional assessment score corresponding to the high confidence modality as the dominant basis. For example, in the case of ideal lighting and the face directly facing the camera, the confidence of facial expression analysis will be very high, and even if the voice has a low confidence due to a noisy environment, the system will mainly rely on facial expressions to judge the emotion.
[0078] When the confidences of all modalities are at a high and similar level, a weighted average method based on the confidences can be used, i.e., a weighted weight is assigned to each emotional assessment score, which can be based on the size of the confidence corresponding to the emotional assessment score (e.g., the weighted weight is equal to the square of the confidence and is normalized). Subsequently, the weighted emotional assessment scores are summed up, and the final fused emotional assessment score can be obtained, and the emotional state recognition result corresponding to the driver or passenger can be determined.
[0079] In the embodiments of the present application, the system can define a sentiment mapping interval in advance, which divides the continuous fusion score into several segments corresponding to specific sentiment states. For example, the overall range of the fusion sentiment evaluation score is [-1, 1], the score in [-1, -0.3] can be mapped to the sentiment state recognition result of "negative / depressed", [-0.3, 0.3] is mapped to the sentiment state recognition result of "neutral / calm", and [0.3, 1] is mapped to the sentiment state recognition result of "positive / excited". The calculated fusion sentiment evaluation score is compared with this predefined interval, and which interval it falls into determines the current sentiment state of the driver or passenger as the type corresponding to the interval.
[0080] In step 240, after successfully obtaining the sentiment state recognition result of the driver or passenger, the system enters the strategy execution phase, dynamically generates and executes a customized vehicle atmosphere adjustment scheme according to the sentiment state recognition result, and realizes intelligent remodeling of the vehicle environment by coordinating the control of multiple atmosphere components.
[0081] Specifically, in the embodiments of the present application, a "sentiment-atmosphere" mapping database can be pre-set, which defines the optimal environment adjustment parameters corresponding to each sentiment state. When the sentiment state recognition result (such as "excited", "irritable", "depressed" or "calm") is determined, the central processor will immediately query the database to generate a vehicle atmosphere adjustment scheme containing specific control instructions. The scheme is a multi-dimensional control set, aiming to intervene through multiple sensory channels such as vision, hearing, and smell. For example, taking the "depressed" emotion as an example, the adjustment scheme generated by the system will contain the following instructions at the same time: control the atmosphere light system to slowly adjust the global light to warm yellow tone (such as 2700K color temperature), and perform soft breathing effect to create a warm and wrapped visual environment; control the sound system to play low-frequency and soothing background music or natural sound effects (such as breeze and drizzle sound), and maintain the volume at a moderate and low level to avoid auditory stimulation; control the intelligent fragrance system to release a specific proportion of invigorating fragrance, such as a mixture of citrus and cedar, to positively intervene through the olfactory channel. For another example, when the passenger is detected to be anxious, the system will automatically adjust the atmosphere light to soft blue, play soothing music, and release a fragrance that helps to relax, helping the passenger to quickly calm down. In addition, the system also provides specific scene modes, such as "meditation mode", which simulates natural scenes (such as forest campfire and sea wave sound) by combining seat vibration, white noise and light and shadow effects, helping passengers to further reduce stress and relax. This system not only improves the driving experience, but also effectively reduces driving stress, enhances driving safety and comfort.
[0082] It can be understood that the vehicle atmosphere adjusting method in the embodiments of the present application aims at the problem that the emotional state of the driver and passenger affects the driving safety and comfort in the related art but lacks effective adjusting means, synchronously collects facial expression images, voice data and body action videos of the driver and passenger in the vehicle; then, respectively calculates corresponding emotional evaluation scores and corresponding confidence thereof based on the above data; then, fuses the confidence and emotional scores of the multi-modal data to generate an accurate emotional state recognition result; automatically generates and executes a vehicle-mounted atmosphere adjusting scheme according to the recognition result, dynamically optimizes the in-vehicle environment by controlling the operating state of components such as atmosphere lamps, sound, fragrance, etc., thereby improving the driving safety and riding comfort. The method can accurately determine the emotional state of the driver and passenger, and intelligently and automatically adjust the environmental atmosphere of the vehicle.
[0083] Specifically, in some embodiments, the collecting facial expression image data, voice data and body action video data of the driver and passenger of the target vehicle comprises:
[0084] continuously detecting physiological signal data of the driver and passenger;
[0085] When it is detected that the change rate of the physiological signal data exceeds a preset threshold, collecting facial expression image data, voice data and body action video data of the driver and passenger of the target vehicle.
[0086] In the embodiments of the present application, the data collection process of step 210 can use an intelligent energy-saving and privacy-enhancing mechanism based on physiological signal triggering. The core of the mechanism is not to continuously and indiscriminately record all data, but to monitor the changes of more bottom-layer and more direct physiological indicators to determine the key moment when the emotional state of the driver and passenger may fluctuate, so as to start the multi-modal data collection with high power consumption on demand, which can significantly improve the efficiency and user experience of the system.
[0087] In the embodiments of the present application, the data collection process of step 210 introduces an intelligent energy-saving and privacy-enhancing mechanism based on physiological signal triggering. The core of the scheme is not to continuously and indiscriminately record all data, but to monitor the changes of more bottom-layer and more direct physiological indicators to determine the key moment when the emotional state of the driver and passenger may fluctuate, so as to start the multi-modal data collection with high power consumption on demand, which significantly improves the efficiency and user experience of the system.
[0088] In particular, in the embodiments of the present application, the steering wheel, seat belt or intelligent seat of the target vehicle can be embedded with non-contact or micro-contact biosensors (such as photoplethysmography (PPG) sensors for monitoring heart rate, or galvanic skin response (GSR) sensors for monitoring skin conductivity). These sensors continuously and passively monitor the physiological signal data (such as heart rate, heart rate variability (HRV), and galvanic skin response) of the driver or passenger with extremely low power consumption. These physiological signals are sensitive indicators of emotional activation states. For example, a sudden increase in heart rate or a sharp fluctuation in galvanic skin response is often closely related to emotional excitement (such as anger, fear or excitement). The system can calculate the rate of change of these physiological signals in real time (for example, calculate the rising slope of heart rate per unit time). When the absolute value of the rate of change exceeds a pre-set threshold, it is determined that the driver or passenger may be experiencing a significant emotional change event. At this time, it will immediately wake up and trigger high-power devices such as camera and microphone array, and start targeted data collection to record facial expression image data, voice data and body movement video data at this moment.
[0089] It can be understood that the intelligent energy-saving and privacy-enhancing mechanism triggered based on physiological signals in the embodiments of the present application has at least the following advantages: first, it avoids unnecessary audio and video recording when the user's emotions are stable, greatly reduces the overall power consumption of the system, and meets the energy efficiency requirements of vehicle-mounted devices; second, it minimizes the potential invasion of user privacy by continuous recording, and only records when something happens, embodying the principle of privacy design; third, it triggers through physiological signals, which are difficult to disguise subjectively, ensuring that the timing of data collection accurately corresponds to emotional physiological events, and providing higher-quality and more valuable raw data samples for subsequent emotional analysis.
[0090] In particular, in some embodiments, the collection of facial expression image data, voice data and body movement video data of the driver or passenger of the target vehicle comprises:
[0091] Collecting driving parameters of the target vehicle;
[0092] According to the driving parameters, determining the driving state of the driver of the target vehicle;
[0093] When the driving state meets the pre-set target state, collecting facial expression image data, voice data and body movement video data of the driver or passenger of the target vehicle.
[0094] In the embodiments of the present application, the data collection process of step 210 can also be deeply coupled with the actual operating conditions of the vehicle. Specifically, through the CAN bus or various sensors of the target vehicle, driving parameters can be continuously collected, including but not limited to vehicle speed, longitudinal / lateral acceleration, steering angle, gear position, throttle / brake opening, and trigger state of advanced driving assistance system (ADAS) (such as lane departure warning, forward collision warning activation). These parameters reflect the motion posture of the vehicle and the complexity of the driving environment in real time.
[0095] In the embodiments of the present application, these driving parameters can be comprehensively analyzed to determine the driving state of the driver. This module can have a pre-defined state machine or classifier built-in. For example, when the vehicle speed is zero and the gear is in "P" gear, the state is determined to be "static parking"; when the vehicle speed is low (such as < 30 km / h) and the acceleration changes dramatically, combined with navigation information, it can be determined to be "congestion crawling"; when the vehicle speed is high and the acceleration is stable, the lane is kept stable, and it is determined to be "high-speed cruising"; and when emergency braking, large and rapid steering, or ADAS warning is detected, it is determined to be "emergency operation".
[0096] In the embodiments of the present application, the system is pre-set to start multi-modal data collection in target states. For example, "static parking" and "congestion crawling" are considered as ideal target states. Because in these two states, the driving task load is low, the camera and microphone can be safely enabled for comprehensive data collection for deep emotional analysis or atmosphere adjustment, without distracting the driver's attention or interfering with the running performance of the vehicle. On the contrary, when the system determines that the current state is "high-speed cruising" or "emergency operation", data collection will be suspended or significantly reduced (such as only retaining the lowest frequency of physiological signal monitoring), to prioritize driving safety, and complete collection will be restored when the vehicle state returns to the target state.
[0097] Specifically, in some embodiments, the determining, according to the facial expression image data, of a first emotional assessment score corresponding to the driver or passenger and a first confidence corresponding to the first emotional assessment score comprises:
[0098] obtaining a pre-established image emotion recognition model;
[0099] inputting the facial expression image data into the image emotion recognition model, predicting an emotional tendency of the driver or passenger through the image emotion recognition model, and obtaining a first emotional assessment score corresponding to the driver or passenger and a first confidence corresponding to the first emotional assessment score.
[0100] Specifically, in some embodiments, the obtaining a pre-established image emotion recognition model comprises:
[0101] obtain a batch of training data, wherein the training data comprises sample facial expression image data of sample persons and emotion labels corresponding to the sample facial expression image data, the emotion labels being used to represent real emotional tendencies of the sample persons;
[0102] input the sample facial expression image data into an initialized image emotion recognition model, and predict emotional tendencies of the sample persons by the image emotion recognition model to obtain predicted emotion evaluation scores corresponding to the sample persons;
[0103] determine a loss value of training according to the emotion labels and the predicted emotion evaluation scores;
[0104] update parameters of the image emotion recognition model according to the loss value to obtain an established image emotion recognition model.
[0105] As described previously, the prediction of the emotion evaluation score can be realized based on a related artificial intelligence model. Illustratively, taking facial expression image data as an example, the system can call and load a pre-established image emotion recognition model to realize the related prediction. The image emotion recognition model is obtained from an offline and large-scale training process. In the training process, a batch of training data is obtained, which is composed of a large amount of “sample facial expression image data” and corresponding “emotion labels”. The sample facial expression image data covers rich expression pictures under different races, ages, genders, light conditions and head poses, and the emotion labels are annotated by experts or recorded after experimental induction, accurately representing real emotional tendencies (such as anger, happiness, sadness, etc.) of the sample persons when shooting.
[0106] Subsequently, the training algorithm inputs the sample facial expression image data into an initialized image emotion recognition model (which can be a deep convolutional neural network such as ResNet or VGG, etc.). The model performs layer-by-layer feature extraction and abstraction on the input image, and outputs a predicted emotion evaluation score representing the emotional tendency. Subsequently, the predicted emotion evaluation score of the model is compared with the real emotion label, and a loss value of this prediction is calculated by a specific loss function (such as mean square error MSE or cross-entropy loss), which quantifies the gap between the model prediction and the real situation.
[0107] In the embodiments of the present application, a back propagation algorithm can be used to iteratively update all parameters (i.e., weights and biases) of the model according to the calculated loss value, aiming to continuously reduce the loss value and make the prediction ability of the model continuously approach the real emotional distribution. This "forward prediction-loss calculation-backward update" loop process is iterated multiple times until the performance of the model on the independent validation set tends to be stable and reaches the required accuracy. At this time, the established image emotion recognition model is obtained, which can be deployed in a vehicle-mounted system.
[0108] In actual vehicle-mounted application scenarios, only the real-time collected facial expression image data of the driver and passenger needs to be directly input into the trained and solidified model, and the corresponding first emotional evaluation score of the driver and passenger can be predicted. The first confidence corresponding to the first emotional evaluation score can be determined by the information entropy of the output probability distribution or the calibrated confidence score, and the present application does not limit this.
[0109] In the embodiments of the present application, a vehicle atmosphere adjusting device is also provided, which comprises:
[0110] The collection unit is configured to collect facial expression image data, voice data and body movement video data of a driver and passenger of a target vehicle.
[0111] The evaluation unit is configured to determine a first emotional evaluation score corresponding to the driver and passenger and a first confidence corresponding to the first emotional evaluation score according to the facial expression image data, determine a second emotional evaluation score corresponding to the driver and passenger and a second confidence corresponding to the second emotional evaluation score according to the voice data, and determine a third emotional evaluation score corresponding to the driver and passenger and a third confidence corresponding to the third emotional evaluation score according to the body movement video data.
[0112] The analysis unit is configured to perform fusion analysis on the first emotional evaluation score, the second emotional evaluation score and the third emotional evaluation score according to the first confidence, the second confidence and the third confidence, and determine an emotional state recognition result corresponding to the driver and passenger.
[0113] The execution unit is configured to determine a vehicle-mounted atmosphere adjusting scheme of the target vehicle according to the emotional state recognition result, and control the operation state of an atmosphere component in the target vehicle according to the vehicle-mounted atmosphere adjusting scheme, wherein the atmosphere component comprises an atmosphere lamp, a sound and a fragrance.
[0114] Reference Figure 3 The embodiments of the present application provide an electronic device, which comprises:
[0115] at least one processor 310;
[0116] at least one memory 320, configured to store at least one program;
[0117] The at least one program, when executed by the at least one processor 310, enables the at least one processor 310 to implement the vehicle atmosphere adjustment method described above.
[0118] Similarly, the contents in the above method embodiments are all applicable to the present electronic device embodiments, the present electronic device embodiments specifically implement the same functions as the above method embodiments, and achieve the same beneficial effects as the above method embodiments.
[0119] The present application also provides a computer readable storage medium, which stores a program executable by the processor 310, and the program executable by the processor 310 is used to execute the vehicle atmosphere adjustment method described above when executed by the processor 310.
[0120] Similarly, the contents in the above method embodiments are all applicable to the present computer readable storage medium embodiments, the present computer readable storage medium embodiments specifically implement the same functions as the above method embodiments, and achieve the same beneficial effects as the above method embodiments.
[0121] The present application also provides a computer program product, which includes a computer program stored in a computer readable storage medium, and a processor of a computer device reads the computer program from the computer readable storage medium, and the processor executes the computer program, so that the computer device executes the vehicle atmosphere adjustment method described above.
[0122] In some alternative embodiments, the functions / operations mentioned in the block diagram can not occur in the order mentioned in the operation diagram. For example, depending on the functions / operations involved, two blocks shown in succession can actually be executed substantially simultaneously or the blocks can sometimes be executed in reverse order. In addition, the embodiments presented and described in the flowcharts of the present application are provided by way of example, with the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated in which the order of various operations is changed and in which sub-operations described as part of larger operations are independently executed.
[0123] Furthermore, although the present application is described in the context of functional modules, it is understood that one or more of the functions and / or features can be integrated in a single physical device and / or software module, or one or more functions and / or features can be implemented in separate physical devices or software modules. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary to an understanding of the present application. Rather, the actual implementation is within the routine skill of those in the art, given the nature of the property, function and internal relationships of the various functional modules disclosed herein. Therefore, the present application is not limited to the specific embodiments described herein, but only by the claims, with equivalents of the scope of these claims being permitted. It is also understood that the specific
[0124] If the functions are implemented in software, the functions can be stored in or implemented as one or more software modules on a computer-readable storage medium. In terms of this understanding, the technical solutions of the present application, in essence, or the part of the technical solutions that make a contribution to the prior art, or the part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing an apparatus (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0125] The logic and / or steps represented in the flowcharts, or otherwise described herein, for example, can be considered as a list of executable instructions for implementing logic functions, which can be specifically embodied in any computer-readable medium for use by or in conjunction with an instruction execution system, device or apparatus, such as a computer-based system, a system including a processor, or other system that can fetch and execute instructions from the instruction execution system, device or apparatus. For the purpose of this specification, the "computer-readable medium" can be any device that can contain, store, communicate, propagate or transport a program for use by or in conjunction with an instruction execution system, device or apparatus, or in conjunction with these instruction execution systems, devices or apparatus.
[0126] More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection (electronic) having one or more wires, a portable computer diskette (magnetic), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can also be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, for example, via optical scanning of the paper or other medium, then compiled, interpreted, or otherwise processed in a suitable manner, if necessary, and then stored in a computer memory.
[0127] It should be understood that aspects of the application can be implemented in hardware, software, firmware or combinations thereof. In the above embodiments, various steps or methods can be implemented in software or firmware that is stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any of the following technologies, or combinations thereof, can be used: a discrete logic circuit having logic gates for implementing logic functions upon data signals, an application specific integrated circuit having appropriate combinational logic gates, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0128] In the above description of the present specification, the description referring to the terms "one embodiment", "another embodiment", or "certain embodiments" or the like means that a specific feature, structure, material or characteristic described in connection with the embodiment or example is included in at least one embodiment or example of the present specification. The illustrative expressions of the above terms do not necessarily refer to the same embodiment or example throughout the present specification. Also, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0129] Although the embodiments of the present application have been shown and described, it will be appreciated by those skilled in the art that changes, modifications, alternatives and variations to these embodiments can be made without departing from the principles and spirit of the application, the scope of which is defined by the claims and their equivalents.
[0130] The above is a specific description of the preferred embodiments of the present application, but the present application is not limited to the embodiments, and those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present application, and these equivalent modifications or substitutions are included in the scope defined by the claims of the present application.
Claims
1. A vehicle atmosphere adjusting method characterized by, The method comprises: Collect facial expression image data, voice data and body action video data of a driver and a passenger of a target vehicle; Determine a first emotional evaluation score corresponding to the driver and passenger and a first confidence degree corresponding to the first emotional evaluation score according to the facial expression image data, determine a second emotional evaluation score corresponding to the driver and passenger and a second confidence degree corresponding to the second emotional evaluation score according to the voice data, and determine a third emotional evaluation score corresponding to the driver and passenger and a third confidence degree corresponding to the third emotional evaluation score according to the body action video data; Perform fusion analysis on the first emotional evaluation score, the second emotional evaluation score and the third emotional evaluation score according to the first confidence degree, the second confidence degree and the third confidence degree, and determine an emotional state recognition result corresponding to the driver and passenger; Determine a vehicle atmosphere adjustment scheme of the target vehicle according to the emotional state recognition result, and control an operation state of an atmosphere component in the target vehicle according to the vehicle atmosphere adjustment scheme; wherein the atmosphere component comprises an atmosphere lamp, a sound and a fragrance.
2. The vehicle atmosphere adjusting method according to claim 1, characterized by, The collection of the facial expression image data, the voice data and the body action video data of the driver and the passenger of the target vehicle comprises: Continuously detect physiological signal data of the driver and the passenger; When a change rate of the physiological signal data exceeds a preset threshold, collect the facial expression image data, the voice data and the body action video data of the driver and the passenger of the target vehicle.
3. The vehicle atmosphere adjusting method according to claim 1, characterized by, The collection of the facial expression image data, the voice data and the body action video data of the driver and the passenger of the target vehicle comprises: Collect driving parameters of the target vehicle; Determine a driving state of a driver of the target vehicle according to the driving parameters; When the driving state meets a preset target state, collect the facial expression image data, the voice data and the body action video data of the driver and the passenger of the target vehicle.
4. The vehicle atmosphere adjusting method according to claim 1, characterized by, The determination of the first emotional evaluation score corresponding to the driver and passenger and the first confidence degree corresponding to the first emotional evaluation score according to the facial expression image data comprises: Obtain a pre-established image emotional recognition model; Input the facial expression image data into the image emotional recognition model, predict an emotional tendency of the driver and the passenger through the image emotional recognition model, and obtain the first emotional evaluation score corresponding to the driver and passenger and the first confidence degree corresponding to the first emotional evaluation score.
5. The vehicle atmosphere adjusting method according to claim 4, characterized by, The obtaining of the pre-established image emotional recognition model comprises: Obtain a batch of training data; wherein the training data comprises sample expression image data of a sample person and an emotional label corresponding to the sample expression image data, and the emotional label is used to represent a real emotional tendency of the sample person; Input the sample expression image data into an initialized image emotional recognition model, predict an emotional tendency of the sample person through the image emotional recognition model, and obtain a predicted emotional evaluation score corresponding to the sample person; Determine a training loss value according to the emotional label and the predicted emotional evaluation score; According to the loss value, the image emotion recognition model is updated in parameters to obtain an established image emotion recognition model.
6. The vehicle atmosphere adjusting method according to claim 5, characterized by, The loss value of the training is determined according to the emotion label and the predicted emotion evaluation score, including: The loss value of the training is determined according to the emotion label and the predicted emotion evaluation score by a cross-entropy loss function.
7. The vehicle atmosphere adjusting method according to any one of claims 1 to 6, characterized by, The first emotion evaluation score, the second emotion evaluation score and the third emotion evaluation score are fused and analyzed according to the first confidence, the second confidence and the third confidence to determine the emotion state recognition result corresponding to the driver and passenger. The first emotion evaluation score, the second emotion evaluation score and the third emotion evaluation score are determined according to the first confidence, the second confidence and the third confidence to determine the weighted weight corresponding to the first emotion evaluation score, the second emotion evaluation score and the third emotion evaluation score. The first emotion evaluation score, the second emotion evaluation score and the third emotion evaluation score are weighted and fused according to the weighted weight to obtain the fused emotion evaluation score corresponding to the driver and passenger. The emotion state recognition result corresponding to the driver and passenger is determined according to the fused emotion evaluation score.
8. A vehicle atmosphere adjusting device characterized by comprising: The device comprises: The acquisition unit is configured to acquire facial expression image data, voice data and body movement video data of a driver and passenger of a target vehicle. The evaluation unit is configured to determine a first emotion evaluation score corresponding to the driver and passenger according to the facial expression image data and a first confidence corresponding to the first emotion evaluation score, determine a second emotion evaluation score corresponding to the driver and passenger according to the voice data and a second confidence corresponding to the second emotion evaluation score, and determine a third emotion evaluation score corresponding to the driver and passenger according to the body movement video data and a third confidence corresponding to the third emotion evaluation score. The analysis unit is configured to fuse and analyze the first emotion evaluation score, the second emotion evaluation score and the third emotion evaluation score according to the first confidence, the second confidence and the third confidence to determine an emotion state recognition result corresponding to the driver and passenger. The execution unit is configured to determine a vehicle atmosphere adjustment scheme of the target vehicle according to the emotion state recognition result, and control an operation state of an atmosphere component in the target vehicle according to the vehicle atmosphere adjustment scheme, wherein the atmosphere component comprises an atmosphere lamp, a sound and a fragrance.
9. An electronic device, comprising: comprises: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the vehicle atmosphere adjustment method according to any one of claims 1-7.
10. A computer readable storage medium having stored therein a program that is executable by a processor, characterized in that, The program executable by the processor when executed by the processor is used to implement the vehicle atmosphere adjustment method according to any one of claims 1-7.