An interactive artificial intelligence driven facial and voice analysis method and system

By collecting and analyzing the driver's facial images and voice data in real time, and using deep learning models to identify emotional states and automatically adjust the in-vehicle environment, the shortcomings of vehicle interaction systems in recognizing and responding to driver emotions are solved, thereby improving the driving experience and safety.

CN119705326BActive Publication Date: 2025-12-30HEALTH HOPE (BEIJING) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411823953.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-12-30
Estimated Expiration
2044-12-12

AI Technical Summary

Technical Problem

Existing vehicle interaction systems cannot effectively recognize the driver's emotional state, especially in the face of challenges such as temperature and humidity changes in the automotive environment. This results in the inability to automatically adjust the in-vehicle environment to alleviate the driver's emotions, thereby increasing driving risks.

Method used

By collecting users' facial images and voice data in real time, and using deep learning models for emotion recognition, the system automatically adjusts in-vehicle environmental parameters, such as temperature and humidity, according to the driver's emotional state to adapt to changes in the driver's emotions.

Benefits of technology

It enables real-time adjustments based on the user's emotional state, enhancing the driving experience, reducing driving risks, and providing a more intelligent and personalized driving environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119705326B_ABST
    Figure CN119705326B_ABST
Patent Text Reader

Abstract

The application provides an interactive artificial intelligence driven face and voice analysis method and system. The interactive artificial intelligence driven face and voice analysis method comprises: collecting face image data of a user according to a current vehicle driving state to obtain the face image data of the user in the vehicle; collecting voice data of the user according to the current vehicle driving state to obtain the voice data in the vehicle; performing emotion recognition processing on the face image data and the voice data pair through a deep learning model to obtain current emotion state information of the user; feeding back the current emotion state information of the user to a vehicle interaction system; and the vehicle interaction system retrieves a vehicle environment parameter scheme matched with the current emotion state of the user from a database and adjusts the vehicle environment according to the vehicle environment parameter scheme. The system comprises modules corresponding to the method steps.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention proposes an interactive, AI-driven facial and voice analysis method and system, belonging to the field of emotion recognition technology. Background Technology

[0002] With the rapid development of automotive technology, modern cars have evolved from simple transportation tools into mobile living spaces integrating numerous intelligent functions. Among these, the Human-Machine Interface (HMI), as a crucial bridge connecting the driver and the vehicle, is becoming increasingly intelligent and personalized. Traditional vehicle interaction systems primarily focus on providing basic information services such as navigation, entertainment, and communication, but they are significantly lacking in understanding and responding to the driver's emotional state. In recent years, the rise of artificial intelligence (AI) technology has brought revolutionary changes to vehicle interaction systems, especially the application of deep learning technology, which enables systems to more accurately identify and analyze users' complex emotions and needs. Emotion recognition technology, as an important branch of artificial intelligence, can achieve real-time and accurate judgment of users' emotional states by analyzing users' facial expressions, voice tone, heart rate, and other physiological and behavioral characteristics. However, existing AI-based emotion recognition systems are mostly applied to smart homes and mobile devices, with relatively few applications specifically for the automotive environment. In the automotive environment, due to factors such as temperature and humidity changes during driving, the collection of users' facial images and voice data faces greater challenges. Furthermore, effectively translating these emotion recognition results into actual responses from vehicle interaction systems to create a more comfortable and safer driving environment is also a pressing issue. Specifically, existing vehicle interaction systems often lack dynamic response mechanisms to changes in the user's emotional state. For example, when a driver is fatigued or stressed, the system cannot automatically adjust the in-vehicle environment (such as temperature and humidity) to alleviate their emotions, potentially increasing driving risks. Therefore, developing an interactive AI-driven method that can collect user facial images and voice data in real time, accurately identify the user's emotional state, and automatically adjust the in-vehicle environment based on that emotional state is of great significance for improving the driving experience and ensuring driving safety.

[0003] In summary, the present invention aims to provide an interactive, AI-driven facial and voice analysis method that overcomes the shortcomings of existing technologies and achieves the goal of automatically adjusting the in-vehicle environment based on the user's emotional state, thereby providing the driver with a more intelligent and personalized driving experience. Summary of the Invention

[0004] This invention provides an interactive, AI-driven facial and voice analysis method and system to solve the technical problems in the prior art. The technical solution adopted is as follows:

[0005] An interactive AI-driven facial and speech analysis method, the interactive AI-driven facial and speech analysis method comprising:

[0006] Based on the current vehicle driving status, collect the user's facial image data to obtain the user's facial image data inside the car;

[0007] Collect user voice data based on the current vehicle driving status to obtain voice data inside the car;

[0008] The facial image data and voice data are processed by a deep learning model to obtain the user's current emotional state information.

[0009] The user's current emotional state information is fed back to the vehicle interaction system;

[0010] The vehicle interaction system retrieves in-vehicle environment parameter schemes from the database that match the user's current emotional state, and adjusts the in-vehicle environment according to the in-vehicle environment parameter schemes.

[0011] Furthermore, based on the current vehicle driving status, the user's facial image data is collected, including:

[0012] Real-time monitoring of the vehicle's current speed;

[0013] The current vehicle speed is compared with a preset first speed threshold and a second speed threshold;

[0014] When the current vehicle speed does not exceed the preset first speed threshold, the user's facial image data is collected according to the preset initial collection frequency.

[0015] When the current vehicle speed exceeds a preset first speed threshold but does not exceed a preset second speed threshold, the first acquisition frequency of the user's facial image data is set according to the current vehicle speed, and the user's facial image data is acquired according to the set first acquisition frequency.

[0016] When the current vehicle speed exceeds the preset second driving speed threshold, the second acquisition frequency of the user's facial image data is set according to the current vehicle speed, and the user's facial image data is acquired according to the set second acquisition frequency.

[0017] Furthermore, based on the current vehicle speed, a first collection frequency for the user's facial image data is set, including:

[0018] When the current vehicle's speed exceeds a preset first speed threshold but does not exceed a preset second speed threshold, the current vehicle's speed is collected in real time during a first preset monitoring period starting from the moment the current vehicle's speed exceeds the preset first speed threshold, and is used as the first speed data; wherein, the value range of the first preset monitoring period is 5min-12min.

[0019] A first driving speed coefficient is obtained using the first driving speed data; wherein, the first driving speed coefficient is obtained by the following formula:

[0020]

[0021] Among them, S 01 V represents the first driving speed coefficient; 01 V represents the preset first driving speed threshold; 02 This represents the preset second driving speed threshold; V represents the current vehicle speed; α1 represents the maximum acceleration occurring during the first preset monitoring time period; α max T represents the maximum acceleration that a vehicle can achieve. s1 T represents the cumulative duration of the system in accelerated mode within the first preset monitoring time period; 01 This indicates the duration of the first preset monitoring time period;

[0022] The first acquisition frequency of the user's facial image data is set using the first driving speed coefficient; wherein, the first acquisition frequency of the user's facial image data is obtained by the following formula:

[0023]

[0024] Among them, F 01 F0 represents the first acquisition frequency of the user's facial image data; S represents the preset initial acquisition frequency; 01 T represents the first driving speed coefficient; s1 V represents the cumulative duration of the accelerated state within the first preset monitoring time period; z1 V represents the vehicle speed at the end of the first preset monitoring period; q1 This indicates the vehicle speed corresponding to the start time of the first preset monitoring period.

[0025] Furthermore, a second collection frequency for the user's facial image data is set based on the current vehicle speed, including:

[0026] When the current vehicle's speed exceeds a preset second speed threshold, the current vehicle's speed is collected in real time during a second preset monitoring period starting from the moment the current vehicle's speed exceeds the preset second speed threshold, and is used as the second speed data; wherein, the value range of the second preset monitoring period is 3min-8min.

[0027] A second driving speed coefficient is obtained using the second driving speed data; wherein the second driving speed coefficient is obtained by the following formula:

[0028]

[0029] Among them, S 02 V represents the second driving speed coefficient; 02 This represents the preset second driving speed threshold; V represents the current vehicle speed; α2 represents the maximum acceleration occurring during the second preset monitoring period; α max T represents the maximum acceleration that a vehicle can achieve. s2 This indicates the cumulative duration of the device in accelerated mode within the second preset monitoring time period; T 02 This indicates the duration of the second preset monitoring time period;

[0030] The second acquisition frequency of the user's facial image data is set using the second driving speed coefficient; wherein, the second acquisition frequency of the user's facial image data is obtained by the following formula:

[0031]

[0032] Among them, F 02 The second acquisition frequency represents the user's facial image data; F0 represents the preset initial acquisition frequency; S 02 T represents the second driving speed coefficient; s2 This indicates the cumulative duration of the accelerated state within the second preset monitoring time period; V z2 This indicates the vehicle speed at the end of the second preset monitoring period; V q2 This indicates the vehicle speed corresponding to the start time of the second preset monitoring period.

[0033] Furthermore, based on the current vehicle driving status, the user's voice data is collected, including:

[0034] Real-time monitoring of the vehicle's current speed;

[0035] The current vehicle speed is compared with a preset third speed threshold; wherein the third speed threshold is greater than the first speed threshold, but less than the second speed threshold.

[0036] When the current vehicle speed does not exceed the preset third driving speed threshold, the user's voice data is collected according to the preset initial wake-up duration.

[0037] When the current vehicle's speed exceeds a preset third speed threshold, the frequency of collecting the user's facial image data is extracted.

[0038] The wake-up duration is set by combining the current vehicle speed with the frequency of facial image data collection from the user.

[0039] The user's voice data is collected according to the set wake-up duration.

[0040] Furthermore, the wake-up duration is set by combining the current vehicle speed with the frequency of user facial image data collection, including:

[0041] When the current vehicle's speed exceeds a preset third speed threshold, the current vehicle's speed is extracted;

[0042] Extract the frequency of facial image data collection for the user corresponding to the current vehicle speed;

[0043] Retrieve the driving speed coefficient corresponding to the acquisition frequency of the user's facial image data;

[0044] The wake-up duration is obtained by combining the current vehicle speed with the frequency of the user's facial image data collection and its corresponding speed coefficient.

[0045] The wake-up duration is obtained using the following formula:

[0046]

[0047] Among them, T c T0 represents the initial wake-up duration; F represents the current vehicle speed corresponding to the frequency of facial image data collection; S represents the speed coefficient corresponding to the current frequency of facial image data collection; F0 represents the initial collection frequency; V 03 This indicates the preset third driving speed threshold; V represents the current vehicle speed.

[0048] Furthermore, the facial image data and voice data pairs are processed using a deep learning model to perform emotion recognition processing, thereby obtaining the user's current emotional state information, including:

[0049] The facial image data is preprocessed to obtain preprocessed facial image data, wherein the image preprocessing includes image noise reduction, image enhancement and image normalization.

[0050] The speech data is subjected to audio preprocessing to obtain preprocessed speech data, wherein the audio preprocessing includes audio noise reduction, audio frame segmentation, and volume enhancement.

[0051] The preprocessed facial image data is input into a trained deep learning model to obtain facial feature vectors; wherein the deep learning model adopts a convolutional neural network surface model.

[0052] The preprocessed audio speech data is processed using the MFCC feature extraction method to obtain the speech feature matrix corresponding to the speech data; wherein, the dimension of the speech feature matrix is ​​N×S, where N represents the number of frames contained in the speech feature matrix; and S represents the speech feature dimension of each frame.

[0053] When each frame of the audio data contains image data with the same frame time after the audio data is segmented, the image data with the same frame time as each frame of the audio data is used as the target image data.

[0054] When there is no image data with the same frame time for each frame of audio after the audio data of the speech data is divided into frames, the image data corresponding to the image acquisition time with the closest time distance to the frame time of each frame of audio is used as the target image data.

[0055] The facial feature vector of the target image data is concatenated with its corresponding speech feature matrix to obtain a comprehensive vector matrix;

[0056] The comprehensive vector matrix is ​​input into the emotion state analysis model for emotion recognition processing to obtain the user's current emotion state information.

[0057] Furthermore, the structure of the sentiment state analysis model is as follows:

[0058] The input layer is used to receive the integrated vector matrix as input.

[0059] The first fully connected layer is used to perform preliminary processing and nonlinear transformation on the input composite vector matrix. The first fully connected layer has 128 units and the activation function is the ReLU function.

[0060] The second fully connected layer is used for further feature extraction of the vector matrix obtained after processing by the first fully connected layer. The second fully connected layer has 256 units.

[0061] The Dropout layer is used to perform regularization processing on the data output by the second fully connected layer, discarding a preset proportion of data information, wherein the preset proportion ranges from 0.43 to 0.52.

[0062] The third fully connected layer is used to summarize the feature data in the data obtained by the Dropout layer and obtain the summarized feature data.

[0063] The output layer is used to perform sentiment classification based on the summarized feature data, obtain sentiment classification results, and output the sentiment classification results; wherein, the sentiment classification results refer to the probabilities corresponding to seven preset emotions, and the seven emotions include happiness, sadness, anger, surprise, fear, disgust, and calmness.

[0064] Furthermore, the vehicle interaction system retrieves an in-vehicle environment parameter scheme from the database that matches the user's current emotional state, and adjusts the in-vehicle environment according to the in-vehicle environment parameter scheme, including:

[0065] After receiving emotional state information, the vehicle interaction system extracts the emotional type corresponding to the highest probability as the user's current emotional state.

[0066] Retrieve in-vehicle environment parameter schemes that match the user's current emotional state from the database;

[0067] The in-vehicle environment parameters are extracted from the in-vehicle environment parameter scheme, wherein the in-vehicle environment parameters include temperature parameters and humidity parameters;

[0068] Adjust the current in-vehicle environment according to the stated temperature and humidity parameters.

[0069] An interactive AI-driven facial and voice analysis system, the interactive AI-driven facial and voice analysis system comprising:

[0070] The facial data acquisition module is used to collect the user's facial image data based on the current vehicle driving status, and to obtain the user's facial image data inside the car.

[0071] The voice data acquisition module is used to collect user voice data based on the current vehicle driving status and obtain voice data inside the car.

[0072] The emotional state information acquisition module is used to perform emotion recognition processing on the facial image data and voice data pair through a deep learning model to obtain the user's current emotional state information.

[0073] The emotional state information feedback module is used to feed back the user's current emotional state information to the vehicle interaction system;

[0074] The environmental parameter adjustment module is used by the vehicle interaction system to retrieve an in-vehicle environmental parameter scheme that matches the user's current emotional state from the database, and to adjust the in-vehicle environment according to the in-vehicle environmental parameter scheme.

[0075] Beneficial effects of this invention:

[0076] This invention provides an interactive, AI-driven facial and voice analysis method and system. By collecting and analyzing users' facial images and voice data in real time, the system can accurately identify users' emotional states and automatically adjust the in-car environment accordingly. This personalized adjustment will make users feel more comfortable and pleasant while driving, thereby enhancing the driving experience. When users are fatigued or stressed, the system can automatically adjust the in-car environment to alleviate their emotions, thereby reducing driving risks. Furthermore, by monitoring users' emotional states in real time, the system can also promptly detect and remind drivers to pay attention to driving safety. The application of this technology will drive the development of automobiles towards greater intelligence and personalization. In the future, with continuous technological advancements and the deepening of applications, automobiles will become a more intelligent and human-centered mobile living space. Attached Figure Description

[0077] Figure 1 This is a flowchart of the method described in this invention;

[0078] Figure 2 This is a system block diagram of the system described in this invention. Detailed Implementation

[0079] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0080] This invention proposes an interactive, AI-driven facial and voice analysis method, such as... Figure 1 As shown, the interactive AI-driven facial and voice analysis method includes:

[0081] S1. Collect user facial image data based on the current vehicle driving status, and obtain user facial image data inside the car.

[0082] S2. Collect user voice data based on the current vehicle driving status to obtain voice data inside the car;

[0083] S3. Perform emotion recognition processing on the facial image data and voice data pair using a deep learning model to obtain the user's current emotional state information;

[0084] S4. Feed back the user's current emotional state information to the vehicle interaction system;

[0085] S5. The vehicle interaction system retrieves an in-vehicle environment parameter scheme that matches the user's current emotional state from the database, and adjusts the in-vehicle environment according to the in-vehicle environment parameter scheme.

[0086] The working principle of the above technical solution is as follows: The system first determines whether it is suitable to collect facial image data based on the current driving status of the vehicle (such as speed, whether it is in a congested area, etc.). When the conditions are met, the system uses the in-vehicle camera to capture the user's facial image data. This data will be used for subsequent emotion recognition processing. Similar to facial image data collection, the system also determines whether it is suitable to collect voice data based on the vehicle's driving status. When the conditions are met, the system captures the user's voice data through the in-vehicle microphone. This data will be used together with the facial image data for emotion recognition processing. The system inputs the collected facial image data and voice data into a deep learning model. This deep learning model has been trained on a large amount of data and can accurately identify the user's emotional state. By analyzing the user's facial expressions, voice tone, and other features, the model can output the user's current emotional state information. The system feeds back the identified user emotional state information to the vehicle interaction system. The vehicle interaction system uses this information to determine whether the in-vehicle environment needs to be adjusted to adapt to the user's emotional state. The vehicle interaction system retrieves an in-vehicle environment parameter scheme that matches the user's current emotional state from the database. The system adjusts the in-vehicle environment accordingly based on these parameters to create a more comfortable and safer driving environment.

[0087] The aforementioned technical solution achieves the following results: By collecting and analyzing users' facial images and voice data in real time, the system can accurately identify the user's emotional state and automatically adjust the in-car environment accordingly. This personalized adjustment will make users feel more comfortable and enjoyable while driving, thereby enhancing the driving experience. When users are fatigued or stressed, the system can automatically adjust the in-car environment to alleviate their emotions, thus reducing driving risks. Furthermore, by monitoring the user's emotional state in real time, the system can also promptly detect and remind the driver to pay attention to driving safety. The application of this technical solution will drive the development of automobiles towards greater intelligence and personalization. In the future, with continuous technological advancements and the deepening of applications, automobiles will become a more intelligent and human-centered mobile living space.

[0088] In conclusion, this interactive AI-driven facial and voice analysis method has significant technical effects and application prospects, which will provide drivers with a more intelligent and personalized driving experience and promote the intelligent development of the automotive industry.

[0089] One embodiment of the present invention involves collecting facial image data of a user based on the current vehicle driving status, including:

[0090] S101. Real-time monitoring of the current vehicle speed;

[0091] S102. Compare the current vehicle speed with a preset first speed threshold and a second speed threshold;

[0092] S103. When the current vehicle speed does not exceed the preset first driving speed threshold, the user's facial image data is collected according to the preset initial collection frequency.

[0093] S104. When the current vehicle speed exceeds a preset first driving speed threshold, but does not exceed a preset second driving speed threshold, the first acquisition frequency of the user's facial image data is set according to the current vehicle speed, and the user's facial image data is acquired according to the set first acquisition frequency of facial image data.

[0094] S105. When the current vehicle speed exceeds the preset second driving speed threshold, the second acquisition frequency of the user's facial image data is set according to the current vehicle speed, and the user's facial image data is acquired according to the set second acquisition frequency of facial image data.

[0095] The working principle of the above technical solution is as follows: The vehicle's speed is monitored in real time using sensors or a GPS system. The monitored speed is compared with two preset speed thresholds (a first speed threshold and a second speed threshold). These two thresholds can be set according to road regulations, safety standards, and user experience. When the speed does not exceed the first speed threshold, the vehicle is considered to be at low speed or stationary, and facial image data is collected at a preset initial collection frequency (which may be relatively high). When the speed exceeds the first speed threshold but does not exceed the second speed threshold, the vehicle is considered to be at medium speed, and the collection frequency is dynamically adjusted based on the current speed, set to the first collection frequency. This frequency is usually lower than the initial collection frequency to reduce data volume and processing burden. When the speed exceeds the second speed threshold, the vehicle is considered to be at high speed, and the collection frequency is further reduced, set to the second collection frequency. This is because frequent facial image collection at high speeds may interfere with the driver, and facial changes at high speeds may not be as noticeable as at low speeds.

[0096] The advantages of the above technical solution are as follows: By dynamically adjusting the acquisition frequency, the amount of facial image data collected can be flexibly adjusted according to the vehicle's driving status, ensuring data accuracy while avoiding unnecessary resource waste. At low speeds or when stationary, a higher acquisition frequency ensures the integrity and accuracy of facial image data, which is beneficial for subsequent emotion recognition or other applications. At high speeds, reducing the acquisition frequency minimizes driver interference, improving driving safety and comfort. This solution automatically adjusts the acquisition frequency according to different driving speeds, enhancing the system's adaptability and robustness. The system maintains stable performance regardless of the vehicle's driving status. By reducing unnecessary acquisition operations, the power consumption of the in-vehicle system can be reduced, extending the equipment's lifespan.

[0097] In summary, this technical solution achieves flexible data collection and processing under different driving conditions by dynamically adjusting the acquisition frequency of facial image data, which not only ensures the accuracy and integrity of the data, but also improves user experience and system performance.

[0098] One embodiment of the present invention includes setting a first acquisition frequency for a user's facial image data based on the current vehicle speed, comprising:

[0099] S1041. When the current vehicle's speed exceeds a preset first speed threshold but does not exceed a preset second speed threshold, the current vehicle's speed is collected in real time during a first preset monitoring period starting from the moment the current vehicle's speed exceeds the preset first speed threshold, and is used as the first speed data; wherein, the value range of the first preset monitoring period is 5min-12min.

[0100] S1042. Obtain a first driving speed coefficient using the first driving speed data; wherein, the first driving speed coefficient is obtained by the following formula:

[0101]

[0102] Among them, S 01 V represents the first driving speed coefficient; 01 V represents the preset first driving speed threshold; 02 This represents the preset second driving speed threshold; V represents the current vehicle speed; α1 represents the maximum acceleration occurring during the first preset monitoring time period; α max T represents the maximum acceleration that a vehicle can achieve. s1 T represents the cumulative duration of the system in accelerated mode within the first preset monitoring time period; 01 This indicates the duration of the first preset monitoring time period;

[0103] S1043. Set the first acquisition frequency of the user's facial image data using the first driving speed coefficient; wherein, the first acquisition frequency of the user's facial image data is obtained by the following formula:

[0104]

[0105] Among them, F 01 F0 represents the first acquisition frequency of the user's facial image data; S represents the preset initial acquisition frequency; 01 T represents the first driving speed coefficient; s1 V represents the cumulative duration of the accelerated state within the first preset monitoring time period; z1 V represents the vehicle speed at the end of the first preset monitoring period; q1 This indicates the vehicle speed corresponding to the start time of the first preset monitoring period.

[0106] The working principle of the above technical solution is as follows: When the vehicle speed exceeds a preset first speed threshold (but does not exceed a second speed threshold), the system begins to collect real-time vehicle speed data within a first preset monitoring period (5min-12min). This data is referred to as the first speed data. Using the collected first speed data, combined with the preset first speed threshold, second speed threshold, maximum vehicle acceleration, maximum acceleration within the first preset monitoring period, and the cumulative duration of acceleration, a first speed coefficient (S01) is calculated using a specific formula. This coefficient reflects the vehicle's speed characteristics within the first preset monitoring period. Based on the calculated first speed coefficient, the preset initial acquisition frequency, the cumulative duration of acceleration within the first preset monitoring period, and the vehicle speed corresponding to the start and end times of the first preset monitoring period, the first acquisition frequency (F) of the user's facial image data is calculated using another specific formula. 01 This frequency will be used for subsequent acquisition of facial image data.

[0107] The above technical solution achieves the following effects: By dynamically adjusting the facial image data acquisition frequency based on vehicle speed, the system can more flexibly adapt to different driving scenarios and speed changes, thereby ensuring data acquisition quality while reducing unnecessary resource waste. Dynamically adjusting the acquisition frequency helps the system better balance data processing capabilities and real-time requirements. At higher vehicle speeds, increasing the acquisition frequency ensures the system can promptly capture changes in the user's face; conversely, at lower speeds, decreasing the acquisition frequency reduces the data processing burden and improves overall system performance. By rationally setting the acquisition frequency, the system can effectively collect and analyze the user's facial image data without interfering with driving, providing a more personalized and intelligent service experience. To a certain extent, by monitoring and analyzing the user's facial image data, the system can promptly detect potential risks such as driver fatigue and distraction, thereby reminding the driver to pay attention to safety and enhancing driving safety.

[0108] On the other hand, this solution can dynamically adjust the frequency of facial image data collection based on the vehicle's speed. When the vehicle speed is between a preset first speed threshold and a preset second speed threshold, the current vehicle speed is collected in real time, and a first speed coefficient is calculated based on this speed data, thereby adjusting the collection frequency. This dynamic adjustment allows the collection frequency to better adapt to different driving conditions. By calculating the first speed coefficient and the first collection frequency of the user's facial image data using a formula, the collection frequency can be flexibly adjusted according to the actual situation, improving the system's flexibility and response speed. This solution collects the current vehicle speed in real time within a first preset monitoring period and calculates the first speed coefficient, thus more accurately reflecting the vehicle's driving status within that period. This helps improve the accuracy and reliability of data collection. By dynamically adjusting the collection frequency, unnecessary collection times can be reduced while ensuring data collection quality, thereby improving data collection efficiency. Especially at high speeds, reducing the collection frequency can reduce the system load and improve system stability. When the vehicle is traveling at high speeds, frequent collection of the user's facial image data may cause discomfort to the user. This solution, by dynamically adjusting the collection frequency, can reduce interference to the user while ensuring data collection needs are met, improving user comfort. Dynamically adjusting the data collection frequency can also reduce the infringement of user privacy to some extent. When frequent data collection is not required, the system can reduce the collection frequency, thereby reducing the exposure of user privacy. This solution can collect the current vehicle speed in real time and dynamically adjust the collection frequency based on the speed data. This helps the system to obtain the user's driving status in a timely manner, improving the system's real-time performance and response speed. By monitoring and analyzing vehicle speed, the system can detect potential driving risks in advance, such as speeding and fatigued driving, and take corresponding early warning measures to improve system safety. This technical solution demonstrates good adaptability, flexibility, accuracy, efficiency, comfort, privacy protection, and security in detailed performance indicators. These effects collectively improve the overall system performance and user experience.

[0109] In summary, this technical solution achieves multiple technical benefits, including improved data acquisition flexibility, optimized system performance, enhanced user experience, and improved driving safety, by dynamically adjusting the acquisition frequency of facial image data according to vehicle speed.

[0110] One embodiment of the present invention includes setting a second acquisition frequency for the user's facial image data based on the current vehicle speed, comprising:

[0111] S1051. When the current vehicle's driving speed exceeds a preset second driving speed threshold, the current vehicle's driving speed is collected in real time during a second preset monitoring period starting from the moment the current vehicle's driving speed exceeds the preset second driving speed threshold, and is used as the second driving speed data; wherein, the value range of the second preset monitoring period is 3min-8min.

[0112] S1052. Obtain a second driving speed coefficient using the second driving speed data; wherein, the second driving speed coefficient is obtained by the following formula:

[0113]

[0114] Among them, S 02 V represents the second driving speed coefficient; 02 This represents the preset second driving speed threshold; V represents the current vehicle speed; α2 represents the maximum acceleration occurring during the second preset monitoring period; α max T represents the maximum acceleration that a vehicle can achieve. s2 This indicates the cumulative duration of the device in accelerated mode within the second preset monitoring time period; T 02 This indicates the duration of the second preset monitoring time period;

[0115] S1053. Set the second acquisition frequency of the user's facial image data using the second driving speed coefficient; wherein, the second acquisition frequency of the user's facial image data is obtained by the following formula:

[0116]

[0117] Among them, F 02 The second acquisition frequency represents the user's facial image data; F0 represents the preset initial acquisition frequency; S 02 T represents the second driving speed coefficient; s2 This indicates the cumulative duration of the accelerated state within the second preset monitoring time period; V z2 This indicates the vehicle speed at the end of the second preset monitoring period; V q2 This indicates the vehicle speed corresponding to the start time of the second preset monitoring period.

[0118] The working principle of the above technical solution is as follows: When the vehicle's speed exceeds a preset second driving speed threshold, the system will collect the vehicle's speed in real time during the next second preset monitoring period (3-8 minutes). This data is referred to as second driving speed data. Using the collected second driving speed data, a second driving speed coefficient is calculated using a specific formula. This coefficient considers multiple factors, including the preset second driving speed threshold, the current vehicle speed, the maximum acceleration occurring during the first preset monitoring period, the maximum acceleration the vehicle can achieve, and the cumulative duration of acceleration during the second preset monitoring period. The various parameters in the formula work together to calculate the second driving speed coefficient, reflecting the vehicle's dynamic characteristics at the current speed. Based on the calculated second driving speed coefficient, a second acquisition frequency for the user's facial image data is set using another formula. This frequency is not only related to the second driving speed coefficient but is also affected by the preset initial acquisition frequency, the cumulative duration of acceleration during the second preset monitoring period, and the vehicle's speed at the start and end of the monitoring period. In this way, the system can flexibly adjust the acquisition frequency of facial image data according to the vehicle's dynamic driving state.

[0119] The effects of the above technical solution are as follows: By collecting and analyzing vehicle speed data in real time, the system can more accurately reflect the vehicle's dynamic state, thereby allowing for more precise setting of the facial image data collection frequency. The system can dynamically adjust the collection frequency based on parameters such as vehicle speed and acceleration, enhancing its dynamic adaptability and robustness. By reducing unnecessary data collection operations, system power consumption and data processing burden can be reduced, optimizing resource utilization. At high speeds, lowering the collection frequency reduces driver interference, improving driving safety and comfort. Through more precise control of the data collection frequency, the system can better monitor the driver's state, promptly detect potential safety hazards, and thus enhance system safety.

[0120] Simultaneously, by collecting real-time vehicle speed data and analyzing it within a second preset monitoring period, this solution can accurately capture the vehicle's high-speed driving state. This helps ensure the system can respond quickly when the vehicle speed exceeds a certain threshold. When the vehicle speed exceeds the preset second speed threshold, the system can immediately initiate real-time data collection within the second preset monitoring period, thus achieving continuous and dynamic monitoring of the speed. This dynamic response capability helps the system more accurately assess the vehicle's driving state. By using a second speed coefficient to set the second collection frequency of the user's facial image data, this solution can reduce unnecessary collection attempts while ensuring data collection quality. This helps improve the overall efficiency of the system and reduce resource waste. Frequent collection of facial image data at high speeds can consume significant resources. By dynamically adjusting the collection frequency, this solution can optimize resource usage while ensuring data collection needs are met, improving system stability and reliability. Timely and accurate acquisition of user facial image data at high speeds is crucial for ensuring driving safety. This solution, by dynamically adjusting the collection frequency, ensures that user facial information is obtained at critical moments, thereby contributing to improved driving safety. While maintaining efficient data collection, this solution also considers user experience. By reducing unnecessary data collection times, user interference can be minimized, improving user satisfaction and comfort. The formulas for the second driving speed coefficient and the second data collection frequency in this scheme can be adjusted and optimized based on actual conditions. This flexibility allows the scheme to adapt to different application scenarios and needs.

[0121] In summary, this technical solution demonstrates superior performance in terms of accuracy, dynamic response, efficiency, resource optimization, security, user experience, flexibility, and scalability across detailed performance metrics. These effects collectively enhance the overall system performance and user experience, providing strong protection for vehicle driving safety. By dynamically adjusting the acquisition frequency of facial image data, this solution enables flexible data acquisition and processing under different driving conditions, ensuring both data accuracy and integrity while improving user experience and system performance.

[0122] One embodiment of the present invention involves collecting user voice data based on the current vehicle driving status, including:

[0123] S201. Real-time monitoring of the current vehicle speed;

[0124] S202. Compare the current vehicle speed with a preset third driving speed threshold; wherein the third driving speed threshold is greater than the first driving speed threshold, but less than the second driving speed threshold.

[0125] S203. When the current vehicle speed does not exceed a preset third driving speed threshold, the user's voice data is collected according to a preset initial wake-up duration; wherein, the wake-up duration refers to the duration of continuous voice collection performed by the voice data collection module after the user's voice command stops; if no voice data is collected in the car by the end of the wake-up duration, the voice data collection module is controlled to go into sleep mode until the next voice command issued by the user wakes it up for voice collection.

[0126] S204. When the current vehicle speed exceeds the preset third driving speed threshold, the frequency of collecting the user's facial image data is extracted.

[0127] S205. Set the wake-up duration by combining the current vehicle speed with the frequency of collecting the user's facial image data;

[0128] S206. Collect the user's voice data according to the set wake-up duration.

[0129] The working principle of the above technical solution is as follows: The system monitors the vehicle's current speed in real time through the vehicle's sensors or GPS. The monitored speed is compared with a preset third speed threshold. When the speed does not exceed the preset third speed threshold, the system collects the user's voice data according to a preset initial wake-up duration. The wake-up duration refers to the length of time the voice data acquisition module continues to collect voice data after the user's voice command has stopped. If no new voice data is collected within this time, the voice data acquisition module enters a sleep state until the user issues a voice command to wake it up again. When the speed exceeds the preset third speed threshold, the system extracts the current user's facial image data collection frequency. This frequency may be dynamically adjusted based on the previous vehicle driving status, reflecting the current user's attention and driving state. Using the current vehicle speed and the user's facial image data collection frequency, the system dynamically sets a new wake-up duration. This new duration may take into account the impact of vehicle speed on the user's attention and the user's current driving state. The system collects the user's voice data according to the set wake-up duration, ensuring accurate reception and processing of the user's voice commands when necessary.

[0130] The effects of the above technical solution are as follows: By dynamically adjusting the wake-up duration, the system can flexibly adjust the voice data acquisition strategy according to the vehicle's driving status and the user's driving state, ensuring that the user's voice commands are acquired at the appropriate time. At low speeds or when stationary, a longer wake-up duration ensures the system can accurately receive the user's voice commands, improving the system's response speed and accuracy. At high speeds, a shorter wake-up duration reduces interference with the user, improving driving safety and comfort. This solution can automatically adjust the voice data acquisition strategy according to different driving speeds and the user's driving state, enhancing the system's adaptability and robustness. By reducing unnecessary voice data acquisition operations, the system's power consumption and data processing burden can be reduced, optimizing resource utilization.

[0131] In summary, this technical solution achieves flexible data collection and processing under different driving conditions by dynamically adjusting the wake-up duration of voice data acquisition. This ensures data accuracy and integrity while improving user experience and system performance. Furthermore, the solution considers the user's driving state and attention level, further enhancing the system's safety and adaptability.

[0132] One embodiment of the present invention sets the wake-up duration by combining the current vehicle speed with the frequency of user facial image data collection, including:

[0133] S2051. When the current vehicle's speed exceeds a preset third speed threshold, extract the current vehicle's speed.

[0134] S2052. Extract the acquisition frequency of the user's facial image data corresponding to the current vehicle's driving speed;

[0135] S2053. Retrieve the driving speed coefficient corresponding to the acquisition frequency of the user's facial image data;

[0136] S2054. The wake-up duration is obtained by combining the current vehicle speed with the frequency of the user's facial image data collection and its corresponding speed coefficient.

[0137] The wake-up duration is obtained using the following formula:

[0138]

[0139] Among them, T c T0 represents the initial wake-up duration; F represents the current vehicle speed corresponding to the frequency of facial image data collection; S represents the speed coefficient corresponding to the current frequency of facial image data collection; F0 represents the initial collection frequency; V03 This indicates the preset third driving speed threshold; V represents the current vehicle speed.

[0140] The working principle of the above technical solution is as follows: When the vehicle's speed exceeds a preset third speed threshold, the system first extracts the current vehicle speed. This speed is obtained in real time through the vehicle's sensors or GPS system. Next, the system extracts the acquisition frequency of the user's facial image data corresponding to the current vehicle speed. This frequency may be dynamically adjusted based on the previous vehicle driving status and user behavior, reflecting the current user's attention and driving state. Based on the acquisition frequency of the current user's facial image data, the system retrieves the corresponding driving speed coefficient. This coefficient may be an empirical value or obtained based on big data analysis, used to reflect the impact of vehicle speed on user attention at different acquisition frequencies. Finally, the system uses the current vehicle speed, the acquisition frequency of the user's facial image data, and the corresponding driving speed coefficient to calculate the wake-up duration using a specific formula. This formula comprehensively considers multiple factors, including the preset initial wake-up duration, the ratio of the current acquisition frequency to the initial acquisition frequency, the driving speed coefficient, and the ratio of the current driving speed to the third speed threshold.

[0141] The above technical solution achieves the following effects: By comprehensively considering vehicle speed, the frequency of facial image data acquisition from the user, and the speed coefficient, the system can more accurately calculate the wake-up duration, ensuring that the voice data acquisition module is woken up at the appropriate time, reducing unnecessary interference and power consumption. This solution can automatically adjust the wake-up duration according to different driving speeds and the user's driving state, enhancing the system's dynamic adaptability and robustness. Regardless of the vehicle's driving state, the system maintains stable performance. At high speeds, a shorter wake-up duration reduces interference to the user, improving driving safety and comfort. Simultaneously, by accurately calculating the wake-up duration, the system can ensure timely response to the user's voice commands when necessary, improving the user experience. By reducing unnecessary voice data acquisition operations, system power consumption can be reduced, extending the device's lifespan.

[0142] Meanwhile, this solution can dynamically adjust the wake-up duration based on the vehicle's current speed and the frequency of facial image data collection. This dynamic adjustment mechanism allows the system to maintain higher alertness at high speeds or in complex road conditions, thus responding promptly to potential driving risks. By setting the wake-up duration in relation to driving speed and data collection frequency, the system can quickly activate at critical moments, providing the driver with necessary assistance or warning information. This helps improve driving safety and reduce the risk of traffic accidents.

[0143] Under low-speed driving or safe road conditions, the system can reduce unnecessary system wake-ups and data processing by decreasing the wake-up duration, thereby reducing system energy consumption and extending service life. By accurately calculating the wake-up duration, the system can optimize resource utilization and improve overall system efficiency while ensuring driving safety. When frequent system wake-ups are unnecessary, this solution reduces system interference with the user, enhancing the driving experience. By dynamically adjusting the wake-up duration and providing timely and effective auxiliary information, the system can increase user trust and improve user acceptance of intelligent driving technology. This solution can flexibly adjust the wake-up duration according to different driving scenarios (such as urban roads, highways, etc.) and driving speeds to adapt to different driving environments and needs. The design of this technical solution has high flexibility and scalability, facilitating subsequent functional expansion and performance upgrades to meet the future development needs of intelligent driving technology.

[0144] In summary, this technical solution demonstrates significant improvements in driving safety, system resource utilization, user experience, and system flexibility. These improvements collectively enhance the overall performance and reliability of the intelligent driving system, providing drivers with a safer, more comfortable, and convenient driving experience. Furthermore, by dynamically adjusting the wake-up duration based on multiple factors, this solution enables flexible control of the voice data acquisition module under different driving conditions, ensuring data accuracy and integrity while simultaneously improving user experience and system performance.

[0145] In one embodiment of the present invention, an emotion recognition process is performed on the facial image data and voice data pair using a deep learning model to obtain the user's current emotional state information, including:

[0146] S301. Perform image preprocessing on the facial image data to obtain preprocessed facial image data, wherein the image preprocessing includes image noise reduction, image enhancement, and image normalization.

[0147] S302. Perform audio preprocessing on the speech data to obtain preprocessed speech data, wherein the audio preprocessing includes audio noise reduction processing, audio frame segmentation processing, and volume enhancement processing.

[0148] S303. Input the preprocessed facial image data into a trained deep learning model to obtain facial feature vectors; wherein the deep learning model adopts a convolutional neural network surface model.

[0149] S304. Process the pre-processed audio speech data using the MFCC feature extraction method to obtain the speech feature matrix corresponding to the speech data; wherein, the dimension of the speech feature matrix is ​​N×S, where N represents the number of frames contained in the speech feature matrix; and S represents the speech feature dimension of each frame.

[0150] S305. When each frame of audio data has image data with the same frame time after the audio data is divided into frames, the image data with the same frame time as each frame of audio data is used as the target image data.

[0151] S306. When there is no image data with the same frame time for each frame of audio after the audio data of the voice data is divided into frames, the image data corresponding to the image acquisition time with the closest time distance to the frame time of each frame of audio is taken as the target image data.

[0152] S307. Concatenate the facial feature vector of the target image data with its corresponding speech feature matrix to obtain a comprehensive vector matrix;

[0153] S308. Input the comprehensive vector matrix into the emotion state analysis model for emotion recognition processing to obtain the user's current emotion state information.

[0154] The structure of the emotion state analysis model is as follows:

[0155] The input layer is used to receive the integrated vector matrix as input.

[0156] The first fully connected layer is used to perform preliminary processing and nonlinear transformation on the input composite vector matrix. The first fully connected layer has 128 units and the activation function is the ReLU function.

[0157] The second fully connected layer is used for further feature extraction of the vector matrix obtained after processing by the first fully connected layer. The second fully connected layer has 256 units.

[0158] The Dropout layer is used to perform regularization processing on the data output by the second fully connected layer, discarding a preset proportion of data information, wherein the preset proportion ranges from 0.43 to 0.52.

[0159] The third fully connected layer is used to summarize the feature data in the data obtained by the Dropout layer and obtain the summarized feature data.

[0160] The output layer is used to perform sentiment classification based on the summarized feature data, obtain sentiment classification results, and output the sentiment classification results; wherein, the sentiment classification results refer to the probabilities corresponding to seven preset emotions, and the seven emotions include happiness, sadness, anger, surprise, fear, disgust, and calmness.

[0161] The working principle of the above technical solution is as follows: Image noise reduction, image enhancement, and image normalization are performed on the acquired facial image data to obtain preprocessed facial image data. These preprocessing steps aim to improve image quality, reduce noise interference, and standardize the image format for easier subsequent processing. Audio noise reduction, audio framing, and volume enhancement are performed on the acquired speech data to obtain preprocessed speech data. Audio preprocessing aims to improve the clarity of the speech signal, reduce background noise, and segment the speech signal into frames that are easier to process.

[0162] The preprocessed facial image data is input into a pre-trained deep learning model (using a convolutional neural network model) to obtain facial feature vectors. These feature vectors can represent key information in the facial image.

[0163] The preprocessed speech data is processed using MFCC (Mel-frequency cepstral coefficients) feature extraction to obtain the corresponding speech feature matrix. MFCC is a commonly used speech feature extraction method that reflects the spectral characteristics of the speech signal. Based on the audio framing results of the speech data, the target image data corresponding to each frame is determined. If an image with the same frame exists at that time, it is used directly; otherwise, the image data corresponding to the image acquisition time closest to the frame time is selected as the target image data. The facial feature vector of the target image data is concatenated with its corresponding speech feature matrix to obtain a comprehensive vector matrix. This matrix integrates key information from the facial image and speech signal, providing rich feature data for subsequent emotion recognition. The comprehensive vector matrix is ​​input into the emotion state analysis model for emotion recognition processing. This model includes an input layer, a first fully connected layer, a second fully connected layer, a Dropout layer, a third fully connected layer, and an output layer. Through layer-by-layer processing, the model performs preliminary processing, feature extraction, regularization, feature summarization, and emotion classification on the input feature data, ultimately outputting the user's current emotion state information. The emotion classification results include the probabilities corresponding to seven preset emotions: happiness, sadness, anger, surprise, fear, disgust, and calmness.

[0164] The aforementioned technical solution achieves the following results: by fusing key information from facial images and speech signals, it can more comprehensively capture the user's emotional state, thereby improving the accuracy of emotion recognition. This solution employs deep learning models and the MFCC feature extraction method, which have wide applications and validation in image processing and speech signal processing, enhancing the system's robustness and stability. By accurately identifying the user's emotional state, the system can provide more personalized services and interactive experiences, such as adjusting interaction methods based on the user's emotions and recommending suitable content, thus improving the user experience. This technical solution provides new ideas and methods for the field of affective computing, promoting its development and application. As an important branch of artificial intelligence, affective computing has broad application prospects and significant research value.

[0165] In summary, this technical solution integrates key information from facial images and speech signals, and employs a deep learning model and MFCC feature extraction method for emotion recognition processing, achieving accurate acquisition and output of the user's current emotional state information. This solution demonstrates significant technical effectiveness in improving emotion recognition accuracy, enhancing system robustness, optimizing user experience, and promoting the development of emotion computing technology.

[0166] In one embodiment of the present invention, the vehicle interaction system retrieves an in-vehicle environment parameter scheme matching the user's current emotional state from a database, and adjusts the in-vehicle environment according to the in-vehicle environment parameter scheme, including:

[0167] S501. After receiving the emotional state information, the vehicle interaction system extracts the emotional type corresponding to the maximum probability as the user's current emotional state.

[0168] S502. Retrieve in-vehicle environment parameter schemes that match the user's current emotional state from the database;

[0169] S503. Extract in-vehicle environmental parameters from the in-vehicle environmental parameter scheme, wherein the in-vehicle environmental parameters include temperature parameters and humidity parameters;

[0170] S504. Adjust the current in-vehicle environment according to the temperature and humidity parameters.

[0171] The working principle of the above technical solution is as follows: After receiving emotional state information, the vehicle interaction system first analyzes this information. Typically, emotional state information includes multiple emotional types and their corresponding probabilities. The system extracts the emotional type corresponding to the highest probability, which is taken as the user's current emotional state. This step is the foundation for subsequent adjustments to the in-vehicle environment. Based on the identified current emotional state, the system retrieves a matching in-vehicle environment parameter scheme from a pre-established database. This database may contain the correspondence between various emotional states and in-vehicle environment parameter schemes, each designed to create a more comfortable or emotionally suitable in-vehicle environment for the user. The system further extracts specific in-vehicle environment parameters from the retrieved schemes, typically including temperature and humidity parameters. These parameters form the basis for adjusting the in-vehicle environment. Finally, based on the extracted temperature and humidity parameters, the system adjusts the current in-vehicle environment by controlling the vehicle's air conditioning, humidification / dehumidification devices, and other equipment. The adjusted in-vehicle environment will better meet the user's emotional needs, thereby enhancing the user's driving experience.

[0172] The effects of the above technical solution are as follows: By adjusting the in-car environment according to the user's emotional state, the system can provide a more personalized driving experience. For example, when a user feels tense or anxious, the system can lower the in-car temperature and increase the humidity to create a quieter and more comfortable environment. This technical solution demonstrates the intelligence level of the vehicle interaction system. The system can perceive the user's emotional state in real time and make corresponding adjustments accordingly, making the vehicle more closely aligned with the user's needs and expectations. A comfortable in-car environment helps reduce driver fatigue and stress, thereby improving driving safety. For example, suitable temperature and humidity can reduce driver irritability, allowing them to focus more on driving tasks. This technical solution provides a new approach to the application of affective computing technology in the automotive field. By combining affective computing with the vehicle interaction system, the system can better understand the user's needs and feelings, thereby providing more considerate services.

[0173] In summary, this technical solution achieves multiple benefits by adjusting in-vehicle environmental parameters based on the user's emotional state, including improved user experience, enhanced vehicle intelligence, improved driving safety, and promotion of the application of affective computing technology. This solution not only improves the vehicle's intelligence level but also provides users with a more comfortable and personalized driving experience.

[0174] This invention proposes an interactive, AI-driven facial and voice analysis system, such as... Figure 2 As shown, the interactive AI-driven facial and voice analysis system includes:

[0175] The facial data acquisition module is used to collect the user's facial image data based on the current vehicle driving status, and to obtain the user's facial image data inside the car.

[0176] The voice data acquisition module is used to collect user voice data based on the current vehicle driving status and obtain voice data inside the car.

[0177] The emotional state information acquisition module is used to perform emotion recognition processing on the facial image data and voice data pair through a deep learning model to obtain the user's current emotional state information.

[0178] The emotional state information feedback module is used to feed back the user's current emotional state information to the vehicle interaction system;

[0179] The environmental parameter adjustment module is used by the vehicle interaction system to retrieve an in-vehicle environmental parameter scheme that matches the user's current emotional state from the database, and to adjust the in-vehicle environment according to the in-vehicle environmental parameter scheme.

[0180] The working principle of the above technical solution is as follows: The system first determines whether it is suitable to collect facial image data based on the current driving status of the vehicle (such as speed, whether it is in a congested area, etc.). When the conditions are met, the system uses the in-vehicle camera to capture the user's facial image data. This data will be used for subsequent emotion recognition processing. Similar to facial image data collection, the system also determines whether it is suitable to collect voice data based on the vehicle's driving status. When the conditions are met, the system captures the user's voice data through the in-vehicle microphone. This data will be used together with the facial image data for emotion recognition processing. The system inputs the collected facial image data and voice data into a deep learning model. This deep learning model has been trained on a large amount of data and can accurately identify the user's emotional state. By analyzing the user's facial expressions, voice tone, and other features, the model can output the user's current emotional state information. The system feeds back the identified user emotional state information to the vehicle interaction system. The vehicle interaction system uses this information to determine whether the in-vehicle environment needs to be adjusted to adapt to the user's emotional state. The vehicle interaction system retrieves an in-vehicle environment parameter scheme that matches the user's current emotional state from the database. The system adjusts the in-vehicle environment accordingly based on these parameters to create a more comfortable and safer driving environment.

[0181] The aforementioned technical solution achieves the following results: By collecting and analyzing users' facial images and voice data in real time, the system can accurately identify the user's emotional state and automatically adjust the in-car environment accordingly. This personalized adjustment will make users feel more comfortable and enjoyable while driving, thereby enhancing the driving experience. When users are fatigued or stressed, the system can automatically adjust the in-car environment to alleviate their emotions, thus reducing driving risks. Furthermore, by monitoring the user's emotional state in real time, the system can also promptly detect and remind the driver to pay attention to driving safety. The application of this technical solution will drive the development of automobiles towards greater intelligence and personalization. In the future, with continuous technological advancements and the deepening of applications, automobiles will become a more intelligent and human-centered mobile living space.

[0182] In conclusion, this interactive AI-driven facial and voice analysis method has significant technical effects and application prospects, which will provide drivers with a more intelligent and personalized driving experience and promote the intelligent development of the automotive industry.

[0183] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. An interactive artificial intelligence driven face and voice analysis method, characterized in that, The interactive artificial intelligence driven face and voice analysis method comprises: According to the current vehicle driving state, the user's face image data is collected, and the user's face image data in the car is obtained; According to the current vehicle driving state, the user's voice data is collected, and the voice data in the car is obtained; Through the deep learning model, the face image data and the voice data are subjected to emotional recognition processing, and the current emotional state information of the user is obtained; The current emotional state information of the user is fed back to the vehicle interaction system; The vehicle interaction system retrieves the in-vehicle environment parameter scheme matched with the current emotional state of the user from the database, and adjusts the in-vehicle environment according to the in-vehicle environment parameter scheme; According to the current vehicle driving state, the user's face image data is collected, which comprises: The driving speed of the current vehicle is monitored in real time; the driving speed of the current vehicle is compared with the preset first driving speed threshold and the second driving speed threshold; when the driving speed of the current vehicle does not exceed the preset first driving speed threshold, the face image data of the user is collected at a preset initial collection frequency; When the driving speed of the current vehicle exceeds the preset first driving speed threshold, but does not exceed the preset second driving speed threshold, the first collection frequency of the face image data of the user is set according to the current vehicle driving speed, and the face image data of the user is collected at the set first collection frequency of the face image data; When the driving speed of the current vehicle exceeds the preset second driving speed threshold, the second collection frequency of the face image data of the user is set according to the current vehicle driving speed, and the face image data of the user is collected at the set second collection frequency of the face image data; According to the current vehicle driving state, the user's voice data is collected, which comprises: The driving speed of the current vehicle is monitored in real time; the driving speed of the current vehicle is compared with the preset third driving speed threshold; wherein the third driving speed threshold is greater than the first driving speed threshold, but less than the second driving speed threshold; When the driving speed of the current vehicle does not exceed the preset third driving speed threshold, the voice data of the user is collected at a preset initial wake-up duration; When the driving speed of the current vehicle exceeds the preset third driving speed threshold, the driving speed of the current vehicle and the collection frequency of the face image data of the user are extracted; the driving speed coefficient corresponding to the current collection frequency of the face image data of the user is retrieved; the wake-up duration is obtained by combining the driving speed of the current vehicle with the collection frequency of the face image data of the user and the corresponding driving speed coefficient; wherein the wake-up duration is obtained by the following formula: wherein, T c represents a wake-up duration; T 0 represents a preset initial wake-up duration; F represents a collection frequency of the user's face image data corresponding to the current vehicle driving speed; S represents a driving speed coefficient corresponding to the current collection frequency of the user's face image data; F 0 represents a preset initial collection frequency; V 03 represents a preset third driving speed threshold; V represents a current vehicle driving speed; The voice data of the user is collected at the set wake-up duration.

2. The interactive artificial intelligence driven facial and voice analysis method as claimed in claim 1, wherein, According to the current vehicle driving speed, the first collection frequency of the face image data of the user is set, which comprises: When the driving speed of the current vehicle exceeds the preset first driving speed threshold but does not exceed the preset second driving speed threshold, real-time driving speed of the current vehicle is collected as first driving speed data within a first preset monitoring time period from the moment when the driving speed of the current vehicle exceeds the preset first driving speed threshold; wherein the value range of the first preset monitoring time period is 5min-12min; A first driving speed coefficient is obtained by using the first driving speed data; wherein the first driving speed coefficient is obtained by the following formula: wherein, S 01 represents a first driving speed coefficient; V 01 represents a preset first driving speed threshold value; V 02 represents a preset second driving speed threshold value; V represents a current driving speed of the vehicle; α 1 represents a maximum acceleration of the occurrence of the first preset monitoring time period; α max represents a maximum acceleration that the vehicle can reach; T s1 represents a cumulative duration in the acceleration state within the first preset monitoring time period; T 01 represents a first preset monitoring time period corresponding duration; A first collection frequency of the facial image data of the user is set by using the first driving speed coefficient; wherein the first collection frequency of the facial image data of the user is obtained by the following formula: wherein, F 01 a first collection frequency of face image data of the user; F 0 represents a preset initial collection frequency; S 01 a first driving speed coefficient; T s1 a cumulative duration of the vehicle in an accelerating state within a first preset monitoring time period; V z1 a vehicle driving speed corresponding to a termination time of the first preset monitoring time period; V q1 a vehicle driving speed corresponding to a starting time of the first preset monitoring time period.

3. The interactive artificial intelligence driven facial and voice analysis method as claimed in claim 1, wherein, A second collection frequency of the facial image data of the user is set according to the driving speed of the current vehicle, comprising: When the driving speed of the current vehicle exceeds the preset second driving speed threshold, real-time driving speed of the current vehicle is collected as second driving speed data within a second preset monitoring time period from the moment when the driving speed of the current vehicle exceeds the preset second driving speed threshold; wherein the value range of the second preset monitoring time period is 3min-8min; A second driving speed coefficient is obtained by using the second driving speed data; wherein the second driving speed coefficient is obtained by the following formula: wherein, S 02 represents a second driving speed coefficient; V 02 represents a preset second driving speed threshold value; V represents a driving speed of the current vehicle; α 2represents a maximum acceleration of the occurrence of the second preset monitoring time period; α max represents a maximum acceleration that the vehicle can reach; T s2 represents a cumulative duration in the acceleration state within the second preset monitoring time period; T 02 represents a corresponding duration of the second preset monitoring time period; A second collection frequency of the facial image data of the user is set by using the second driving speed coefficient; wherein the second collection frequency of the facial image data of the user is obtained by the following formula: wherein, F 02 a second collection frequency of the face image data of the user; F 0 represents a preset initial collection frequency; S 02 a second driving speed coefficient; T s2 a cumulative duration in the acceleration state within the second preset monitoring time period; V z2 a vehicle driving speed corresponding to the end time of the second preset monitoring time period; V q2 a vehicle driving speed corresponding to the start time of the second preset monitoring time period.

4. The interactive artificial intelligence driven facial and voice analysis method as claimed in claim 1, wherein, Emotion state information of the user is obtained by performing emotion recognition processing on the facial image data and the voice data pair by using a deep learning model, comprising: Image preprocessing is performed on the facial image data to obtain preprocessed facial image data, wherein the image preprocessing comprises image noise reduction processing, image enhancement processing and image normalization processing; Audio preprocessing is performed on the voice data to obtain audio preprocessed voice data, wherein the audio preprocessing comprises audio noise reduction processing, audio frame processing and volume enhancement processing; The preprocessed facial image data is input into a trained deep learning model to obtain a facial feature vector; wherein the deep learning model adopts a convolutional neural network face model; The audio preprocessed speech data is processed by using an MFCC feature extraction mode to obtain a speech feature matrix corresponding to the speech data; wherein a dimension of the speech feature matrix is N × S wherein, N represents a number of frames contained in the speech feature matrix; S represents a speech feature dimension of each frame. When each frame of audio after audio frame processing of the voice data has image data at the same frame time as the frame, the image data at the same frame time as each frame of audio is taken as target image data; When each frame of audio after audio frame processing of the voice data does not have image data at the same frame time as the frame, the image data corresponding to the image acquisition time closest in time to the frame time of each frame of audio is taken as target image data; The facial feature vector of the target image data is spliced with the corresponding voice feature matrix to obtain a comprehensive vector matrix; The comprehensive vector matrix is input into an emotion state analysis model for emotion recognition processing to obtain emotion state information of the user.

5. The interactive artificial intelligence driven facial and voice analysis method as claimed in claim 4, wherein, The structure of the emotion state analysis model is as follows: An input layer is used to receive the input comprehensive vector matrix; The first full connection layer is configured to perform preliminary processing and nonlinear transformation on the input integrated vector matrix, wherein the number of units of the first full connection layer is 128, and the activation function is a ReLU function. The second full connection layer is configured to perform further feature extraction on the vector matrix obtained after the processing of the first full connection layer, wherein the number of units of the second full connection layer is 256. The Dropout layer is configured to perform regularization processing on the data output by the second full connection layer, and discard a preset proportion of data information, wherein the preset proportion is in the range of 0.43-0.

52. The third full connection layer is configured to summarize the feature data in the data obtained by the Dropout layer, and obtain the summarized feature data. The output layer is configured to perform sentiment classification according to the summarized feature data, obtain a sentiment classification result, and output the sentiment classification result, wherein the sentiment classification result refers to the probability corresponding to a preset seven kinds of emotions, and the seven kinds of emotions include happiness, sadness, anger, surprise, fear, disgust, and calmness.

6. The interactive artificial intelligence driven facial and voice analysis method of claim 1, wherein, The vehicle interaction system retrieves a vehicle interior environment parameter scheme matching the current emotional state of the user from the database, and adjusts the vehicle interior environment according to the vehicle interior environment parameter scheme, including: After receiving the emotional state information, the vehicle interaction system extracts the emotional category corresponding to the maximum probability as the current emotional state of the user; Retrieve a vehicle interior environment parameter scheme matching the current emotional state of the user from the database; Extract vehicle interior environment parameters from the vehicle interior environment parameter scheme, wherein the vehicle interior environment parameters include temperature parameters and humidity parameters; Adjust the current vehicle interior environment according to the temperature parameters and humidity parameters.

7. An interactive artificial intelligence driven facial and voice analysis system characterized in that, The interactive artificial intelligence driven face and voice analysis system includes: A face data acquisition module is configured to acquire face image data of a user according to a current vehicle driving state, and obtain the face image data of the user in the vehicle; A voice data acquisition module is configured to acquire voice data of a user according to a current vehicle driving state, and obtain the voice data in the vehicle; An emotional state information acquisition module is configured to perform emotional recognition processing on the face image data and voice data through a deep learning model, and obtain the current emotional state information of the user; An emotional state information feedback module is configured to feed back the current emotional state information of the user to a vehicle interaction system; An environment parameter adjustment module is configured to retrieve a vehicle interior environment parameter scheme matching the current emotional state of the user from the database by the vehicle interaction system, and adjust the vehicle interior environment according to the vehicle interior environment parameter scheme; Wherein, the acquisition of the face image data of the user according to the current vehicle driving state includes: Real-time monitoring of the driving speed of the current vehicle; comparing the driving speed of the current vehicle with a preset first driving speed threshold and a second driving speed threshold; when the driving speed of the current vehicle does not exceed the preset first driving speed threshold, the face image data of the user is collected at a preset initial collection frequency. When the driving speed of the current vehicle exceeds the preset first driving speed threshold, but does not exceed the preset second driving speed threshold, a first collection frequency of the facial image data of the user is set according to the driving speed of the current vehicle, and the facial image data of the user is collected at the set first collection frequency of the facial image data; When the driving speed of the current vehicle exceeds the preset second driving speed threshold, a second collection frequency of the facial image data of the user is set according to the driving speed of the current vehicle, and the facial image data of the user is collected at the set second collection frequency of the facial image data; The collection of the voice data of the user according to the driving state of the current vehicle comprises: The driving speed of the current vehicle is monitored in real time, and the driving speed of the current vehicle is compared with a preset third driving speed threshold; wherein the third driving speed threshold is greater than the first driving speed threshold, but less than the second driving speed threshold; When the driving speed of the current vehicle does not exceed the preset third driving speed threshold, the voice data of the user is collected at a preset initial wake-up duration; When the driving speed of the current vehicle exceeds the preset third driving speed threshold, the driving speed of the current vehicle and the collection frequency of the facial image data of the user are extracted; the driving speed coefficient corresponding to the current collection frequency of the facial image data of the user is called; the wake-up duration is obtained by combining the driving speed of the current vehicle, the collection frequency of the facial image data of the user and the corresponding driving speed coefficient; wherein the wake-up duration is obtained by the following formula: wherein, T c represents the wake-up duration; T 0 represents a preset initial wake-up duration; F represents the collection frequency of the user's face image data corresponding to the current vehicle driving speed; S represents the driving speed coefficient corresponding to the current collection frequency of the user's face image data; F 0 represents a preset initial collection frequency; V 03 represents a preset third driving speed threshold; V represents the current vehicle driving speed; The voice data of the user is collected at the set wake-up duration.

Citation Information

Patent Citations

  • Vehicle owner emotion recognition and adjustment method, storage medium and vehicle-mounted system

    CN109190459A