Anonymization processing method and system for emotion data in vehicle

By performing local pre-processing and desensitization of multimodal emotional data on the in-vehicle edge computing platform, the communication delay and privacy leakage problems of traditional in-vehicle emotion recognition systems are solved, and efficient emotion recognition and privacy protection are achieved.

CN120509053APending Publication Date: 2025-08-19SHANGHAI PUFAFEN ELECTRONIC TECH CO LTD

Patent Information

Application Number
CN202510571019.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

Traditional in-vehicle emotion recognition systems have communication delays and user privacy leakage risks, and existing anonymization processing affects data availability or has low recognition accuracy.

Method used

Edge computing architecture is used to collect, preprocess, desensitize and identify emotional data in local nodes in the car. Through a multimodal fusion model, deep emotional feature extraction is performed on images, speech and physiological signals, and anonymous tag output mechanism is adopted to ensure the accuracy of emotion recognition and privacy protection.

Benefits of technology

It improves the system's real-time response capabilities, effectively blocks user identity information, ensures the privacy and security of passengers, and does not affect the accuracy of emotional recognition, and eliminates the transmission of original data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120509053A_ABST
    Figure CN120509053A_ABST
Patent Text Reader

Abstract

The invention provides an anonymization processing method and system for emotion data in a vehicle. The anonymization processing method comprises the steps of collecting multi-mode emotion data; correspondingly carrying out local preprocessing and multi-modal alignment processing on the multi-modal emotion data; corresponding sensitive emotion feature information in the preprocessed multi-modal emotion data is extracted through a lightweight recognition algorithm, and the sensitive emotion feature information at the recognized position is packaged in a unified mode; performing desensitization processing on various types of emotion modal data, and encapsulating and synchronizing desensitization results with labels; deep emotional feature extraction is performed on the image, the voice and the physiological signal through a multi-modal fusion model, and the extracted three types of deep features are fused to obtain an emotional representation vector; and inputting the emotion representation vector into a pre-trained emotion recognition model, and outputting the current emotion state of the passenger, including the specific emotion category and the corresponding confidence coefficient. According to the method, emotion analysis and transmission are performed after data anonymization is realized, and accurate judgment of the system on the emotion state is not influenced while privacy security of the user is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a method and system for anonymizing in-vehicle emotion data. Background Art

[0002] Traditional in-vehicle emotion recognition systems often rely on uploading raw image and voice data to cloud servers for processing. This not only presents communication latency issues but also carries a high risk of user privacy breaches. After emotion recognition is complete, if the output contains high-dimensional features or raw embedded data, it could be used to infer the user's identity, posing a privacy risk.

[0003] Patent document CN114041177A discloses a method for anonymizing vehicle data, particularly a motor vehicle operating method and a network server operating method. Based on detected vehicle data, a data set is created, the data set including information about the data detection location and / or the data detection time. The data set is anonymized based on traffic flow data by masking the location of the data detection location and / or the time of the data detection time. The traffic flow data is based on group information received by the vehicle providing the vehicle data from other vehicles via vehicle-to-vehicle communication.

[0004] However, the anonymization processing in patent document CN114041177A is achieved by shielding the data detection location and time information. This shielding affects the availability of data, and not all vehicles have the technical equipment to upload data, resulting in incomplete data collection.

[0005] Patent document CN118314930A discloses a design method for an intelligent in-vehicle emotion interaction system based on speech emotion recognition technology. At the user input layer, the voice signals and facial expression signals of the driver and passengers are collected through the hardware in the vehicle to form a multimodal emotion data input recognition system for emotion perception; at the technical architecture layer, an improved speech emotion recognition model is adopted as the main emotion recognition method, and the recognizability of the speech signal is improved through silent filters and independent component analysis. The accuracy and robustness of speech emotion recognition are improved by fusing MFCCs and openSMILE features into the machine learning model; at the interactive feedback layer, modules with emotion regulation and emotion monitoring feedback functions are mainly integrated to provide personalized emotional services for drivers and passengers.

[0006] However, the speech emotion recognition technology of patent document CN118314930A may be interfered by background noise and other factors in complex driving environments, affecting recognition accuracy and stability.

[0007] Patent document CN109190459A discloses a method for identifying and regulating the emotions of a vehicle owner, a storage medium, and an in-vehicle system. The method includes: collecting facial data of the vehicle owner and / or driving data of the vehicle; identifying and determining the emotions of the vehicle owner based on the collected facial data and / or driving data of the vehicle to obtain a determination result of the emotion of the vehicle owner; and calling a response device to respond based on the determination result of the emotion of the vehicle owner.

[0008] However, patent document CN109190459A judges the current owner's mood by setting thresholds on the collected facial data, voice data and driving data. Its accuracy still needs to be improved and there is a risk of user privacy leakage. Summary of the Invention

[0009] In view of the defects in the prior art, the purpose of the present invention is to provide a method and system for anonymizing in-vehicle emotion data.

[0010] A method for anonymizing in-vehicle emotion data provided by the present invention includes:

[0011] Step S1: Collect multimodal emotion data;

[0012] Step S2: performing local preprocessing and multimodal alignment processing on the multimodal emotion data;

[0013] Step S3: extracting the corresponding sensitive emotion feature information from the pre-processed multimodal emotion data through a lightweight recognition algorithm, and uniformly encapsulating the sensitive emotion feature information at the identified location;

[0014] Step S4: Desensitize each type of emotional modal data, and encapsulate the desensitization results and synchronize them with the labels;

[0015] Step S5: extracting deep emotion features from images, speech, and physiological signals respectively through a multimodal fusion model, and fusing the three types of extracted deep features to form a unified emotion representation vector;

[0016] Step S6: Input the emotion representation vector into a pre-trained emotion recognition model to output the occupant's current emotional state, including the specific emotion category and the corresponding confidence level.

[0017] Preferably, the multimodal emotion data includes image modality data, voice modality data and physiological modality data;

[0018] The image modality data is collected by multiple cameras in the vehicle, which can capture occupant facial images in real time and support visible light and infrared dual-channel imaging functions;

[0019] The voice modality data is collected through a distributed microphone array system, including voice input signals from the driver, co-driver, and rear seats;

[0020] The physiological modality data is collected by non-invasive physiological sensors integrated in the cockpit to obtain emotion-related physiological signals in real time;

[0021] Each modal data is automatically timestamped at the acquisition end and synchronized based on a time window through the edge node. The synchronized data is uniformly packaged into time-segment data packets for subsequent processing steps.

[0022] Preferably, the preprocessing of the image modality data includes image denoising, grayscale conversion and normalization, face alignment, expression region cropping, and timestamp synchronization;

[0023] The preprocessing of speech modality data includes denoising, silence detection and segmentation, amplitude normalization, short-time frame division and time alignment marking;

[0024] Preprocessing of physiological modality data includes filtering and noise reduction, outlier removal, sliding average, normalization, sampling and alignment to achieve temporal alignment of cross-modality data;

[0025] The multimodal alignment process introduces a unified timestamp mechanism and performs time alignment operations on all data based on a master time source, including sliding window synchronization, interpolation / cropping strategies, and synchronous data packet encapsulation.

[0026] Preferably, the sensitive feature identification of image modality data includes face area detection, facial key point extraction, facial feature vector extraction and emotion-irrelevant background feature removal, to obtain the face area position coordinates, key point index and embedded feature vector for desensitization processing;

[0027] Sensitive feature recognition of speech modal data includes voiceprint feature extraction, spectral component analysis, and speech feature labeling that is unrelated to semantic content. The output includes voiceprint vectors, spectral sensitive segments, and frequency band distribution parameters, which are used to control subsequent voiceprint perturbations or speech masking strategies.

[0028] Sensitive feature recognition of physiological modality data includes resting heart rate feature analysis, GSR skin conductance pattern recognition and heart rate variability HRV time domain / frequency domain features, and respiratory rhythm and amplitude analysis.

[0029] Preferably, the desensitization processing of the image modality data adopts Gaussian blurring, mosaic occlusion or replacing the face area with an embedded vector, including face blurring, area occlusion / replacement, embedding alternative representation and style transfer forgery;

[0030] The desensitization processing of speech modal data includes voiceprint perturbation processing, spectral feature mapping, speech pseudo-sound transformation and low-rank reconstruction de-featureization;

[0031] Desensitization of physiological modality data introduces differential privacy algorithms or dimensionality degradation processing, including normalization transformation, differential privacy perturbation, feature subspace selection, and segmented resampling;

[0032] The desensitization result encapsulation and label synchronization include repackaging the retained valid emotion features and synchronizing the corresponding emotion recognition labels and timestamps.

[0033] Preferably, deep emotion feature extraction from image modality data includes extracting key expression feature embeddings, including facial expression embedding vectors, dynamic expression change indicators, and attention heat maps, through a lightweight convolutional neural network or a visual Transformer model;

[0034] Deep emotion feature extraction from speech modality data involves extracting emotion indicators using a temporal modeling structure, including prosodic features, spectral features, rhythm and punctuation patterns, and emotional speech embedding vectors;

[0035] Deep emotion feature extraction from physiological modality data includes features related to heart rate variability, skin conductance fluctuations and emotional states.

[0036] Preferably, the emotion recognition model includes a multi-layer perceptron MLP structure, a multimodal Transformer structure, or builds an integrated classifier system, and realizes the post-decision integration of different modal recognition results through a Softmax fusion network or a SVM and random forest combination model.

[0037] Preferably, it also includes an anonymization output policy control step, which includes identity information stripping, label recoding, context desensitization synchronization and output permission control;

[0038] Identity information stripping refers to outputting only the desensitized emotion labels and emotion embedding vectors without outputting the original images, voices, and physiological signals;

[0039] The label recoding means that the emotion labels can be converted into numerical codes for external systems or third-party platforms;

[0040] The context desensitization synchronization refers to hiding the desensitization field in the system log or interaction record in combination with the privacy protection flag (privacy_flags) information generated by the current processing task.

[0041] According to the present invention, a system for anonymizing in-vehicle emotional data includes:

[0042] Module M1: Collect multimodal emotion data;

[0043] Module M2: performing local preprocessing and multimodal alignment processing on the multimodal emotion data;

[0044] Module M3: Extracts the corresponding sensitive emotion feature information from the pre-processed multimodal emotion data through a lightweight recognition algorithm, and uniformly encapsulates the sensitive emotion feature information at the identified location;

[0045] Module M4: Desensitizes each type of emotional modal data, and encapsulates the desensitization results and synchronizes them with the labels;

[0046] Module M5: Uses a multimodal fusion model to extract deep emotion features from images, speech, and physiological signals, and fuses the three extracted deep features to form a unified emotion representation vector.

[0047] Module M6: Input the emotion representation vector into the pre-trained emotion recognition model and output the occupant's current emotional state, including the specific emotion category and corresponding confidence level.

[0048] Preferably, the multimodal emotion data includes image modality data, voice modality data and physiological modality data;

[0049] The image modality data is collected by multiple cameras in the vehicle, which can capture occupant facial images in real time and support visible light and infrared dual-channel imaging functions;

[0050] The voice modality data is collected through a distributed microphone array system, including voice input signals from the driver, co-driver, and rear seats;

[0051] The physiological modality data is collected by non-invasive physiological sensors integrated in the cockpit to obtain emotion-related physiological signals in real time;

[0052] Each modal data is automatically timestamped at the acquisition end and synchronized based on the time window through the edge node. The synchronized data is uniformly packaged into time-segment data packets for subsequent processing module calls.

[0053] Compared with the prior art, the present invention has the following beneficial effects:

[0054] 1. This invention adopts an edge computing architecture to realize the entire process of emotional data collection, preprocessing, desensitization and recognition at the local node in the vehicle, avoiding the external transmission of sensitive data. It not only improves the real-time response capability of the system, but also blocks the path of emotional data leakage at the source, effectively protecting the privacy and security of passengers.

[0055] 2. The present invention introduces a modal desensitization mechanism into the emotion perception process, performing targeted de-identification processing on different modal data, thereby effectively shielding the user's identity information without affecting the accuracy of emotion recognition.

[0056] 3. This invention employs an anonymized label output mechanism, outputting only the occupant's current emotion category (e.g., "anxious," "calm," "happy," etc.) and its corresponding confidence score, eliminating the transmission of raw or reversible feature data. This ensures the system maintains high emotion recognition capabilities while minimizing the risk of identity exposure and ensuring full protection of user privacy throughout the entire process. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:

[0058] Figure 1 It is a schematic flow chart of the working method of the present invention. DETAILED DESCRIPTION

[0059] The present invention will be described in detail below with reference to specific embodiments. The following examples will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several changes and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.

[0060] The present invention uses a local multimodal emotion perception module to pre-process, identify features, and perform modal desensitization on images, voice, and physiological data, anonymizing the data before performing emotion analysis and transmission, ensuring user privacy while not affecting the system's accurate judgment of the emotional state. Emotional data usually contains identifiable identity information, such as facial images, voiceprint features, and individual physiological parameters. In order to prevent these data from being abused or reversely identified, the present invention introduces a modal desensitization mechanism into the emotion perception process, and performs targeted de-identification processing on different modal data. For example: image data is anonymized through face blurring, area occlusion, or feature vector extraction; voice data is anonymized by removing individual voice features through voiceprint perturbation or reconstruction; and physiological data uses differential privacy perturbation or dimensionality reduction algorithms to hide individual parameters. The above processing method effectively shields the user's identity information without affecting the accuracy of emotion recognition.

[0061] Example 1

[0062] According to the present invention, a method for anonymizing in-vehicle emotional data is provided. Figure 1 Shown, including:

[0063] Step S1: Collect multimodal emotion data. Comprehensively collect occupant emotion-related data using a variety of onboard sensors. This multimodal emotion data includes image, voice, and physiological modal data. The multimodal raw data is synchronously transmitted to the corresponding edge processing step S2.

[0064] The image modality data is collected by multiple cameras in the car. The camera deployment positions can be selected from the lower part of the A-pillar, above the center console, rearview mirror housing, etc. The camera can capture the facial image of the occupant in real time and supports visible light and infrared dual-channel imaging functions. Among them, infrared imaging is suitable for special working conditions such as night and strong backlight to ensure stable imaging quality. The frame rate of the collected image is not less than 15fps, and the resolution is set to 720p to ensure the recognizability of facial details. All image data is cached by the local edge device after collection and recorded with a timestamp to cooperate with subsequent multi-modal synchronous processing.

[0065] Speech modality data is collected via a distributed microphone array system, including voice input signals from the driver, front passenger, and rear seats. The microphone array captures the natural speech input of passengers and optimizes spatial orientation using beamforming technology to effectively distinguish the speech sources of different passengers. The collected data is either raw PCM or compressed audio streams, with a sampling rate typically set between 16kHz and 48kHz to ensure the integrity of the speech spectrum.

[0066] The physiological modality data is collected by non-invasive physiological sensors integrated in the cockpit, ensuring that emotion-related physiological signals are obtained in real time without interfering with the behavior of the occupants. For example, physiological sensors embedded in the steering wheel, seat backs or seat belts. The physiological signals include heart rate, skin conductance, and respiratory rhythm signals. Specifically, the heart rate signal can be obtained by the built-in electrodes in the hand-held steering wheel, or the capacitive sensors integrated inside the seat cushion and backrest, so as to monitor the heart rate changes of the driver or passenger; the skin conductance (GSR) signal is collected through hand contact electrodes or wearable devices to reflect the emotional excitement level and psychological stress state of the occupants; in addition, the body surface temperature and respiratory rhythm can be indirectly detected with the help of thermal infrared sensors or seat pressure sensors arranged in the cockpit, providing supporting physiological feature data for emotion recognition. These physiological modality data, together with image and voice signals in the present invention, constitute multimodal input to improve the comprehensiveness and accuracy of the emotion recognition model.

[0067] To ensure temporal consistency between image, speech, and physiological data, each modality is automatically timestamped at the acquisition end and synchronized using time-windowed processing (e.g., a sliding alignment algorithm) via edge nodes. The synchronized data is then packaged into time-aligned multi-modal packets (TAMPs) for subsequent processing modules. Depending on privacy and computing power, edge nodes can be configured for local processing or partially encrypted upload to the cloud.

[0068] Step S2: The edge node receives the multimodal emotion data and immediately performs local preprocessing and multimodal alignment on the multimodal emotion data to improve subsequent processing efficiency and model adaptability. This step ensures that all input data has a uniform format and quality for subsequent processing.

[0069] The preprocessing of image modality data includes image denoising, grayscale and normalization, face alignment, expression area cropping and timestamp synchronization. The image denoising includes performing noise reduction processing such as Gaussian filtering and median filtering on the camera image to remove image noise caused by light fluctuations and vehicle vibration. The grayscale and normalization include converting the image into a grayscale image or a unified color channel, and normalizing the pixel values to the [0,1] interval to facilitate neural network processing. The face alignment uses a facial key point detection algorithm (such as MTCNN, Dlib) to extract facial feature points, perform affine transformation or rotation correction, and standardize facial posture. The expression area cropping automatically identifies key areas such as the eyes and mouth, crops out the area used for emotion recognition, and reduces redundant background interference. The timestamp synchronization adds precise acquisition time to each frame of the image, and unifies the time axis with other modalities.

[0070] The preprocessing of speech modal data includes denoising, silence detection and segmentation, amplitude normalization, short-time frame division and time alignment marking. The denoising process uses spectral subtraction to remove environmental noise and improve speech clarity. Silence detection and segmentation are based on energy threshold and zero crossing rate to identify silence segments and extract effective segments of speech signals. Amplitude normalization is to unify the amplitude range of speech signals to avoid the subsequent model being sensitive to speech strength. Short-time frame division is to divide speech into frame segments of fixed duration (such as 25ms frame length, 10ms step length), providing window support for spectrum analysis and feature extraction. Time alignment marking is to record the starting timestamp of each speech frame segment for use by the multimodal fusion model in frame synchronization.

[0071] The preprocessing of physiological modality data includes filtering and denoising, anomaly removal, sliding average, normalization, sampling and alignment to achieve time series alignment of cross-modal data. The filtering and denoising uses a bandpass filter (such as 0.5Hz to 5Hz) to remove power frequency interference and low-frequency drift for signals such as heart rate and GSR. Anomaly removal is based on statistical rules to eliminate outliers and acquisition errors (such as heart rate jumps and signal faults). Sliding average uses a sliding window to smooth signal changes and improve signal stability. Normalization is to standardize the signal according to the user's baseline value (such as resting heart rate) to improve the model's generalization ability for different individuals. Sampling and alignment are to unify the sampling frequency (such as 10Hz or 20Hz) and interpolate to fill in the missing points so that the time axis is consistent with the image and speech modalities.

[0072] In order to ensure that the data of each modality are synchronized in the time dimension, the multimodal alignment process of the present invention introduces a unified timestamp mechanism and performs time alignment operations on all data based on the master time source, including sliding window synchronization, interpolation / cropping strategy and synchronous data packet encapsulation. The sliding window synchronization includes setting the image frames, voice segments and physiological samples within the synchronization window (such as 1 second) as the same emotion perception unit. The interpolation / cropping strategy refers to interpolation filling or cropping of modalities with inconsistent sampling rates to ensure comparability between modalities. The synchronous data packet encapsulation is to encapsulate all aligned modal data into a data packet with a standard structure (such as JSON / binary structure), which is handed over to subsequent modules for unified processing.

[0073] Step S3: Use a lightweight recognition algorithm to extract the corresponding sensitive emotion features from the pre-processed multimodal emotion data and encapsulate the identified sensitive emotion features. The output of this step is a list of sensitive features that may cause privacy leaks, which serves as the basis for subsequent desensitization processing.

[0074] Facial key points and facial feature vectors are identified in image modality data. The recognition results will output the facial region location coordinates, key point indexes, and embedded feature vectors, which serve as the basis for subsequent image desensitization operations. Specifically, face monitoring and key point recognition algorithms are used to identify identity-related information, including facial region detection, facial key point extraction, facial feature vector extraction, and emotion-irrelevant background feature removal. Facial region detection uses lightweight convolutional neural networks (such as MTCNN and YOLO-face) to accurately locate the face. Facial key point extraction includes identifying key areas such as eyebrows, eyes, nose tips, and mouth corners to assist in subsequent facial expression analysis and serve as potential identity-recognizing features. Facial feature vector extraction uses pre-trained facial recognition models (such as FaceNet and ArcFace) to obtain a 128-dimensional or 512-dimensional facial embedding vector. This vector has strong identity discrimination capabilities and needs to be marked as a highly sensitive feature. Emotion-irrelevant background feature removal is to identify non-facial areas in the image and remove data that is not related to privacy.

[0075] Voiceprint parameters and speech spectral features are identified in speech modal data. Outputs include voiceprint vectors, spectral sensitive segments, and frequency band distribution, which are used to control subsequent voiceprint perturbations or speech masking strategies. This means identifying voiceprint information that can be used for speaker identity tracking and speech features not directly related to emotion recognition. This involves voiceprint feature extraction, spectral component analysis, and speech feature markers unrelated to semantic content. Voiceprint feature extraction involves extracting voiceprint embedding representations using spectrogram-based convolutional models or i-vector / x-vector methods to identify speaker-unique characteristics. Spectral component analysis involves analyzing personalized acoustic features in speech, such as fundamental frequency, formant structure, and vocal tract characteristics. These components often carry physiological characteristics of the speaker. Speech feature markers unrelated to semantic content, such as non-emotional components such as speech amplitude and speech rate variation patterns, can serve as candidate features for desensitization optimization.

[0076] Physiological modality data is analyzed for signal features strongly associated with identity, such as resting heart rate and individual breathing patterns. Although physiological signals are not as intuitive as images and voice, some individual-specific physiological patterns still have the potential for identity recognition and need to be identified and labeled. Specifically, this includes resting heart rate feature analysis, GSR (skin conductance) pattern recognition, heart rate variability (HRV) time / frequency domain features, and respiratory rhythm and amplitude analysis. Resting heart rate feature analysis refers to statistically analyzing the average heart rate and heart rate stability within a specific time window, which can vary significantly between individuals. GSR (skin conductance) pattern recognition involves statistical modeling of the amplitude and peak response frequency of skin conductance. Heart rate variability (HRV) time / frequency domain features, such as SDNN, RMSSD, and LF / HF ratio, can be used for health or identity analysis. Respiratory rhythm and amplitude analysis uses respiratory waveforms acquired via pressure sensors or thermal infrared signals, which have a certain degree of individual recognition. The present invention extracts statistical features from the above signals and sets a configurable sensitivity threshold (for example, when the variability exceeds a certain standard value, it is determined to be sensitive), outputting feature fields and sampling time periods.

[0077] The sensitive feature information identified by the above three modalities will be uniformly encapsulated. This structured output will be passed to the next modal desensitization processing step S4 as the basis for the desensitization operation. The encapsulation structure is as follows:

[0078] Sensitive feature information encapsulation structure of image modality data:

[0079] image:

[0080] face_box:[x1,y1,x2,y2],

[0081] landmarks:[...],

[0082] embedding:[...]

[0083] Among them, face_box represents the coordinates of the detected face bounding box, in the format of [x1, y1, x2, y2], corresponding to the coordinates of the upper left corner (x1, y1) and the lower right corner (x2, y2) of the face rectangle, respectively, and is used to locate the face area. Landmarks represents a collection of key feature points on the face, such as the corners of the eyes, the tip of the nose, and the corners of the mouth, which are used to assist in face alignment, feature extraction, and anonymization. Embedding: represents the facial feature vector extracted from the face image. It is a low-dimensional facial representation that can be used for emotion recognition and anonymization, and does not directly save the original image.

[0084] Sensitive feature information encapsulation structure of speech modal data:

[0085] voice:

[0086] voiceprint_vector:[...],

[0087] sensitive_freq_band:[start_freq,end_freq]

[0088] Among them, voiceprint_vector: represents the voiceprint feature vector extracted from the passenger's voice data, which is used to describe the speaker's voice characteristics. The voiceprint can be anonymized through subsequent perturbation or encryption. sensitive_freq_band represents the sensitive frequency range in the voice signal that requires special protection. The format is [start_freq, end_freq], indicating the frequency band in the voice spectrum that is most likely to leak identity information.

[0089] Sensitive feature information encapsulation structure of physiological modality data:

[0090] bio:

[0091] hr_mean: Anonymization of in-car emotional data,

[0092] hrv:{...},

[0093] gsr_pattern:high_variance

[0094] Where hr_mean represents the average heart rate over a period of time (Heart Rate Mean), hrv represents the heart rate variability data (Heart Rate Variability), and gsr_pattern represents the pattern characteristics of the galvanic skin response (GSR).

[0095] Step S4: After identifying sensitive features, each type of emotional modality data is desensitized and the desensitization results are packaged and synchronized with the label. This step ensures that the data does not contain traceable identity information during subsequent processing and output.

[0096] The desensitization processing of image modality data adopts Gaussian blur, mosaic occlusion or replacement of facial areas with embedded vectors, including face blurring, regional occlusion / replacement, embedded replacement representation and style transfer forgery. The face blurring includes applying Gaussian blur or mean filtering to the face detection area, so that the original facial features cannot be directly identified. Regional occlusion / replacement is to use the facial key point positioning results to cover the eyes, nose and mouth areas with mosaics or uniform graphic masks (such as masks, cartoon images). Embedded replacement representation is to directly extract facial embedding vectors (such as FaceNet 128 dimensions) and use them as input for subsequent models, while discarding the original face image and retaining only emotion-related representations. Style transfer forgery is to replace real faces with synthetic faces using image style transfer networks such as GAN while maintaining the expression structure to improve usability and privacy protection compatibility.

[0097] The desensitization of speech modal data is achieved by perturbing voiceprint features, spectral stretching, or extracting MFCC (Mel-Frequency Cepstral Coefficients), and includes voiceprint perturbation processing, spectral feature mapping, speech pseudo-voice conversion, and low-rank reconstruction de-featureization. The voiceprint perturbation processing uses an audio perturbation algorithm to fine-tune the voiceprint components in the original speech, such as spectral perturbation of positions such as formants and fundamental frequencies, so that they deviate from individual characteristics but retain the trend of intonation changes. Spectral feature mapping involves converting the original speech signal into non-speaker identification features such as Mel-Frequency Cepstral Coefficients (MFCC), Chroma, and Spectral Contrast, and then using them as model input. Speech pseudo-voice conversion uses voice conversion technology to map the original speech into "anonymous speech," which sounds close to real speech but cannot be traced back to the identity through voiceprint recognition. Low-rank reconstruction de-featureization involves constructing a time-frequency map matrix for the audio signal and performing low-rank decomposition, retaining only the emotional expression dimension and eliminating individual differences.

[0098] The desensitization of physiological modality data introduces differential privacy algorithms or dimensionality degradation processing to reduce the risk of individual identification, including normalization transformation, differential privacy perturbation, feature subspace selection and segmented resampling. The normalization transformation uses the average value of the occupant in the current session as the benchmark to normalize signals such as HR and GSR, avoiding the use of individual inherent parameters such as resting state. Differential privacy perturbation adds Laplace noise or Gaussian noise to key indicators (such as heart rate mean, HRV, etc.) to achieve statistical irreversibility. Feature subspace selection is to retain only the principal components that are strongly correlated with emotional changes through dimensionality reduction methods such as PCA, filtering out individual difference dimensions. Segmented resampling is to resample and interpolate the signal without changing the time trend, destroying its matching ability with the original individual characteristics.

[0099] The desensitization result encapsulation and label synchronization include repackaging the retained valid multimodal emotion data (such as expression embedding, speech spectrum features, HR change curves), and synchronizing the corresponding emotion recognition labels (such as excitement, fatigue, pleasure, etc.) and timestamps. Specifically, the valid multimodal emotion data refers to the retained multimodal emotion data remaining after the pre-processed multimodal emotion data removes sensitive emotion feature information. The emotion recognition label is manually signed, and the data is classified and labeled according to the values of image_embedding and mfcc_features before training. The encapsulation structure of the processing result includes time information, feature information and privacy protection flag (privacy_flags), and the corresponding format is as follows:

[0100] The time information is as follows:

[0101] time:2025-04-18T15:20:00Z

[0102] "Time" indicates the timestamp of data collection, using the standard ISO 8601 time format (for example, "2025-04-18T15:20:00Z", which means 15:20:00 on April 18, 2025, Coordinated Universal Time (UTC). This is used for subsequent tracking and synchronous analysis of emotional states at different time points.

[0103] The characteristic information is as follows:

[0104] features:

[0105] image_embedding:[...]

[0106] mfcc_features:[...]

[0107] normalized_hr:[...]

[0108] emotion_label:anxious

[0109] Image_embedding represents the anonymized image feature vector (embedding) extracted from facial images, used for emotion recognition or classification while ensuring that the original facial image is not directly exposed. MFCc_features represents the Mel-Frequency Cepstral Coefficient (MFCC) features extracted from speech signals, used to analyze the emotional state of the occupant's speech and suitable for retaining sufficient emotional information after the speech data is anonymized. Emotion_label represents the emotion label inferred based on the current multimodal features. In this example, "anxious" is used to classify the emotional state.

[0110] The privacy protection flags (privacy_flags) are as follows:

[0111] privacy_flags:

[0112] face_removed:true,

[0113] voice_perturbed:true,

[0114] bio_masked:true

[0115] Among them, face_removed is a Boolean value. If it is true, it means that the original facial image data has been removed during processing, and only anonymized features (such as embeddings) are retained to protect privacy. voice_perturbed is a Boolean value. If it is true, it means that the original voice data has been perturbed or noised to prevent the occupant's identity from being restored through voiceprint recognition. bio_masked is a Boolean value. If it is true, it means that the original physiological data (such as heart rate and skin conductivity) has been masked or encrypted to prevent physiological features from revealing the individual's identity.

[0116] Step S5: After desensitization is completed, deep emotion features are extracted from images, voice, and physiological signals respectively through a multimodal fusion model, and finally the three types of features are fused to form a unified emotion representation vector.

[0117] Deep emotion feature extraction from image modality data includes extracting key expression feature embeddings through a lightweight convolutional neural network or a visual Transformer model. Specifically, facial information that can identify the identity has been removed from the image modality, and areas that reflect changes in expression (such as eyes, eyebrows, corners of the mouth, etc.) are retained. The present invention uses a lightweight convolutional neural network (CNN) or a visual Transformer model to extract the following key expression features. The key expression feature embedding includes a facial expression embedding vector, a dynamic expression change index, and an attention heat map. The facial expression embedding vector is obtained by inputting a desensitized image or a key point heat map into a convolutional network, and outputting an expression vector with a dimension of 64 to 256, which represents emotions such as anger, surprise, fatigue, sadness, etc. The dynamic expression change index is obtained by performing a differential operation on the emotion embedding changes of consecutive image frames to extract an "emotion mutation index" or a "micro-expression dynamic index." The attention heat map assigns weights to image areas through a multi-head attention mechanism to identify local areas with the most significant emotional expressions.

[0118] Deep emotion feature extraction from speech modal data includes the use of time series modeling structures (such as LSTM) to extract emotional indicators such as intonation and rhythm. In the speech modality, acoustic features related to emotional expression are extracted from the desensitized audio (such as MFCC features, spectral perturbation results), including prosodic features, spectral features, rhythm and punctuation patterns, and emotional speech embedding vectors. The prosodic features include pitch contour, volume intensity, and speaking rate. Spectral features include Mel-frequency cepstral coefficients (MFCC), resonance peak distribution, and spectral energy density. Rhythm and punctuation patterns are used to extract pause duration, speech block distribution, etc., which are used to detect states such as anxiety, hesitation, and excitement. The emotional speech embedding vector is to input speech features into a trained emotional audio recognition network (such as CRNN, Wav2Vec2.0) to output a unified emotional embedding representation.

[0119] Deep emotion feature extraction from physiological modality data includes features related to emotional states, such as heart rate variability and skin conductance fluctuations. After desensitization, physiological signals still contain dynamic fluctuation features that reflect emotional changes, which are extracted using the following methods: Heart rate change features: extract the average heart rate per second, heart rate fluctuation range, and rapid change trends. HRV (heart rate variability) features: such as SDNN (standard deviation), RMSSD (root mean square deviation), LF / HF (low-frequency to high-frequency energy ratio). GSR (skin conductance) response features: such as the number of fluctuations per unit time, response time, and peak amplitude. Comprehensive physiological stress index: Generate a stress level score through multi-parameter fusion (heart rate + GSR + breathing) to assist in identifying states such as anxiety, fright, and tension.

[0120] The feature fusion method includes early fusion, which concatenates the three types of modal features before model input, and performs weighted normalization to form a joint input vector. Intermediate fusion, after processing different modalities separately in the intermediate layer, performs cross-modal association modeling through a shared attention mechanism or a gating mechanism. Late fusion: After independently identifying emotions for the three modalities, decision fusion is performed based on weighted voting or confidence averaging. After feature fusion, an emotion embedding vector of uniform length (for example, a dimension of 128 or 256) is output to represent the comprehensive emotional state of the current occupant.

[0121] Step S6: Emotion recognition and anonymization output. The emotion representation vector is input into the pre-trained emotion recognition model to output the occupant's current emotional state, including specific emotion categories (such as happiness, anger, sadness, etc.) and corresponding confidence levels. The model output only contains anonymized emotion results and does not contain any identity features. This result can be used for intelligent responses of in-vehicle systems, such as air conditioning adjustment, lighting changes, voice prompt optimization, etc., to achieve an emotion-driven human-computer interaction experience, and ensure data privacy and security throughout the process. Step S6 includes the following steps:

[0122] Step S6.1: The emotion recognition module receives a unified feature vector derived from the fusion of image, speech, and physiological signals and inputs it into a trained multimodal emotion recognition model for classification and recognition. This model can employ a multi-layer perceptron (MLP) architecture, suitable for processing aligned and concatenated low-dimensional feature vectors, offering advantages such as high computational efficiency and fast response. Alternatively, a multimodal Transformer architecture can be employed, using a self-attention mechanism to model deep semantic connections between image, speech, and physiological signals, thereby improving the accuracy of recognizing complex emotional states. Furthermore, an integrated classifier system can be constructed, employing a Softmax fusion network or a combined SVM and random forest model to enable post-decision integration of recognition results from different modalities, further improving the robustness and stability of the overall discrimination. Ultimately, the model outputs a specific emotion category label (e.g., calm, happy, angry, anxious, tired, surprised, etc.), along with an emotion confidence score between 0 and 1, which measures the credibility of the recognition result and provides an effective basis for subsequent in-vehicle environment adjustment or interaction strategies.

[0123] Step S6.2: Structural packaging of the recognition results, in the following format:

[0124] "timestamp":"2025-04-18T15:48:22Z",

[0125] "emotion_label":"angry",

[0126] "confidence":0.87,

[0127] "emotion_vector":[...], / / emotion embedding representation

[0128] "source_modalities":["image","voice","bio"]

[0129] Here, emotion_label represents the final determined occupant emotion category, confidence represents the confidence of the current judgment, emotion_vector is the traceable embedded feature after desensitization, and source_modalities records the modal sources involved in this judgment for later analysis.

[0130] Step S6.3: Anonymization control and output strategy. To further protect user privacy and security, this step sets an anonymous output strategy control step in the output link, including identity information stripping, label recoding, context desensitization synchronization, and output permission control. Identity information stripping refers to outputting only the desensitized emotion labels and emotion embedding vectors, without outputting the original images, voice, and physiological signals. Label recoding means that for external systems or third-party platforms, emotion labels can be converted into numerical codes (such as Angry→Code_03) to avoid the leakage of emotion-sensitive labels. Context desensitization synchronization refers to hiding the desensitized fields in system logs or interaction records in combination with the privacy protection flag (privacy_flags) information generated by this processing task. Output permission control refers to setting the output channel permissions for local and cloud, for example, the local area can obtain complete labels, and the cloud only receives aggregated results.

[0131] The purpose of the present invention is to effectively strip and desensitize sensitive information related to personal identity while ensuring the accuracy of multimodal emotion recognition.

[0132] Example 2

[0133] Based on Example 1, the collected multimodal emotion data is anonymized in the edge computing device to ensure that the data is irreversibly associated with an individual's identity. This can also be achieved through the following steps:

[0134] Step 1: For image modality data, the MediaPipe model is used to extract the coordinates of feature points for sentiment analysis calculation through the facial key point detection model, and the areas around the feature points are anonymized.

[0135] Step 2: For the speech modality data, based on the pre-trained DCN clustering model, the mixed features are separated into voiceprint feature vectors and emotion feature vectors.

[0136] Step 3: For physiological modality data, a homomorphic encryption algorithm based on the CKKS scheme is used to support the model to directly calculate the emotion index in the ciphertext state. The CKKS-based ciphertext emotion calculation includes the following steps:

[0137] Step 3.1: Normalize the physiological signal into a numerical value, and then use the CKKS public key to encrypt the normalized numerical vector to generate the ciphertext c_data.

[0138] Step 3.2: Calculate the ciphertext sentiment and homomorphic mean: Perform homomorphic addition and plaintext scalar multiplication on the c_data ciphertext to obtain the encrypted mean c_mean. Calculate the homomorphic standard deviation: Homomorphic variance C_var = C_mean_square (homomorphically calculated square mean) - (C_mean)2. Calculate the encrypted standard deviation c_std using Taylor expansion. Sentiment value = 0.6*c_mean + 0.4*c_std

[0139] Step 3.3: The authorized air conditioning control system module can call the CKKS private key decryption through the HSM to obtain the plaintext emotion value.

[0140] Step 4: The encrypted data is divided into blocks and stored in a secure storage unit on the vehicle. The encryption key is managed by the vehicle's TPM chip.

[0141] Example 3

[0142] Based on Example 1 and Example 2, after the anonymization process, the method further includes real-time monitoring of abnormal access based on the results of the anonymization process. When an abnormal access occurs, it is determined whether the current abnormal access is a misjudgment. If not, dynamic blocking is performed, that is, active defense and attack countermeasures are performed; if so, normal user operations that are judged to be abnormal are extracted from the security logs generated during the monitoring process, and the tracking watermark records triggered by the attacker are captured as dynamically optimized data, the attack behavior and the type of false positive are automatically labeled, the fuzzy samples are reconfirmed during manual review, and the training is optimized through incremental learning. Specifically, normal user operations that are judged to be abnormal by the LSTM model are extracted from the security logs, and the tracking watermark records triggered by the attacker are captured as dynamically optimized data, the attack behavior and the type of false positive are automatically labeled, the fuzzy samples are reconfirmed during manual review, and the behavior detection model of the LSTM is optimized through incremental learning. The incremental learning includes inputting historical behavior data and new samples, freezing the underlying network of the LSTM, updating only the fully connected layer parameters, and regularly training new models offline every week, or triggering real-time online learning when a new attack pattern is detected.

[0143] Abnormal access includes unauthorized IP access, high-frequency requests, illegal protocols, and low-frequency crawling. The active defense and attack countermeasures include the following steps:

[0144] Step 1: Inject fake emotional data containing a tracking watermark into the data stream. When an attacker triggers the watermark, a warning is triggered and the attacker's access history is recorded. The watermark ID includes the IP address, device fingerprint, and timestamp. The fake emotional data includes virtual physiological signals and virtual facial / voice information. The virtual physiological signals are false data that conforms to a normal distribution but has abnormal parameters. The virtual facial / voice information is generated using a GAN.

[0145] Step 2: The data transmission link is superimposed with random noise packets. The noise packets are consistent with the length and transmission frequency of the real data packets, thereby interfering with the attacker's data parsing and cracking process.

[0146] Step 3: After detecting an attack, the data encryption key and watermark ID are automatically updated, and other data are checked for leakage, and the affected nodes are automatically isolated.

[0147] Example 4

[0148] The present invention also provides an anonymization processing system for in-vehicle emotional data. The anonymization processing system for in-vehicle emotional data can be implemented by executing the process steps of the anonymization processing method for in-vehicle emotional data. That is, those skilled in the art can understand the anonymization processing method for in-vehicle emotional data as a preferred implementation of the anonymization processing system for in-vehicle emotional data.

[0149] According to the present invention, a system for anonymizing in-vehicle emotional data includes:

[0150] Module M1: Collects multimodal emotion data; the multimodal emotion data includes image modal data, voice modal data and physiological modal data; the image modal data is collected by multiple cameras in the car, which can capture the facial images of the occupants in real time and support visible light and infrared dual-channel imaging functions; the voice modal data is collected through a distributed microphone array system, including voice input signals from the driver, co-driver and rear seats; the physiological modal data is collected through non-invasive physiological sensors integrated in the cockpit to obtain emotion-related physiological signals in real time; each modal data is automatically timestamped at the collection end, and time window-based synchronization processing is performed through the edge node. The synchronized data will be uniformly packaged into time segment data packets for subsequent processing module calls.

[0151] Module M2: local preprocessing and multimodal alignment processing are performed on the multimodal emotion data; the preprocessing of image modal data includes image denoising, grayscale and normalization, face alignment, expression area cropping and timestamp synchronization; the preprocessing of speech modal data includes denoising, silence detection and segmentation, amplitude normalization, short-time frame division and time alignment marking; the preprocessing of physiological modal data includes filtering and noise reduction, anomaly removal, sliding average, normalization and sampling and alignment to achieve time series alignment of cross-modal data; the multimodal alignment processing is to introduce a unified timestamp mechanism and perform time alignment operations on all data based on the master time source, including sliding window synchronization, interpolation / cropping strategy and synchronous data packet encapsulation.

[0152] Module M3: The corresponding sensitive emotional feature information in the pre-processed multimodal emotional data is extracted through a lightweight recognition algorithm, and the sensitive emotional feature information at the identification point is uniformly encapsulated; the sensitive feature recognition of image modal data includes face area detection, facial key point extraction, facial feature vector extraction and emotion-irrelevant background feature removal, and the face area position coordinates, key point index and embedded feature vector are obtained as the basis for desensitization processing; the sensitive feature recognition of speech modal data includes voiceprint feature extraction, spectral component analysis, and speech feature labeling that is unrelated to semantic content. The output includes voiceprint vectors, spectral sensitive segments, and frequency band distribution parameters, which are used to control subsequent voiceprint perturbations or speech masking strategies; the sensitive feature recognition of physiological modal data includes resting heart rate feature analysis, GSR skin conductance pattern recognition and heart rate variability HRV time domain / frequency domain features, respiratory rhythm and amplitude analysis.

[0153] Module M4: Desensitize each type of emotional modal data separately, and encapsulate the desensitization results and synchronize with the labels; the desensitization of image modal data uses Gaussian blur, mosaic occlusion or replaces the face area with an embedded vector, including face blurring, area occlusion / replacement, embedded replacement representation and style transfer forgery; the desensitization of speech modal data includes voiceprint perturbation processing, spectral feature mapping, speech pseudo-sound transformation and low-rank reconstruction defeatureization; the desensitization of physiological modal data introduces differential privacy algorithms or dimensionality degradation processing, including normalization transformation, differential privacy perturbation, feature subspace selection and segmented resampling; the desensitization result encapsulation and label synchronization include repackaging the retained valid emotional features and synchronizing the corresponding emotion recognition labels and timestamps.

[0154] Module M5: Deep emotion feature extraction is performed on images, speech and physiological signals respectively through a multimodal fusion model, and the three types of extracted deep features are fused to form a unified emotion representation vector; deep emotion feature extraction of image modal data includes extracting key expression feature embeddings through lightweight convolutional neural networks or visual Transformer models, including facial expression embedding vectors, dynamic expression change indicators and attention heat maps; deep emotion feature extraction of speech modal data includes using time series modeling structures to extract emotion indicators, including rhythmic features, spectral features, rhythm and punctuation patterns and emotional speech embedding vectors; deep emotion feature extraction of physiological modal data includes heart rate variability, skin conductance fluctuations and features related to emotional state.

[0155] Module M6: Input the emotion representation vector into the pre-trained emotion recognition model, and output the current emotional state of the occupant, including the specific emotion category and the corresponding confidence. The emotion recognition model includes a multi-layer perceptron MLP structure, a multimodal Transformer structure, or builds an integrated classifier system, and realizes the post-decision integration of different modal recognition results through a Softmax fusion network or a SVM and random forest combination model. It also includes an anonymized output strategy control module, which includes identity information stripping, label recoding, context desensitization synchronization, and output authority control; the identity information stripping refers to outputting only the desensitized emotion label and emotion embedding vector, without outputting the original image, voice, or physiological signal; the label recoding refers to the fact that the emotion label can be converted into a numerical code for an external system or a third-party platform; the context desensitization synchronization refers to hiding the desensitized field in the system log or interaction record in combination with the privacy protection flag (privacy_flags) information generated by the current processing task.

[0156] Those skilled in the art will appreciate that, in addition to implementing the system and its various devices, modules, and units provided by the present invention in purely computer-readable program code, it is entirely possible to implement the same functions of the system and its various devices, modules, and units provided by the present invention in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; the devices, modules, and units for implementing various functions can also be considered as both software modules implementing the method and structures within the hardware component.

[0157] The above describes specific embodiments of the present invention. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art may make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. The embodiments of this application and the features in the embodiments may be combined with each other in any manner unless there is a conflict.

Claims

1. A method for anonymizing in-vehicle emotional data, characterized in that: include: Step S1: Collect multimodal emotion data; Step S2: performing local preprocessing and multimodal alignment processing on the multimodal emotion data; Step S3: extracting the corresponding sensitive emotion feature information from the pre-processed multimodal emotion data through a recognition algorithm, and uniformly encapsulating the identified sensitive emotion feature information; Step S4: Desensitize each type of emotional modal data, and encapsulate the desensitization results and synchronize them with the labels; Step S5: extracting deep emotion features from the desensitized multimodal emotion data using a multimodal fusion model, and fusing the extracted deep emotion features to form a unified emotion representation vector; Step S6: Input the emotion representation vector into a pre-trained emotion recognition model to output the occupant's current emotional state, including the specific emotion category and the corresponding confidence level.

2. The method for anonymizing in-vehicle emotional data according to claim 1, characterized in that: The multimodal emotion data includes image modality data, voice modality data and physiological modality data; The image modality data is collected by multiple cameras in the vehicle, which can capture occupant facial images in real time and support visible light and infrared dual-channel imaging functions; The voice modality data is collected through a distributed microphone array system, including voice input signals from the driver, co-driver, and rear seats; The physiological modality data is collected by non-invasive physiological sensors integrated in the cockpit to obtain emotion-related physiological signals in real time; Each modal data is automatically timestamped at the acquisition end and synchronized based on a time window through the edge node. The synchronized data is uniformly packaged into time-segment data packets for subsequent processing steps.

3. The method for anonymizing in-vehicle emotional data according to claim 1, characterized in that: The preprocessing of image modality data includes image denoising, grayscale conversion and normalization, face alignment, expression region cropping, and timestamp synchronization; The preprocessing of speech modality data includes denoising, silence detection and segmentation, amplitude normalization, short-time frame division and time alignment marking; Preprocessing of physiological modality data includes filtering and noise reduction, outlier removal, sliding average, normalization, sampling and alignment to achieve temporal alignment of cross-modality data; The multimodal alignment process introduces a unified timestamp mechanism and performs time alignment operations on all data based on a master time source, including sliding window synchronization, interpolation / cropping strategies, and synchronous data packet encapsulation.

4. The method for anonymizing in-vehicle emotional data according to claim 1, characterized in that: Sensitive feature recognition of image modality data includes face region detection, facial key point extraction, facial feature vector extraction, and emotion-irrelevant background feature removal. The facial region position coordinates, key point index, and embedded feature vector are obtained as the basis for desensitization processing. Sensitive feature recognition of speech modal data includes voiceprint feature extraction, spectral component analysis, and speech feature labeling that is unrelated to semantic content. The output includes voiceprint vectors, spectral sensitive segments, and frequency band distribution parameters, which are used to control subsequent voiceprint perturbations or speech masking strategies. Sensitive feature recognition of physiological modality data includes resting heart rate feature analysis, GSR skin conductance pattern recognition and heart rate variability HRV time domain / frequency domain features, and respiratory rhythm and amplitude analysis.

5. The method for anonymizing in-vehicle emotional data according to claim 1, characterized in that: Desensitization of image modality data uses Gaussian blurring, mosaic occlusion, or replacing face regions with embedded vectors, including face blurring, region occlusion / replacement, embedding alternative representations, and style transfer forgery; The desensitization processing of speech modal data includes voiceprint perturbation processing, spectral feature mapping, speech pseudo-sound transformation and low-rank reconstruction de-featureization; Desensitization of physiological modality data introduces differential privacy algorithms or dimensionality degradation processing, including normalization transformation, differential privacy perturbation, feature subspace selection, and segmented resampling; The desensitization result encapsulation and label synchronization include repackaging the retained valid multimodal emotion data and synchronizing the corresponding emotion recognition labels and timestamps.

6. The method for anonymizing in-vehicle emotional data according to claim 1, characterized in that: Deep emotion feature extraction from image modality data involves extracting key expression feature embeddings, including facial expression embedding vectors, dynamic expression change indicators, and attention heat maps, through lightweight convolutional neural networks or visual Transformer models. Deep emotion feature extraction from speech modality data involves extracting emotion indicators using a temporal modeling structure, including prosodic features, spectral features, rhythm and punctuation patterns, and emotional speech embedding vectors; Deep emotion feature extraction from physiological modality data includes features related to heart rate variability, skin conductance fluctuations and emotional states.

7. The method for anonymizing in-vehicle emotional data according to claim 1, characterized in that: The emotion recognition model includes a multi-layer perceptron (MLP) structure, a multimodal Transformer structure, or an integrated classifier system, and realizes the post-decision integration of different modal recognition results through a Softmax fusion network or a SVM and random forest combination model.

8. The method for anonymizing in-vehicle emotional data according to claim 1, characterized in that: It also includes an anonymization output policy control step, which includes identity information stripping, label recoding, context desensitization synchronization and output permission control; Identity information stripping refers to outputting only the desensitized emotion labels and emotion embedding vectors without outputting the original images, voices, and physiological signals; The label recoding means that the emotion labels can be converted into numerical codes for external systems or third-party platforms; The context desensitization synchronization refers to hiding the desensitized fields in the system log or interaction record in combination with the privacy protection flag information generated by the current processing task.

9. A system for anonymizing in-vehicle emotional data, characterized in that: include: Module M1: Collect multimodal emotion data; Module M2: performing local preprocessing and multimodal alignment processing on the multimodal emotion data; Module M3: Extracts the corresponding sensitive emotion feature information from the pre-processed multimodal emotion data through a lightweight recognition algorithm, and uniformly encapsulates the sensitive emotion feature information at the identified location; Module M4: Desensitizes each type of emotional modal data, and encapsulates the desensitization results and synchronizes them with the labels; Module M5: Uses a multimodal fusion model to extract deep emotion features from images, speech, and physiological signals, and fuses the three extracted deep features to form a unified emotion representation vector. Module M6: Input the emotion representation vector into the pre-trained emotion recognition model and output the occupant's current emotional state, including the specific emotion category and corresponding confidence level.

10. The in-vehicle emotional data anonymization processing system according to claim 9, characterized in that: The multimodal emotion data includes image modality data, voice modality data and physiological modality data; The image modality data is collected by multiple cameras in the vehicle, which can capture occupant facial images in real time and support visible light and infrared dual-channel imaging functions; The voice modality data is collected through a distributed microphone array system, including voice input signals from the driver, co-driver, and rear seats; The physiological modality data is collected by non-invasive physiological sensors integrated in the cockpit to obtain emotion-related physiological signals in real time; Each modal data is automatically timestamped at the acquisition end and synchronized based on the time window through the edge node. The synchronized data is uniformly packaged into time-segment data packets for subsequent processing module calls.

Citation Information

Patent Citations

  • Vehicle owner emotion recognition and adjustment method, storage medium and vehicle-mounted system

    CN109190459A

  • Method for anonymizing vehicle data

    CN114041177A

  • Intelligent vehicle-mounted emotion interaction system design method based on voice emotion recognition technology

    CN118314930A

Cited By

  • Multi-mode driver emotion recognition method and system in real vehicle environment

    CN120950889A

  • Driver emotion monitoring and early warning method

    CN121147884A

  • Driver emotion recognition method and system based on multi-modal feature fusion

    CN121167647A

  • Psychological counseling real-time speech recognition method based on multi-modal data

    CN121337358A

  • Edge-cloud collaborative architecture-based emotion recognition privacy protection system and method

    CN121435274A