Coevolution driver cognitive load personalized quantification method

Through semi-supervised collaborative pseudo-labeling and supervised contrastive learning, combined with a domain adversarial module, the real-time monitoring problem of driver cognitive load in L2/L3 autonomous driving is solved, high-precision and fine-grained driver cognitive load quantification is achieved, and the robustness and cross-scenario applicability of the model are improved, meeting the real-time decision-making needs of the autonomous driving system.

CN120705485APending Publication Date: 2025-09-26GUANGZHOU MARITIME INST
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510941804.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

The existing real-time monitoring methods of driver cognitive load in L2/L3 autonomous driving rely on a large amount of manually labeled data, have high computational redundancy, poor noise robustness, discrete outputs, and are unable to meet the needs of accurate decision-making. In addition, they lack cross-scenario and cross-individual generalization.

Method used

Semi-supervised collaborative pseudo-labeling and supervised contrastive learning are adopted, combined with a domain adversarial module, and multimodal physiological signal fusion to achieve continuous quantization, improve label quality and generalization ability, and output fine-grained cognitive load values.

Benefits of technology

It achieves high-precision quantification of driver cognitive load with a small amount of manual labeling, improves the robustness of the model and its applicability across scenarios and individuals, provides a smooth decision-making basis, and meets the real-time needs of the autonomous driving system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705485A_ABST
    Figure CN120705485A_ABST
Patent Text Reader

Abstract

The invention discloses a coevolution driver cognitive load personalized quantification method, which relates to the technical field of traffic safety, and comprises the following steps: collecting a multi-modal physiological signal and processing the multi-modal physiological signal into structured feature data; constructing a labeled sample set D1 and an unlabeled sample set Du based on the structured feature data to perform semi-supervised collaborative pseudo labeling training, generating pseudo labels for unlabeled samples, and forming a self-labeled data set S; and constructing a training set and an enhanced training set by using S and D1, carrying out supervised contrast learning feature extraction, and carrying out model tuning in combination with an MAML algorithm. And finally, inputting any sample into the optimized model to generate continuous cognitive load value prediction. Wherein only a small amount of manual labeling is needed, the label quality can be improved, and the personalized quantification efficiency of the cognitive load can be improved; multi-modal information is fused, so that the robustness and the accuracy are improved; a continuous and fine-grained quantification mode is provided, and the accurate automatic driving decision-making requirement is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of traffic safety technology, and in particular to a co-evolutionary personalized quantification method for driver cognitive load. Background Art

[0002] In scenarios where vehicles are partially automated (such as adaptive cruise control and lane keeping) but require driver re-takeover at any time, the driver's cognitive load directly impacts driving safety. Excessive cognitive load can lead to delayed takeover, such as distraction from secondary tasks; while insufficient load can reduce situational awareness, such as decreased attention due to a monotonous environment. Both approaches can potentially lead to accidents. However, traditional approaches that rely on contact physiological sensors (such as EEG) or purely visual single-frame analysis suffer from high deployment costs, significant computational redundancy, and a lack of temporal dynamic modeling capabilities. Furthermore, the scarcity of professionally labeled data further limits the model's generalization capabilities.

[0003] Regarding the real-time monitoring of driver cognitive load in L2 / L3 autonomous driving, the paper "Cognitive Workload Estimation in Conditionally Automated Vehicles: A Weakly Supervised Contrastive Learning Approach" proposes an innovative framework that combines weakly supervised learning and contrastive learning. This framework uses contextual information of the driving scene (such as automatically inferred scene complexity) as a weak supervisory signal to drive the contrastive learning model to learn powerful feature representations from multimodal physiological / behavioral data. This approach achieves high-precision, non-intrusive, real-time cognitive load estimation while significantly reducing the reliance on finely labeled data. However, this approach still has the following drawbacks:

[0004] (1) Existing methods often rely on a large amount of manually labeled data, which is expensive and difficult to quickly promote in new driving scenarios or new driver groups. The above solutions rely heavily on weak labels. Among them, high-frequency "mental calculation vs. no task" labels are easy to obtain, but there is a lack of such explicit triggers on real roads. Human-designed tasks or manual labeling are still required for migration.

[0005] (2) Existing methods often use only a single physiological signal, which is easily affected by noise and individual differences and difficult to accurately distinguish different load levels. The above scheme only uses three channels: ECG, EDA, and RESP, and does not introduce signals such as eye movements. Its robustness is limited when encountering noise such as light or electrode detachment.

[0006] (3) Existing methods only perform three-level discrete classification: low / medium / high, and cannot provide fine-grained and continuous load output, making it difficult to meet the requirements of autonomous driving systems for accurate decision-making;

[0007] (4) The output of the above solution is only binary. It can only judge "high / low" and cannot provide a continuous score from 0 to 100, making it difficult to drive progressive HMI or refined takeover logic;

[0008] (5) Existing models lack generalizability across scenarios (e.g., city / highway) and individuals (different drivers), and are prone to performance degradation when deployed online. Summary of the Invention

[0009] In response to the problems existing in the existing technology, the present invention provides a co-evolutionary personalized quantification method for driver cognitive load. By integrating semi-supervised collaborative pseudo-labeling and supervised comparative learning, only a small amount of manual labeling is required and the label quality can be improved. The domain adversarial module is introduced to enhance the generalization ability across scenarios and individuals. A continuous and fine-grained quantification method is provided to meet the needs of accurate decision-making. Multimodal information fusion improves robustness and accuracy.

[0010] The technical solution of the present invention is achieved as follows:

[0011] A co-evolutionary personalized quantification method for driver cognitive load includes the following steps:

[0012] Step 1: Data preprocessing: Collecting and preprocessing multimodal physiological signals from the driver to generate a plurality of structured feature data; collecting corresponding cognitive load values ​​for some of the structured feature data, where the collected cognitive load values ​​are 1 or 100;

[0013] Specifically, multimodal physiological signals include physiological indicators such as the driver's eye movements, head movements, heart rate (HR), and galvanic skin response (GSR).

[0014] Preprocessing includes bandpass / bandstop filtering of the original signal, ICA artifact removal (EEG), peak detection / heart rate variability analysis (ECG / PPG), low-pass filtering and SCR separation (GSR), and pupil diameter normalization.

[0015] Specifically, the load labels are collected by using the NASA-TLX score form actively filled out by the driver.

[0016] The maximum and minimum values ​​of cognitive load are represented by multi-view features. Specifically, based on the relationship between each cognitive feature indicator and cognitive load, the lowest and highest cognitive workloads are labeled as 1 and 100, respectively. If the feature corresponds to the maximum cognitive load, the sample is labeled as 100. Conversely, if it corresponds to the minimum cognitive load, it is labeled as 1. This range of cognitive load is determined based on the NASA-TLX study.

[0017] The other unlabeled samples are obtained through collaborative training and contrastive learning in the following ensemble learning.Compared with dividing cognitive workload into several discrete levels, the cognitive load labeling method provides significant information gain.

[0018] Step 2: Semi-supervised collaborative pseudo-labeling: Construct a labeled sample set D based on the structured feature data l and the unlabeled sample set D u ;D l and D u The samples in D all contain the structured feature data; l All samples in contain the stated cognitive load values;

[0019] Based on D l Train two gradient decision tree regressors R1 and R2; use R1 and R2 to u Perform collaborative pseudo-labeling iterative training to generate a self-labeled dataset S; R1 and R2 predict and generate a predicted load value and confidence score for the input sample; the predicted load value is used as a pseudo-label and has a value range of [1,100];

[0020] S includes several samples, and each sample includes the structured feature data, the pseudo label and the confidence score;

[0021] Step 3: Supervised contrastive learning feature extraction: S and D l Merge and construct an enhanced data set; input the enhanced data set into a preset encoder for feature extraction, and obtain a spatial vector z through projection head mapping; one sample corresponds to one spatial vector z; and construct a loss function based on the spatial vector z;

[0022] The projection head is a neural network that projects the hidden layer features of the model into another vector space.

[0023] Step 4: Model tuning: Repeat steps 1 to 3 and use the MAML algorithm for optimization iterations, namely:

[0024] Where i represents task i, i.e., the optimization iteration of round i; θ is the initial parameter of the model; θ′ is the parameter after internal update of task i; θ * is the parameter after external update of task i; α and β are the preset internal update learning rate and external update learning rate respectively; L i (·) is the loss function for task i; f (·) Indicates the model corresponding to the selected parameters;

[0025] MAML stands for Model-Agnostic Meta-Learning, which means model-independent meta-learning. With a small number of samples and iterations, the model can quickly adapt to new tasks.

[0026] Step 5, quantized output: For any sample, input R1 and R2 respectively to obtain the two predicted load values; perform weighted calculation on the two predicted load values ​​and output the final continuous load value; the value range of the continuous load value is [1,100].

[0027] The continuous load value is the final predicted output of the cognitive load value for a sample. "Continuous" means taking a value between "1, 100." Existing load value quantization uses levels to divide into multiple categories, that is, using a discrete number of numerical representations. However, the "continuous" quantization in this solution uses a numerical representation of 1-100, which is more refined. This continuous quantization method provides a smoother basis for intervention / warning decisions for the autonomous driving system, avoiding the rigid switching caused by discrete classification, and improving the human-machine collaborative experience.

[0028] This solution combines the two stages of "semi-supervised collaborative training" and "supervised contrastive loss" and incorporates personalized fine-tuning, achieving a complete closed-loop process from a small amount of manual labeling → automatically expanding large-scale high-quality pseudo-labels → enhancing feature separability through contrastive learning → personalized model calibration → ultimately outputting personalized driver load quantification values. This strategy not only improves label quality in the pseudo-labeling stage, but also enhances feature separability in the contrastive learning stage, and achieves precise adaptation through personalized fine-tuning. The overall performance significantly exceeds that of single-stage or purely supervised or unsupervised methods.

[0029] As a further optimization of the above scheme, the method further includes step 6, verification and evaluation: inputting a new sample; the new sample includes a cognitive load value, and the cognitive load value has a value range of [1,100]; inputting the new sample into the encoder, and finally generating the continuous load value;

[0030] The cognitive load value and the continuation load value are compared to evaluate the prediction accuracy of the model.

[0031] The two load values—the former is a collected subjective value, the latter a predicted value—can be compared to accurately assess the load quantification accuracy of the personalized model for structured features. This not only directly reveals the model's prediction bias but also guides subsequent optimization efforts, such as adjusting feature weights or calibrating algorithm parameters for specific driving scenarios, thereby improving the generalization and reliability of personalized cognitive load modeling.

[0032] As a further optimization of the above scheme, the collaborative pseudo-labeling iterative training is:

[0033] A1. Construct an empty self-labeled dataset S;

[0034] A2, D u The samples in are input into R1 and R2 for prediction; after prediction, D u Each sample in obtains the predicted load value and the mean square error improvement; the mean square error improvement is used as D u Confidence score of the sample in ;

[0035] R1 and R2 sort the samples based on the confidence scores, and select several samples with the highest confidence scores and add them to S; at the same time, the predicted load values ​​of the samples selected by R1 are added as pseudo labels to the training parameters of R2; the predicted load values ​​of the samples selected by R2 are added as pseudo labels to the training parameters of R1;

[0036] A3. Get D l and S, and use the collection to update and adjust R1 and R2;

[0037] A4. Repeat steps A2 and A3 until the preset iteration end condition is reached.

[0038] The predicted load value is the prediction of the continuous cognitive load value of a sample. Specifically, R1 uses XGBoost and R2 uses LightGBM.

[0039] XGBoost is the abbreviation of "Extreme Gradient Boosting". The XGBoost algorithm is a type of synthetic algorithm that combines basis functions and weights to form a good data fitting effect.

[0040] LightGBM, the full name of which is Light Gradient Boosting Machine, is a gradient boosting framework based on decision trees, provided by Microsoft.

[0041] The two lightweight GBDT models used above can generate pseudo labels in offline batch processing, which is convenient and fast to be deployed on the vehicle edge computing unit, achieving a single-sample inference time as low as 0.001s, meeting the needs of real-time load monitoring.

[0042] As a further optimization of the above scheme, in step 3, S and D l Merge to obtain a training set; perform random data augmentation on each sample in the training set to generate an enhanced data set;

[0043] The random data enhancement operation includes time series signal noise addition, channel masking and time scaling.

[0044] As a further optimization of the above solution, the structured feature data includes pupil diameter, fixation duration, heart rate, heart rate index, dispersion, maximum GSR, blink duration, and minimum RR interval; the structured feature data is formed into a 1×8 dimensional vector;

[0045] The encoder adopts a variational autoencoder (VAE); the projection head consists of two fully connected layers, namely [FC→ReLU→FC], where FC is the fully connected layer and ReLU is a rectified linear unit.

[0046] The samples in the training set and the enhanced data set are input into VAE to generate 1×128-dimensional feature vectors, where one feature vector corresponds to one sample; wherein the feature vector corresponding to the enhanced data set is input into the projection head, mapped to generate spatial vectors, where one spatial vector corresponds to one feature vector.

[0047] Variational Auto-Encoders (VAEs) can effectively offset the noise and artifacts of a single physiological signal through a multimodal information cross-validation mechanism, making cognitive load assessment more stable and accurate. This has significant advantages, especially in complex scenarios such as nighttime driving and traffic congestion.

[0048] As a further optimization of the above solution, the loss function includes a first loss calculation, which is expressed as L1, namely:

[0049] Wherein, N represents the total number of feature vectors corresponding to the training set in the same task, z i 、z j represents the i-th and j-th eigenvectors; represents the i-th of all the space vectors in the same task;

[0050] τ is a preset temperature parameter used to control the penalty intensity for negative samples that are difficult to classify; exp(·) represents exponential calculation; sim(·) represents similarity calculation.

[0051] Contrastive loss is designed for contrastive prediction tasks and aims to maximize the consistency between pairs of positive examples. This loss function reduces the distance between the representations of similar examples (positive pairs) while increasing the distance between dissimilar examples (negative pairs). This results in closer clustering of similar instances and more distinct separation of instances in the latent space, thereby improving the discriminability and generalization of the representations.

[0052] Furthermore, the similarity is calculated as: ‖z i‖ means taking vector z i The norm of .

[0053] As a further optimization of the above solution, the encoder is further connected to a domain discrimination network; the domain discrimination network receives the feature vector and outputs a prediction of the driving scene; the driving scene includes an environment category and a driver category.

[0054] As a further optimization of the above solution, in step 1, contextual information is also collected for some of the structured feature data. The contextual information includes scene information and driver information. Scene information includes information related to driving scenarios and environments, such as nighttime, rainy days, interference, tunnel entrances and exits, and traffic congestion. Driver information includes information related to the driver's driving style, such as aggressive, moderate, and normal.

[0055] In the step 2, based on D l Training domain discriminator R3;

[0056] During the collaborative pseudo-labeling iterative training, D u The samples in are input into R3 for prediction; after prediction, D u Each sample in obtains predicted scene information and predicted driver information; each sample in S includes the predicted scene information and the predicted driver information;

[0057] The loss function includes a second loss calculation, which is denoted as L2, namely:

[0058]

[0059] Wherein, x represents a sample; h(x) represents a deep regression network, which is used to convert the sample into the feature vector; ∑ (x) (·) is the cumulative calculation based on x; a is the domain label, D(·) (·) is a domain discriminator, used to output probability distribution, ID d and ID s Respectively represent the predicted scene information and the predicted driver information associated with the sample; concat(·) represents character splicing;

[0060] The sum of the results of the first loss calculation and the second loss calculation is the calculation result of the loss function.

[0061] The domain discriminator is a multi-classification network whose task is to determine which domain the sample belongs to. The domain is defined by the combination of "driver ID + scene ID". Its role is to promote domain invariance across individuals and scenes, thereby achieving the purpose of personalized quantification of cognitive load. d and ID sExpressed in the form of integer identifiers, the combination is mapped into an integer number index through concat (function), and an integer number index corresponds to a probability distribution.

[0062] The network structure of the domain discriminator R3 is [Linear→BatchNorm→ReLU→Dropout→

[0063] Linear→ReLU→Linear→Softmax].

[0064] Linear achieves linear transformation of input features through weight matrix and bias vector; BatchNorm represents batch normalization layer, which standardizes the input data (subtracts mean, divides standard deviation) and introduces learnable scaling parameters and offset parameters to stabilize the training process; ReLU represents rectified linear unit activation function, which enhances the model expression ability through nonlinear mapping; Dropout represents random inactivation layer, which temporarily discards neuron output with a set probability to prevent overfitting; after the first Linear layer receives the input parameters, the data is sequentially normalized by BatchNorm → ReLU activation → Dropout regularization → the second Linear layer performs another linear transformation on the features → ReLU secondary activation → the final Linear layer outputs the score (logit) corresponding to each category, and finally Softmax converts the original score into a probability distribution output.

[0065] By introducing domain discrimination networks and adversarial training, the encoder can remove the interference of individual driver differences and scene differences while learning to distinguish different workloads, achieving stronger applicability across drivers and road scenarios, that is, improving the generalization ability across scenarios and individuals.

[0066] Compared with the prior art, the present invention achieves the following beneficial effects:

[0067] (1) This invention combines the two stages of "semi-supervised collaborative training" and "supervised contrastive loss" and incorporates personalized fine-tuning, achieving a complete closed-loop process from a small amount of manual labeling, automatic expansion of large-scale high-quality pseudo-labels, contrastive learning to enhance feature separability, personalized model calibration, and finally output of the driver's personalized load quantization value. This strategy not only improves label quality in the pseudo-labeling stage, but also enhances feature separability in the contrastive learning stage and achieves precise adaptation through personalized fine-tuning, significantly surpassing the performance of single-stage or purely supervised / purely unsupervised methods.

[0068] (2) Existing load quantification uses levels to divide load values ​​into multiple categories, that is, using discrete numerical representations. However, the "continuous" quantification in this solution uses numerical representations from 1 to 100, which is a more refined classification. The continuous quantification method can provide a smoother basis for intervention / warning decision-making for the autonomous driving system, avoid the rigid switching caused by discrete classification, and improve the human-machine collaborative experience.

[0069] (3) The two lightweight GBDT models used can generate pseudo labels in offline batch processing, which is convenient and fast to be deployed on the vehicle edge computing unit, achieving a single sample inference time as low as 0.001s, meeting the real-time load monitoring needs.

[0070] (4) By introducing domain discrimination networks and adversarial training, the encoder can remove the interference of individual driver differences and scene differences while learning to distinguish different workloads, achieving stronger cross-driver and cross-road scene applicability, that is, improving the generalization ability across scenes and individuals.

[0071] (5) The multimodal information cross-validation mechanism can effectively offset the noise and artifact interference of a single physiological signal, making the cognitive load assessment more stable and accurate, especially in complex scenarios such as night driving and traffic congestion. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 This is a flow chart of a co-evolutionary personalized quantification method for driver cognitive load provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0073] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0074] like Figure 1 As shown, this embodiment provides a co-evolutionary personalized quantification method for driver cognitive load, including the following steps:

[0075] Step 1: Data preprocessing: Collect and preprocess multimodal physiological signals from the driver to generate several structured feature data.

[0076] Specifically, the structured feature data includes pupil diameter, fixation duration, heart rate, heart rate index, dispersion, maximum GSR, blink duration, and minimum RR interval; the structured feature data is formed into a 1×8 dimensional vector. Preprocessing includes bandpass / bandstop filtering of the raw signal, artifact removal by ICA (EEG), peak detection / heart rate variability analysis (ECG / PPG), low-pass filtering and SCR separation (GSR), and pupil diameter normalization.

[0077] For some structured feature data, corresponding cognitive load values ​​and contextual information are collected, with the collected cognitive load values ​​being either 1 or 100. Contextual information includes scene information and driver information. Scene information includes information related to the driving scene and environment, such as nighttime, rainy days, interference, tunnel entrances and exits, and traffic congestion. Driver information includes information related to the driver's driving style, such as aggressive, moderate, and normal.

[0078] Step 2: Semi-supervised collaborative pseudo-labeling: Construct a labeled sample set D based on structured feature data l and the unlabeled sample set D u ;D l and D u The samples in D all contain structured feature data; l All samples in the dataset include cognitive load values.

[0079] Based on D l Train two gradient decision tree regressors, R1 and R2, and a domain discriminator, R3. Specifically, use XGBoost for R1 and LightGBM for R2. XGBoost, short for "Extreme Gradient Boosting," is a synthetic algorithm that combines basis functions and weights to achieve excellent data fit.

[0080] LightGBM, the full name of which is Light Gradient Boosting Machine, is a gradient boosting framework based on decision trees, provided by Microsoft.

[0081] The two lightweight GBDT models used above can generate pseudo labels in offline batch processing, which is convenient and fast to be deployed on the vehicle edge computing unit, achieving a single-sample inference time as low as 0.001s, meeting the needs of real-time load monitoring.

[0082] The network structure of the domain discriminator R3 is [Linear→BatchNorm→ReLU→Dropout→Linear→ReLU→Linear→Softmax].

[0083] Use R1 and R2 to Du Perform collaborative pseudo-annotation iterative training to generate a self-annotated dataset S; specifically:

[0084] A1. Construct an empty self-labeled dataset S;

[0085] A2, D u The samples in are input into R1 and R2 for prediction; after prediction, D u Each sample in obtains the predicted load value and the mean square error improvement; the mean square error improvement is used as D u The confidence score of the sample in the prediction load value is in the range of [1,100]. The prediction load value is the prediction of the continuous cognitive load value of a sample. In addition, R3 outputs the predicted scene information and predicted driver information about the sample.

[0086] The predicted sample and its predicted value constitute a new sample to be added to S. R1 and R2 sort the samples based on the confidence score, and select several samples with the highest confidence score to add to S; at the same time, the predicted load value of the sample selected by R1 is added as a pseudo label to the training parameters of R2; the predicted load value of the sample selected by R2 is added as a pseudo label to the training parameters of R1;

[0087] A3. Get D l and S, and use the collection to update and adjust R1 and R2;

[0088] A4. Repeat steps A2 and A3 until the preset iteration end condition is reached.

[0089] S includes several samples, and each sample includes structured feature data, pseudo labels and confidence scores, as well as predicted scene information and predicted driver information.

[0090] Step 3: Supervised Contrastive Learning Feature Extraction:

[0091] S and D l Merge to obtain a training set; perform random data augmentation operations on each sample in the training set to generate an enhanced data set; random data augmentation operations include time series signal noise addition, channel occlusion, and time scaling.

[0092] The augmented dataset is fed into the preset encoder for feature extraction. The spatial vector z is obtained by projection head mapping.

[0093] Specifically, the encoder uses a variational autoencoder (VAE). The projection head is a neural network layer that projects the hidden layer features of the model into another vector space. The projection head consists of two fully connected layers, namely [FC→ReLU→FC], where FC is a fully connected layer and ReLU is a rectified linear unit.

[0094] The samples in the training set and the enhanced data set are input into the VAE to generate a 1×128-dimensional feature vector, where one feature vector corresponds to one sample; among them, the feature vector corresponding to the enhanced data set is input into the projection head, mapped to generate a spatial vector z, where one spatial vector corresponds to one feature vector.

[0095] Construct a loss function based on the space vector z. In this embodiment, the loss function includes a first loss calculation and a second loss calculation. The first loss calculation is represented by L1, that is:

[0096] Among them, N represents the total number of feature vectors corresponding to the training set in the same task, z i 、z j represents the i-th and j-th eigenvectors; Represents the i-th of all space vectors in the same task;

[0097] τ is the preset temperature parameter (used to control the penalty intensity for negative samples that are difficult to classify); exp(·) represents exponential calculation; sim(·) represents similarity calculation.

[0098] The second loss calculation is expressed as L2, that is:

[0099] Where x represents a sample; h(x) represents a deep regression network, which is used to convert the sample into a feature vector; ∑ (x) (·) is the cumulative calculation based on x; a is the domain label, D(·) (·) is a domain discriminator, used to output probability distribution, ID d and ID s Respectively represent the predicted scene information and predicted driver information associated with the sample; concat(·) represents character concatenation;

[0100] The sum of the results of the first loss calculation and the second loss calculation is the calculation result of the loss function.

[0101] Contrastive loss is designed for contrastive prediction tasks and aims to maximize the consistency between pairs of positive examples. This loss function reduces the distance between the representations of similar examples (positive pairs) while increasing the distance between dissimilar examples (negative pairs). This results in closer clustering of similar instances and more distinct separation of instances in the latent space, thereby improving the discriminability and generalization of the representations.

[0102] The domain discrimination network receives the feature vector and outputs a prediction of the driving scene; the driving scene includes the environment category and the driver category.

[0103] Environmental categories include nighttime, rainy days, interference, tunnel entrances and exits, and traffic congestion. Driver categories include personalized driver classification, such as identifying drivers based on specific characteristics such as age, occupation, or non-occupation. For example, drivers can be categorized as aggressive, moderate, or average based on their driving style.

[0104] By introducing domain discrimination networks and adversarial training, the encoder can remove the interference of individual driver differences and scene differences while learning to distinguish different workloads, achieving stronger applicability across drivers and road scenarios, that is, improving the generalization ability across scenarios and individuals.

[0105] Step 4: Model tuning: Repeat steps 1 to 3 and use the MAML algorithm for optimization iterations, namely:

[0106] Where i represents task i, i.e., the optimization iteration of round i; θ is the initial parameter of the model; θ′ is the parameter after internal update of task i; θ * is the parameter after external update of task i; α and β are the preset internal update learning rate and external update learning rate respectively; L i (·) is the loss function for task i; f (·) Indicates the model corresponding to the selection parameters.

[0107] MAML stands for Model-Agnostic Meta-Learning, which means model-independent meta-learning. With a small number of samples and iterations, the model can quickly adapt to new tasks.

[0108] Step 5. Quantified output: For any sample, input R1 and R2 respectively to obtain two predicted load values; perform weighted calculation on the two predicted load values ​​with weights of 0.5 and 0.5 respectively, and output the final continuous load value.

[0109] Step 6: Verification and evaluation: Input a new sample; the new sample contains a cognitive load value, and the cognitive load value range is [1,100]. Input the new sample into the encoder and finally generate a continuous load value.

[0110] Compare the cognitive load values ​​and the continuation load values ​​to evaluate the prediction accuracy of the model.

[0111] The two load values—the former is a collected subjective value, the latter a predicted value—can be compared to accurately assess the load quantification accuracy of the personalized model for structured features. This not only directly reveals the model's prediction bias but also guides subsequent optimization efforts, such as adjusting feature weights or calibrating algorithm parameters for specific driving scenarios, thereby improving the generalization and reliability of personalized cognitive load modeling.

[0112] Based on the disclosure and teachings of the above description, those skilled in the art may also make changes and modifications to the above embodiments. Therefore, the present invention is not limited to the specific embodiments disclosed and described above, and modifications and variations of the present invention should also fall within the scope of protection of the claims of the present invention. In addition, although certain specific terms are used in this description, these terms are only for convenience of description and do not constitute any limitation to the present invention.

Claims

1. A co-evolutionary personalized quantification method for driver cognitive load, characterized by: The steps include: Step 1: Data preprocessing: Collecting and preprocessing multimodal physiological signals from the driver to generate a plurality of structured feature data; collecting corresponding cognitive load values ​​for some of the structured feature data, where the collected cognitive load values ​​are 1 or 100; Step 2: Semi-supervised collaborative pseudo-labeling: Construct a labeled sample set D based on the structured feature data l and the unlabeled sample set D u ;D l and D u The samples in all contain the structured feature data; D l All samples in contain the stated cognitive load values; Based on D l Train two gradient decision tree regressors R1 and R2; use R1 and R2 to u Perform collaborative pseudo-labeling iterative training to generate a self-labeled dataset S; wherein R1 and R2 predict and generate a predicted load value and a confidence score for the input sample; the predicted load value is used as a pseudo-label and has a value range of [1,100]; S includes several samples, and each sample includes the structured feature data, the pseudo label and the confidence score; Step 3: Supervised contrastive learning feature extraction: S and D l Merge and construct an enhanced data set; input the enhanced data set into a preset encoder for feature extraction, and obtain a spatial vector z through projection head mapping; one sample corresponds to one spatial vector z; and construct a loss function based on the spatial vector z; Step 4: Model tuning: Repeat steps 1 to 3 and use the MAML algorithm for optimization iterations, namely: Where i represents task i, i.e., the optimization iteration of round i; θ is the initial parameter of the model; θ′ is the parameter after internal update of task i; θ * is the parameter after external update of task i; α and β are the preset internal update learning rate and external update learning rate respectively; L i (·) is the loss function for task i; f (·) Indicates the model corresponding to the selected parameters; Step 5, quantized output: For any sample, input R1 and R2 respectively to obtain the two predicted load values; perform weighted calculation on the two predicted load values ​​and output the final continuous load value; the value range of the continuous load value is [1,100].

2. The co-evolutionary personalized quantification method for driver cognitive load according to claim 1, characterized in that: The method further includes step 6, verification evaluation: inputting a new sample; the new sample includes a cognitive load value, and the cognitive load value has a value range of [1,100]; inputting the new sample into the encoder, and finally generating the continuous load value; The cognitive load value and the continuation load value are compared to evaluate the prediction accuracy of the model.

3. The co-evolutionary personalized quantification method for driver cognitive load according to claim 1, characterized in that: The collaborative pseudo-labeling iterative training is: A1. Construct an empty self-labeled dataset S; A2, D u The samples in are input into R1 and R2 for prediction; after prediction, D u Each sample in obtains the predicted load value and the mean square error improvement; the mean square error improvement is used as D u Confidence score of the sample in ; R1 and R2 sort the samples based on the confidence scores, and select several samples with the highest confidence scores and add them to S; at the same time, the predicted load values ​​of the samples selected by R1 are added as pseudo labels to the training parameters of R2; the predicted load values ​​of the samples selected by R2 are added as pseudo labels to the training parameters of R1; A3. Get D l and S, and use the collection to update and adjust R1 and R2; A4. Repeat steps A2 and A3 until the preset iteration end condition is reached.

4. The co-evolutionary personalized quantification method for driver cognitive load according to claim 1, characterized in that: In the step 3, S and D l Merge to obtain a training set; perform random data augmentation on each sample in the training set to generate an enhanced data set; The random data enhancement operation includes time series signal noise addition, channel masking and time scaling.

5. The co-evolutionary personalized quantification method for driver cognitive load according to claim 4, characterized in that: The structured feature data includes pupil diameter, fixation duration, heart rate, heart rate index, dispersion, maximum GSR, blink duration, and minimum RR interval; the structured feature data is formed as a 1×8 dimensional vector; The encoder adopts a variational autoencoder (VAE); the projection head consists of two fully connected layers, namely [FC→ReLU→FC]; The samples in the training set and the enhanced data set are input into VAE to generate 1×128-dimensional feature vectors, where one feature vector corresponds to one sample; wherein the feature vector corresponding to the enhanced data set is input into the projection head, mapped to generate spatial vectors, where one spatial vector corresponds to one feature vector.

6. The co-evolutionary personalized quantification method for driver cognitive load according to claim 5, characterized in that: The loss function includes a first loss calculation, which is denoted as L1, namely: Wherein, N represents the total number of feature vectors corresponding to the training set in the same task, z i 、z j represents the i-th and j-th eigenvectors; represents the i-th of all the space vectors in the same task; τ is the preset temperature parameter; exp(·) represents exponential calculation; sim(·) represents similarity calculation.

7. The co-evolutionary personalized quantification method for driver cognitive load according to claim 6, characterized in that: In the step 1, context information is also collected for part of the structured feature data; the context information includes scene information and driver information; In the step 2, based on D l Training domain discriminator R3; During the collaborative pseudo-labeling iterative training, D u The samples in are input into R3 for prediction; after prediction, D u Each sample in obtains predicted scene information and predicted driver information; each sample in S includes the predicted scene information and the predicted driver information; The loss function includes a second loss calculation, which is denoted as L2, namely: Wherein, x represents a sample; h(x) represents a deep regression network, which is used to convert the sample into the feature vector; ∑ (x) (·) is the cumulative calculation based on x; a is the domain label, D(·) (·) is a domain discriminator, used to output probability distribution, ID d and ID s Respectively represent the predicted scene information and the predicted driver information associated with the sample; concat(·) represents character splicing; The sum of the results of the first loss calculation and the second loss calculation is the calculation result of the loss function.

8. The co-evolutionary personalized quantification method for driver cognitive load according to claim 7, characterized in that: The network structure of the domain discriminator R3 is [Linear→BatchNorm→ReLU→Dropout→Linear→ReLU→Linear→Softmax].

Citation Information

Cited By

  • Driver risk cognition state identification method based on heterogeneous modal knowledge distillation

    CN122220885A

  • A method for identifying driver risk perception state based on heterogeneous modal knowledge distillation

    CN122220885B