Driver distraction behavior detection model construction method based on supervised comparative learning and detection method
By improving the comparison learning loss function based on the method of supervised comparison learning, the driver's distraction behavior detection model is constructed, and the problem of difficulty in accurately identifying and opening-set detection in the existing technology is solved, and the precise identification of typical distraction behavior is achieved, driving safety and passenger trust are improved.
Patent Information
- Application Number
- CN202510314434.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-22
AI Technical Summary
The existing driver distraction detection methods are difficult to achieve precise identification and open set detection at the same time, especially the detection of unknown distraction behavior is poor.
Using a method based on supervised contrast learning, the classic contrast learning loss function is improved by introducing multi-clustering and offset parameters, and a comparison loss function with offset is constructed, and the model is trained to achieve accurate identification and open set detection of specific distraction behaviors.
It realizes accurate identification of typical distraction behaviors such as sending text messages, making phone calls and drinking water in open-set detection, improves the reliability of driving risk assessment and the rationality of human-machine collaborative driving control, and enhances the trust of drivers and passengers in driving cars together with human-machine cars.
Smart Images

Figure CN120356191A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of driver distraction behavior monitoring, and in particular relates to a method for constructing a driver distraction behavior detection model and a detection method based on supervised contrast learning. Background Art
[0002] Driver distraction is an important factor affecting driving safety, as it can significantly distract attention, reduce reaction ability, and increase the risk of accidents. According to data from the National Highway Traffic Safety Administration of the United States, 80% of traffic accidents are caused by distracted driving. Therefore, researchers have carried out a number of studies on this issue to reduce its harm. Current distraction detection methods mainly include three categories: methods based on vehicle driving characteristics, methods based on physiological signals, and methods based on vision. However, due to the insufficient accuracy of vehicle driving characteristic methods, the high cost of physiological signal detection equipment, and the accuracy and low cost of vision detection methods, the vision detection method has become the main research direction.
[0003] Vision-based detection relies on a driver distraction dataset. However, existing datasets usually only cover a few typical distraction behaviors and cannot identify behaviors that do not appear in the training set. In response to this, this solution proposes to use the method of supervised contrast learning to achieve the detection of unknown distraction behaviors, that is, open-set detection. By introducing multi-clustering and offset parameters, the classical contrast learning loss function is improved, and the multi-clustering contrast loss function with offset is proposed. The model is trained using this method, which not only realizes open-set detection but also accurately identifies specific distraction behaviors in the context of open-set detection. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for constructing a driver distraction behavior detection model and a detection method based on supervised contrast learning for the above problems.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0007] A method for constructing a driver distraction behavior detection model based on supervised contrast learning, the method comprising:
[0008] Training samples are labeled as normal driving samples and distracted driving samples;
[0009] Each distracted driving sample is respectively labeled as a specific distracted sample category P1, P2... P Q-1 , and the remaining distracted driving samples are labeled as unknown distracted sample class P Q , Q represents the number of distracted sample categories; the above process is the process of labeling samples.
[0010] Using the labeled training samples and contrast loss functions L1, L2... L QTrain a detection model,
[0011] L M = L1 + L2 +...... + L Q ;
[0012] The function L Q is used to separate normal driving samples and distracted driving samples;
[0013] The functions L1, L2... L Q-1 correspond to specific distracted sample categories P1, P2... P Q-1 respectively, and the functions L1, L2... L Q-1 are respectively used to separate the corresponding specific distracted sample categories from the remaining samples. For example, L1 corresponds to P1, and the function L1 is used to separate the P1 sample category from the remaining samples. If P1 is the behavior of drinking water, then L1 is used to separate the drinking water behavior and other samples.
[0014] In the above method for constructing a driver distraction behavior detection model based on supervised contrast learning, the detection model includes a feature encoding layer and a distraction behavior recognition layer, and the parameters of the feature encoding layer and the distraction behavior recognition layer are continuously updated during the training process.
[0015] In the above method for constructing a driver distraction behavior detection model based on supervised contrast learning, the parameters of the distraction behavior recognition layer updated during the training process include several partitioning thresholds for dividing the training samples into Q + 1 classification intervals. If Q is 6, then it means there are 7 classification intervals, five corresponding to specific distracted sample classes P1 - P5, one corresponding to the unknown distracted sample class P6, and one corresponding to the normal driving sample class P7.
[0016] In the above method for constructing a driver distraction behavior detection model based on supervised contrast learning, the distraction behavior recognition layer further includes a standard vector and a similarity calculation module, and the feature encoding layer is used to output a feature vector for the input sample; during the training process, the parameters of the feature encoding layer are continuously updated, and for the same sample, as the parameters are updated, the output feature vector will change.
[0017] The similarity calculation module is used to calculate the similarity between the feature vector output by the feature encoding layer and the standard vector;
[0018] The distraction behavior recognition layer outputs a classification prediction result according to the classification interval in which the similarity falls. During the training process, the loss function guides the model to update the parameters according to the gap between the prediction result and the true value.
[0019] For example, assuming Q is 6, the partitioning thresholds are γ1, γ2, γ3, γ4, γ5, γ6, γ7, γ8, γ9, γ 10 、γ 11, those with a similarity > γ1 fall into the seventh interval, corresponding to the normal driving sample class P7, and the output prediction result is normal driving; those with a similarity between γ2 and γ3 fall into the fifth interval, corresponding to the distracted driving sample class P5, and the output prediction result is the distracted behavior of class P5; those with a similarity between γ4 and γ5 fall into the fourth interval, corresponding to the distracted driving sample class P4, and the output prediction result is the distracted behavior of class P4;...; those with a similarity between γ 10 ~γ 11 fall into the first interval, corresponding to the distracted driving sample class P1, and the output prediction result is the distracted behavior of class P1. If the similarity is not in the first, second, third, fourth, fifth, or seventh interval, it is considered to fall into the sixth interval, corresponding to the distracted driving sample class P6, and the output prediction result is unknown distracted behavior.
[0020] Through the update of the threshold parameters for the above classification interval division, it is possible to achieve accurate prediction of specific distracted behaviors while predicting the open set of driving distracted behaviors, accurately identify specific distracted behaviors while realizing open set detection, and distinguish different risk levels of distracted behaviors. For example, texting is usually more dangerous than making a phone call, providing detection technical support for improving driving safety.
[0021] In the above method for constructing a driver distracted behavior detection model based on supervised contrast learning, the standard vector is the average vector value of the feature vectors output by the feature encoding layer for normal driving samples. As the feature encoding layer is updated, the standard vector is also continuously updated. After the training is completed, this standard vector is determined.
[0022] In the above method for constructing a driver distracted behavior detection model based on supervised contrast learning, the function L Q achieves the learning goal of separating normal driving samples and distracted driving samples by maximizing the similarity between normal driving sample pairs and minimizing the similarity between normal driving samples and distracted driving samples;
[0023] The functions L1, L1... L Q-1 respectively achieve the learning goals of aggregating the corresponding specific distracted sample categories and separating the corresponding specific distracted sample categories from the remaining samples by minimizing the sample similarity of the corresponding specific distracted sample categories p1, p2... p Q-1 and maximizing the similarity between the samples of the corresponding specific distracted sample categories P1, P2... P Q-1 and the remaining samples.
[0024] In the above method for constructing a driver distracted behavior detection model based on supervised contrast learning, the calculation method of the similarity calculation module is as follows:
[0025]
[0026] In the formula, ysim is the similarity, and f θ (x) is the encoded feature vector after L2-norm normalization of the feature vector, and V n is the standard vector;
[0027] The L2-norm normalization is as follows:
[0028]
[0029] where v i is the i-th element in the multi-dimensional vector, and the feature vector output by the feature encoding layer is provided to the similarity calculation module for similarity calculation after being normalized by the L2-norm.
[0030] In the above method for constructing a driver distraction behavior detection model based on supervised contrast learning, the specific distraction sample categories at least include three specific categories: texting P1, making a call P2, and drinking water P3;
[0031] The loss function L1 for texting P1 is:
[0032]
[0033] P(i) is the set of indices of the vectors in V t except for V ti and A(t) is the set of indices of the vectors in V except for V t , and K t represents K t texting samples;
[0034] The loss function L2 for making a call P2 is:
[0035]
[0036] P(i) is the set of indices of the vectors in V p except for V pi and A(p) is the set of indices of the vectors in V except for V p , and K p represents K p calling samples;
[0037] The loss function L3 for drinking water P3 is:
[0038]
[0039] P(i) is the set of indices of the vectors in V d except for V di and A(d) is the set of indices of the vectors in V except for V d , and K d represents Kd A drinking water sample.
[0040] In the above formula, V t is the set of samples for sending text messages;
[0041] v tj is the j-th sample for sending text messages, and is its transpose;
[0042] V p is the set of samples for making phone calls;
[0043] v pj is the j-th sample for making phone calls, and is its transpose;
[0044] V d is the set of samples for drinking water;
[0045] v dj is the j-th sample for drinking water, and is its transpose;
[0046] v m and τ are adjustable weight parameters.
[0047] In the above method for constructing a driver distraction behavior detection model based on supervised contrast learning, the contrast loss function is introduced with offset parameters:
[0048] L D = L nD + 0.5 * (L1, L2......L Q-1 ) (13)
[0049] L nD is the function L Q after introducing the offset parameters μ1, μ2……μ Q , and μ1, μ2……μ Q correspond to the distraction sample categories P1, P2……P Q respectively.
[0050] A method for detecting distraction behavior using a detection model constructed by the above method, including:
[0051] Obtain a driver image;
[0052] Preprocess the driver image and then input it into the feature encoding layer;
[0053] The feature encoding layer outputs a driver image feature vector;
[0054] The similarity calculation module calculates the similarity between the driver image feature vector and the standard vector. Both vectors can be normalized by the L2 norm;
[0055] Output the classification prediction result according to the classification interval in which the similarity falls.
[0056] The advantages of the present invention are as follows:
[0057] A driver distraction behavior detection architecture based on supervised contrast learning is proposed, which solves the problem that it is difficult to simultaneously achieve accurate recognition and open-set detection in driver distraction behavior detection; the training samples are not only labeled as normal driving samples and distracted driving samples, but further the distracted driving samples are labeled according to the categories of distraction sample categories. For normal samples and distracted samples, a loss function is used for separation. At the same time, for each determined sample category, a corresponding loss function is further corresponding. By introducing multi-clustering and offset parameters, the classical contrast learning loss function is improved, and the offset multi-clustering contrast loss function is proposed, which can realize the open-set detection of distraction behaviors, and further realize the accurate recognition of typical distraction driving behaviors such as texting, making calls, and drinking water in the case of open-set detection, providing a more reliable evaluation basis for the driving risk assessment system, and improving the rationality of the allocation of human-machine collaborative driving control permissions and the trust and acceptance of passengers and drivers in human-machine co-driving vehicles. Brief Description of the Drawings
[0058] Figure 1 It is a schematic diagram of the driver distraction behavior detection architecture of the present invention.
[0059] Figure 2 It is a schematic diagram of the supervised contrast learning training framework of the present invention.
[0060] Figure 3 It is a receiver operating characteristic curve of three loss models of the present invention.
[0061] Figure 4 It is a confusion matrix of the training model under the D-InfoNCE loss of the present invention. Detailed Description of the Invention
[0062] The present invention provides a method for constructing a driver distraction behavior detection model based on supervised contrast learning. The model constructed based on this method can simultaneously achieve open-set detection and accurate recognition of several specific distraction categories, providing a more reliable evaluation basis for the driving risk assessment system, and improving the rationality of the allocation of human-machine collaborative driving control permissions and the trust and acceptance of passengers and drivers in human-machine co-driving vehicles.
[0063] In this embodiment, three typical distraction driving behaviors, namely texting, making calls, and drinking water, are taken as specific distraction types to illustrate the solution in detail.
[0064] Such as Figure 1 and Figure 2As shown, the architecture of the detection method includes an image acquisition layer, an image preprocessing layer, a feature encoding layer, a distracted behavior recognition layer, and a result output layer. The image acquisition layer is connected to the image prediction processing layer and outputs the acquired image to the image preprocessing layer. The image preprocessing layer is connected to the feature encoding layer and outputs the preprocessed image to the feature encoding layer. The feature encoding layer is connected to the distracted behavior recognition layer and outputs a multi-dimensional encoding vector to the distracted behavior recognition layer; the distracted behavior recognition layer is connected to the result output layer and outputs the recognized distracted behavior to the result output layer. The following respectively gives an implementation of the image acquisition layer, the image preprocessing layer, the feature encoding layer, the distracted behavior recognition layer, and the result output layer. When put into use, it is not limited to the example implementation given below.
[0065] The image acquisition layer can obtain the driver image in real time through various RGB cameras, and the sampling frame rate of the camera can be 20 frames per second;
[0066] The image preprocessing layer includes image scaling and mean normalization processing. In this embodiment, the image is finally scaled to 48×48 pixels.
[0067] The feature encoding layer has two models: 3D ResNet 18 and 2D ResNet 18. In this embodiment, the DAD dataset and the State-Farm dataset are used to train and validate the feature encoder. Since the data types of the two datasets are different, their corresponding feature encoder structures are also different. For the DAD dataset, since the images in the DAD dataset are consecutive frame images, a temporal model is required, and the feature encoder uses the 3D ResNet 18 model. The input image to the 3D ResNet 18 model is the 16 preprocessed images newly obtained by the camera. The convolutional layer, pooling layer, and batch normalization in 3D ResNet 18 are all in three-dimensional form. For the State-Farm dataset, the images are single-frame images, and the feature encoder uses the 2D ResNet 18 model. The structure of the 2D ResNet 18 model is the same as that of the 3D ResNet 18 model, but the convolutional layer, pooling layer, and batch normalization in 2D ResNet 18 are all in two-dimensional form. The feature encoding layer finally outputs a 512-dimensional feature vector.
[0068] The distracted behavior recognition layer first performs L2 norm normalization on the 512-dimensional encoding vector output by the feature encoding layer, then calculates its similarity with the standard vector, and compares the obtained similarity with a preset threshold to determine the type of the image.
[0069] The result output layer includes five types of driving behaviors: normal driving, texting, making a call, drinking water, and other distracted categories (i.e., distracted categories other than texting, making a call, and drinking water).
[0070] Specifically, the L2 norm normalization in the distracted behavior recognition layer is as follows:
[0071]
[0072] In the formula, v i is the i-th element in the 512-dimensional vector.
[0073] The similarity calculation method in the distracted behavior recognition layer is as follows:
[0074]
[0075] In the formula, Sim is the similarity, f θ (x) is the 512-dimensional encoded vector after L2 norm normalization, and V n is the standard vector.
[0076] The 512-dimensional feature vector here can be replaced by a 128-dimensional feature vector obtained by dimensionality reduction of the 512-dimensional feature vector by the classifier, as Figure 2 shown.
[0077] Specifically, the standard vector is obtained in the following way:
[0078] Input the normal driving samples in the training set into the feature encoding layer. The feature encoding layer calculates the average of the feature vectors output for the normal driving samples and performs L2 norm normalization on the average vector. This process of L2 norm normalization corresponds to the similarity calculation process. If f θ (x) in the above formula (2) is the feature vector after L2 norm normalization, then L2 norm normalization is also performed on the average vector here. If f θ (x) in the above formula (2) is the feature vector without normalization or normalized by other methods, then no normalization or other normalization methods are performed on the average value here.
[0079] Since this embodiment takes three specific distracted sample categories as examples, that is, there are a total of four distracted categories, Q = 4. Therefore, there should be five classification results here, and there can be 7 division thresholds in the distracted behavior recognition layer: γ1, γ2, γ3, γ4, γ5, γ6, γ7. The relationship between similarity, threshold, and category is:
[0080]
[0081] That is, if the similarity is greater than γ1 and not within the intervals corresponding to the three categories of texting, calling, and drinking, it is recognized as an unknown distracted category.
[0082] Calculate the contrast loss function for the prediction results and update the model parameters through backpropagation to complete the training of the model.
[0083] For a training sample set X, the supervised contrast loss function is as follows:
[0084]
[0085] where I is the set of indices of samples in the data set X, x i is the i-th sample in the data set X, P(i) is the set of indices of the positive example samples of x i , |P(i)| is the number of positive example samples, and A(i) is the set of indices of samples in the data set X other than x i . In supervised contrast learning, in addition to having multiple negative example samples, a sample also has multiple positive example samples. The positive example samples of a sample are the set of samples of the same class as it, and the negative example samples are the set of samples of different classes from it.
[0086] For the contrast loss function used in the training process, assume that in a batch, there are K n normal driving samples, K a distracted driving samples, where K a distracted driving samples are composed of K t texting samples, K p calling samples, K d drinking samples, and K o samples of other distracted behavior categories, that is
[0087] K a = K t + K p + K d + K o (5)
[0088] Inputting a batch of samples into the feature encoder gives the vector H. Denote the vectors corresponding to the normal driving samples, distracted driving samples, texting samples, calling samples, drinking samples, and samples of other distracted behavior categories in H as
[0089] Denote the vectors obtained by the classifier reducing the dimension of H as V respectively as
[0090] Applying the supervised contrast loss function, that is, Equation (4), to a batch of samples, the corresponding loss function can be obtained as follows:
[0091]
[0092] where P(i) is the set of indices of the other vectors in V n except for V ni .
[0093] Subsequent experiments prove that L S loss can better distinguish normal driving behavior from distracted driving behavior, and can identify unknown distracted driving behavior categories, but cannot identify specific distracted driving behaviors. The method adopted in this solution applies a supervised contrast loss function to texting, calling, and drinking behaviors.
[0094] The loss function for texting behavior is:
[0095]
[0096] In the formula, P(i) is the set of indices of other vectors in V t except for V ti and A(t) is the set of indices of other vectors in V except for V t .
[0097] The loss function for calling behavior is:
[0098]
[0099] In the formula, P(i) is the set of indices of other vectors in V p except for V pi and A(p) is the set of indices of other vectors in V except for V p .
[0100] The loss function for drinking behavior is:
[0101]
[0102] In the formula, P(i) is the set of indices of other vectors in V d except for V di and A(d) is the set of indices of other vectors in V except for V d .
[0103] Among them, equation (6) can separate and cluster the vectors of the normal driving category from the vectors of other categories; equation (7) can separate and cluster the vectors of the texting category from the vectors of other categories; equation (8) can separate and cluster the vectors of the calling category from the vectors of other categories; equation (9) can separate and cluster the vectors of the drinking category from the vectors of other categories. Combining equations (6)-(9), the supervised contrast loss function for multi-clustering is obtained:
[0104] L M = L n + L t + L p + L d (10)
[0105] Subsequent experiments prove that L M Loss can cause the vectors corresponding to the three specific distracted driving behaviors of sending text messages, making phone calls, and drinking water to cluster, but the similarity values corresponding to the three behaviors cluster in adjacent regions, and these three behaviors still cannot be recognized. The method of this solution introduces an offset parameter μ on the basis of using multi-clustering, so that the vectors corresponding to the three distracted driving behaviors of sending text messages, making phone calls, and drinking water are separated. Introducing the offset parameter μ into Equation (6) gives:
[0106]
[0107] where μ1 = μ n , μ2 = μ t , μ3 = μ p , μ4 = μ d , μ5 = μ Q = μ o are the offset parameters corresponding to the categories of normal driving, sending text messages, making phone calls, drinking water, and other distracted behaviors respectively.
[0108] Combining Equations (7)-(11), the multi-clustering contrast loss function with offset is obtained:
[0109] L D = L nD + 0.5 * (L t + L p + L d ) (12)
[0110] L D Loss is the contrast loss function finally used in this solution.
[0111] This embodiment further provides a method for detecting distracted behaviors using the detection model trained by the above method, including:
[0112] Obtain the driver's image through the image acquisition layer;
[0113] The image preprocessing layer preprocesses the driver's image and then inputs it into the feature encoding layer;
[0114] The feature encoding layer outputs the driver's image feature vector;
[0115] The similarity calculation module calculates the similarity between the driver's image feature vector after L2 norm normalization and the standard vector;
[0116] Output the classification prediction result according to the classification interval in which the similarity falls. In this example, the prediction result is one of the five classifications in formula (3). During the training process, the loss function guides the model to update the parameters according to the difference between the prediction result and the true value, and several division thresholds will be updated accordingly. When using the trained model for prediction after the training is completed, several division thresholds are determined.
[0117] To verify the effectiveness and accuracy of the proposed driver distraction behavior detection method based on supervised contrast learning, in this embodiment, through the Python 3.7 and PyTorch 1.11 environments, for L S Loss, L M Loss and L D The models trained under the loss were compared and verified, and the DAD dataset and the State-Farm dataset were used to verify the open-set detection ability of the model and the detection effect of specific distraction behaviors respectively.
[0118] Based on the DAD dataset, the driver distraction open-set detection ability of the models under the three losses was experimentally verified. Figure 3 This is the receiver operating characteristic curve (ROC curve) of the three loss models of the present invention. The accuracy of the model under L S is 81.41%, and the AUC value is 0.8893; the accuracy of the model under L M is 81.66%, and the AUC value is 0.8826; the accuracy of the model under L D is 84.40%, and the AUC value is 0.8898. It can be seen that all three losses can achieve open-set detection of driver distraction, and among them, the model trained using L D loss has the best effect.
[0119] Figure 4 This is the confusion matrix of the model trained under the L D loss of this solution. Since there are other types of distraction behavior categories in the test set, the sum of probabilities in each row of the confusion matrix is not 1. It can be seen from the figure that the proposed L D loss can enable the model to well identify three specific distraction behavior categories: sending text messages, making phone calls, and drinking water. Since the similarities between making phone calls and drinking water are relatively close, there are still mutual misidentifications, but the misidentification rate is very low.
[0120] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art of the present invention can make various modifications or - 12 - supplements or use similar methods to replace them, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
[0121] Although terms such as image acquisition layer, image preprocessing layer, feature encoding layer, distraction behavior recognition layer, result output layer, etc. are used more frequently in this article, the possibility of using other terms is not excluded. The use of these terms is only for more convenient description and explanation of the essence of the present invention; interpreting them as any kind of additional limitation is contrary to the spirit of the present invention.
Claims
1. A method for constructing a driver distraction behavior detection model based on supervised contrastive learning, characterized in that, The method includes: Training samples are labeled as normal driving samples and distracted driving samples; Each distracted driving sample is separately labeled as a specific distracted sample category P1, P2... P according to the distraction category Q-1 , and the remaining distracted driving samples are labeled as the unknown distracted sample class P Q , where Q represents the number of distracted sample categories; Use tagged training samples and contrastive loss functions L1, L2... L Q Train the detection model; The function L Q is used to separate normal driving samples and distracted driving samples; Functions L1, L2... L Q-1 correspond to specific distraction sample categories P1, P2... P Q-1 , and functions L1, L2... L Q-1 are respectively used to separate the corresponding specific distraction sample categories from the remaining samples.
2. The method for constructing a driver distraction behavior detection model based on supervised contrastive learning according to claim 1, wherein The detection model includes a feature encoding layer and a distracted behavior recognition layer, and the parameters of the feature encoding layer and the distracted behavior recognition layer are continuously updated during the training process.
3. The method for constructing a driver distraction behavior detection model based on supervised contrastive learning according to claim 2, characterized in that, The parameters updated during the training process of the distracted behavior recognition layer include several partitioning thresholds for partitioning the training samples into Q + 1 classification intervals.
4. The method for constructing a driver distraction behavior detection model based on supervised contrastive learning according to claim 3, wherein The distracted behavior recognition layer further includes a standard vector and a similarity calculation module, and the feature encoding layer is used to output a feature vector for the input sample; The similarity calculation module is used to calculate the similarity between the feature vector output by the feature encoding layer and the standard vector; The distracted behavior recognition layer outputs a classification prediction result according to the classification interval in which the similarity falls.
5. The method for constructing a driver distraction behavior detection model based on supervised contrastive learning according to claim 4, characterized in that, The standard vector is the average vector value of the feature vectors output by the feature encoding layer for normal driving samples.
6. The method for constructing a driver distraction behavior detection model based on supervised contrastive learning according to claim 4, wherein, The function L Q achieves the learning objective of separating normal driving samples and distracted driving samples by maximizing the similarity between normal driving sample pairs and minimizing the similarity between normal driving samples and distracted driving samples; The functions L1, L1... L Q-1 respectively achieve the learning objectives of aggregating the corresponding specific distraction sample categories P1, P2... P Q-1 by minimizing the sample similarity of the corresponding specific distraction sample categories P1, P2... P Q-1 and maximizing the similarity between the samples of the corresponding specific distraction sample categories P1, P2... P and the remaining samples, so as to separate the corresponding specific distraction sample categories from the remaining samples.
7. The method for constructing a driver distraction behavior detection model based on supervised contrastive learning according to claim 4, characterized in that, The calculation method of the similarity calculation module is as follows: where f θ (x) is the encoded feature vector after L2 norm normalization of the feature vector, and V n is the standard vector; The L2 norm normalization is as follows: where, v i is the i-th element in the multi-dimensional vector, and the feature vector output by the feature encoding layer is provided to the similarity calculation module for similarity calculation after being normalized by the L2 norm as described above.
8. The method for constructing a driver distraction behavior detection model based on supervised contrastive learning according to claim 7, wherein The specific distracted sample categories at least include three specific categories: texting P1, making a call P2, and drinking water P3; The loss function L1 for texting P1 is: P(i) is the index set of vectors in V t except for V ti and A(t) is the index set of vectors in V except for V t K t denotes K t text message samples; The loss function L2 for making a call P2 is: P(i) is the index set of vectors in V p except for V pi and A(p) is the index set of vectors in V except for V p K p represents K p calling samples; The loss function L3 for drinking water P3 is: P(i) is the index set of vectors in V d other than V di and A(d) is the index set of vectors in V other than V d other than V d Let K d represent K drinking water samples; It should be noted that the original text seems a bit unclear and may need further clarification for a more accurate and complete understanding. The translation is based on the best interpretation of the provided text. In the above formula, V t is the sample set of sending text messages; v tj is the j-th sample of sending text messages, and is its transpose; V p is a sample set of making phone calls; v pj is the j-th sample of making a phone call, and is its transpose; V d is the sample set of drinking water; v dj is the j-th sample of drinking water, is its transpose; v m , where τ is an adjustable weight parameter.
9. The method for constructing a driver distraction behavior detection model based on supervised contrastive learning according to claim 1, wherein, The contrastive loss function is introduced with an offset parameter: L D = L nD + 0.5 * (L1, L2......L Q-1 ) (13) L nD is the function L after introducing offset parameters μ1, μ2... μ Q , where μ1, μ2... μ Q correspond to distraction sample categories P1, P2... P Q respectively. Q .
10. A method for detecting distracted behavior using a detection model constructed by the method described in any one of claims 1-9, including: Obtaining a driver image; Preprocessing the driver image and then inputting it into the feature encoding layer; The feature encoding layer outputs a driver image feature vector; The similarity calculation module calculates the similarity between the driver image feature vector and the standard vector; Outputting a classification prediction result according to the classification interval in which the similarity falls.