Driving behavior analysis method and related device based on facial emotion recognition

By acquiring and processing driver's facial videos in real time, combining the distribution locations of core facial images and facial image sequences, iteratively improves the emotional recognition results, solving the problem of low emotional recognition accuracy in traditional methods, and achieving accurate recognition of driver's tiny emotional changes and high accuracy in driving behavior safety assessment.

CN119741683BActive Publication Date: 2025-05-09GUIZHOU UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510237761.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-05-09
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

The prior art is difficult to accurately capture the slight emotional changes of drivers during long-term driving, and traditional driving behavior analysis methods ignore the driver's emotional factors, resulting in low emotional recognition accuracy and unable to provide a reliable basis for driving behavior safety assessment.

Method used

By obtaining real-time driver facial videos, segmenting and feature representations, initial emotion recognition results are generated, and the emotional recognition results are iteratively improved by detecting the distribution positions of core facial images and facial image sequences, and finally evaluating driving behavior safety based on the target emotion recognition results.

Benefits of technology

It improves the accuracy of the boundary of the facial image sequence, enhances the accuracy of emotion recognition, can more accurately identify the driver's tiny emotional changes, and improves the accuracy of driving behavior safety assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119741683B_ABST
    Figure CN119741683B_ABST
Patent Text Reader

Abstract

The present invention provides a driving behavior analysis method based on facial emotion recognition and a related device. The real-time driver's facial video is segmented, and the obtained target facial video clip is characterized to obtain a video feature array of the real-time driver's facial video, and based on the video feature array, an initial emotion recognition result corresponding to the real-time driver's facial video is generated, one or more core facial images are detected in the real-time driver's facial video, and based on the core facial images, one or more facial image sequences are detected in the real-time driver's facial video. The video feature array is segmented based on the distribution position to obtain a target video feature array of each facial image sequence, and the initial emotion recognition result is iterated based on the target video feature array to obtain a target emotion recognition result corresponding to the real-time driver's facial video. The present invention can improve the accuracy of driver facial emotion recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of data processing and machine learning, and in particular to a driving behavior analysis method based on facial emotion recognition and a related device. Background Art

[0002] With the rapid development of the automobile industry, driving safety issues have received increasing attention. The emotional state of the driver during driving has a crucial impact on the safety of driving behavior. Bad emotions, such as anger, fatigue, anxiety, etc., may cause the driver to react slowly, lose concentration, make mistakes in judgment, etc., thereby increasing the risk of traffic accidents. At present, traditional driving behavior analysis methods mainly focus on monitoring and analyzing vehicle driving data, such as speed, acceleration, steering angle, etc. Although these data can reflect the safety of driving behavior to a certain extent, they ignore the driver's emotional factors. In addition, some existing emotion recognition technologies are mostly based on static images or short-term video clips, which makes it difficult to accurately capture the driver's subtle emotional changes during long-term driving. Moreover, when processing video data, there are deficiencies in the detection of emotional fluctuation nodes in the video and the accurate division of facial image sequences, resulting in low accuracy of emotion recognition and failure to provide a reliable basis for driving behavior safety assessment. Summary of the invention

[0003] The purpose of the present invention is to provide a driving behavior analysis method based on facial emotion recognition and a related device. The present invention is implemented as follows:

[0004] In a first aspect, the present invention provides a driving behavior analysis method based on facial emotion recognition, comprising: acquiring a real-time driver facial video, and segmenting the real-time driver facial video to obtain a plurality of target facial video segments; performing feature representation on the target facial video segments to obtain a video feature array of the real-time driver facial video, and generating an initial emotion recognition result corresponding to the real-time driver facial video based on the video feature array; detecting one or more core facial images in the real-time driver facial video, and detecting the distribution positions of one or more facial image sequences in the real-time driver facial video based on the core facial images; segmenting the video feature array based on the distribution positions to obtain a target video feature array for each facial image sequence; iterating the initial emotion recognition result based on the target video feature array to obtain a target emotion recognition result corresponding to the real-time driver facial video; and evaluating the safety of the driver's driving behavior based on the target emotion recognition result.

[0005] On the other hand, the present invention provides a behavior analysis device, comprising: one or more processors; a memory; one or more computer programs; wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more computer programs are executed by the processors, the method described above is implemented.

[0006] The beneficial effects of the present invention include at least: after acquiring a real-time driver facial video and segmenting the real-time driver facial video to obtain a plurality of target facial video segments, the present invention performs feature representation on the target facial video segments to obtain a video feature array of the real-time driver facial video, and generates an initial emotion recognition result corresponding to the real-time driver facial video based on the video feature array, detects one or more core facial images in the real-time driver facial video, and detects the distribution positions of one or more facial image sequences in the real-time driver facial video based on the core facial images, then segments the video feature array based on the distribution positions to obtain a target video feature array for each facial image sequence, and iterates the initial emotion recognition result based on the target video feature array to obtain a target emotion recognition result corresponding to the real-time driver facial video. Because when the above process recognizes the image stream, it can not only detect the initial emotion recognition results, but also identify one or more core facial images in the real-time driver facial video, and then detect the distribution position of the facial image sequence in the real-time driver facial video based on the core facial image, thereby improving the boundary accuracy of the facial image sequence, so as to improve the emotion recognition accuracy of the image stream corresponding to the facial image sequence, help to more accurately identify the driver's subtle emotional changes during long-term driving, and improve the accuracy of safe driving assessment. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for describing the embodiments of the present invention are briefly introduced below.

[0008] Figure 1 This is a flowchart of a driving behavior analysis method based on facial emotion recognition provided by an embodiment of the present invention.

[0009] Figure 2 It is a schematic diagram of the composition of a behavior analysis device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0010] The following describes the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. The terms used in the implementation method part of the embodiments of the present invention are only used to explain the specific embodiments of the present invention, and are not intended to limit the present invention.

[0011] In the embodiment of the present invention, the executor of the driving behavior analysis method based on facial emotion recognition is a behavior analysis device, for example, a vehicle-mounted system, such as a device running in an electronic control unit (ECU). The electronic control unit can be connected to the vehicle-mounted camera module via a CAN bus (Controller Area Network), and the camera module can capture images of the driver under authorization and permitted by laws and regulations.

[0012] The embodiment of the present invention provides a driving behavior analysis method based on facial emotion recognition, such as Figure 1 As shown, the method includes:

[0013] Step S100: acquiring a real-time driver's facial video, and segmenting the real-time driver's facial video to obtain a plurality of target facial video segments.

[0014] In step S100, the behavior analysis device obtains a real-time driver's facial video, and segments the real-time driver's facial video to obtain a plurality of target facial video segments. In actual operation, for example, a real-time driver's facial video is obtained with the aid of a vehicle-mounted camera. The vehicle-mounted camera is usually installed in a suitable position inside the vehicle, such as above the dashboard, near the rearview mirror, etc., to ensure that the driver's facial image can be clearly captured. For example, in an ordinary family car, the vehicle-mounted camera is installed above the center of the dashboard, facing the driver's face, so that during the driving of the vehicle, the camera can continuously record the driver's facial dynamics to form a real-time driver's facial video. It should be noted that in the specific implementation, before information collection, it is necessary to obtain the authorization of the relevant object and conduct it under the permission of laws and regulations.

[0015] After acquiring the real-time driver facial video, the behavior analysis device needs to segment it to obtain multiple target facial video segments. The real-time driver facial video here can be regarded as a video stream composed of a series of continuous facial image frames, and the segmentation operation is to divide this continuous video stream into multiple relatively independent video segments according to certain rules.

[0016] In order to achieve video segmentation, one feasible method is to segment based on time intervals. The real-time driver facial video can be evenly divided into multiple target facial video segments according to the preset time intervals. Another segmentation method is to segment based on the changes in video content. The segmentation point can be determined by analyzing the changes in the features of the facial image in the video, such as obvious changes in facial expressions, significant changes in head posture, etc. When a significant change in the features of the facial image is detected, the behavior analysis device segments the video at this point to obtain different target facial video segments.

[0017] For segmentation based on time intervals, for example, the start and end positions of the segmentation are determined by the following formula: Assume that the total duration of the video is T (unit: seconds), the preset time interval is t (unit: seconds), and the segment sequence number is n (n=1, 2, 3,...), then the start time Sn and end time En of the nth target facial video segment can be calculated by the following formula: Sn=(n-1)*t, En=n*t. For example, when T=60 seconds and t=5 seconds, for the third target facial video segment, n=3, the start time S3=(3-1)*5=10 seconds, and the end time E3=3*5=15 seconds, that is, the segment contains video content from the 10th second to the 15th second.

[0018] For segmentation based on content changes, for example, machine learning algorithms are used to detect changes in facial image features. For example, a convolutional neural network (CNN) is used to extract and analyze facial images in a video. CNN can learn patterns of features such as facial expressions and head postures. A CNN model can be pre-trained to enable it to recognize different facial feature changes. In practical applications, each frame of a real-time driver's facial video is input into a trained CNN model, and the model outputs a feature vector for the frame. By comparing the similarity of feature vectors of adjacent frames, if the similarity is lower than a preset threshold, it is considered that a significant feature change has occurred, thereby determining the segmentation point. Assuming that the preset similarity threshold is 0.8, when the similarity of the feature vectors of two adjacent frames is 0.7, the segmentation point is determined between the two frames.

[0019] Step S200: Perform feature representation on the target facial video clip to obtain a video feature array of the real-time driver facial video, and generate an initial emotion recognition result corresponding to the real-time driver facial video based on the video feature array.

[0020] In step S200, the behavior analysis device performs feature representation on the target facial video clip, obtains a video feature array of the real-time driver facial video, and generates an initial emotion recognition result corresponding to the real-time driver facial video based on the video feature array. After obtaining the target facial video clip, the behavior analysis device needs to convert these video clips into a feature form that can be understood and processed by the computer, that is, perform feature representation to obtain a video feature array, and then preliminarily judge the driver's emotional state based on this array.

[0021] The purpose of the behavior analysis device to perform feature representation on the target facial video clip is to extract key features that can reflect the facial features and emotional information in the video for subsequent analysis and recognition. The feature representation process is actually to convert complex video data into simple and effective feature arrays, such as feature vectors, which can accurately describe the facial features and emotional changes in the video.

[0022] In order to realize feature representation, a convolutional neural network (CNN) based on deep learning can be used. CNN can automatically learn the features of facial images in videos, and extract feature information at different levels through multi-layer convolution and pooling operations. For example, each frame of the target facial video clip is input into a pre-trained CNN model, and the model will output the feature vector of the frame. Then, for example, these feature vectors are combined to obtain the feature representation of the entire target facial video clip. For example, for a target facial video clip containing 10 frames of images, the behavior analysis device inputs each frame of the image into the CNN model to obtain 10 feature vectors, and combines these 10 feature vectors according to certain rules to form a new feature vector as the feature representation of the video clip.

[0023] After obtaining the feature representation of the target facial video clip, these feature representations are fused to obtain a video feature array of the real-time driver facial video. During the fusion process, the feature information of different target facial video clips is integrated to form a complete feature representation, so as to more comprehensively describe the real-time driver facial video. For example, feature splicing is used to splice the feature vectors of each target facial video clip in sequence to form a longer feature vector as the video feature array of the real-time driver facial video. In addition to splicing, feature fusion can also be performed by weighted averaging. Specifically, different weights can be assigned to the feature vectors of each target facial video clip according to its importance, and then the weighted feature vectors are averaged to obtain a video feature array.

[0024] After obtaining the video feature array of the real-time driver's facial video, the initial emotion recognition result is generated based on the array. For example, a classifier is used to complete this task. The classifier can be a support vector machine (SVM), a decision tree, a neural network, etc. The behavior analysis device inputs the video feature array into the classifier. The classifier judges the driver's emotional state, such as happiness, sadness, anger, calmness, etc., based on the pre-learned patterns and rules, and outputs the corresponding emotion label as the initial emotion recognition result. For example, the video feature array [4.9, 6.1, 7.3] is input into a trained SVM classifier. After calculation and judgment, the classifier outputs the emotion label "calm", then this "calm" is the initial emotion recognition result corresponding to the real-time driver's facial video.

[0025] When performing initial emotion recognition, some other factors can also be considered to improve the accuracy of recognition. For example, the contextual information of the video, such as the previous and next clips of the video, the driving behavior of the driver, etc., can be combined to assist in judging the emotional state. If the driver shows nervous emotions in the previous clip of the video, and the emotional state shown by the video feature array of the current video clip is unclear, for example, based on the contextual information, it is more inclined to judge that the driver is still in a nervous state.

[0026] In step S200, the behavior analysis device converts complex real-time driver facial videos into simple emotion labels by performing feature representation, feature fusion and initial emotion recognition on the target facial video clips, thereby providing an important basis for subsequent driving behavior safety assessment.

[0027] Step S300: Detecting one or more core facial images in the real-time driver facial video, and detecting the distribution positions of one or more facial image sequences in the real-time driver facial video based on the core facial images.

[0028] In step S300, the behavior analysis device detects one or more core facial images in the real-time driver facial video, and based on these core facial images, detects the distribution positions of one or more facial image sequences in the real-time driver facial video. The core facial image refers to the facial image that can represent the key state of the driver's emotional fluctuation in the video, and the distribution position of the facial image sequence represents the specific range of these facial images related to emotional fluctuations in the entire video.

[0029] When the behavior analysis device detects the core facial image, it selects representative images from the multiple facial images contained in the real-time driver facial video. Specifically, a machine learning algorithm can be used to determine the importance of each facial image. For example, a convolutional neural network (CNN) based on deep learning is used to extract and analyze the features of the facial images in the video. The pre-trained CNN model can identify key features in the facial image, such as subtle changes in facial muscles, obvious changes in facial expressions, etc. When the model detects significant changes in these key features, the corresponding facial image may be determined as a core facial image. Suppose that in a real-time driver facial video, the driver's expression is relatively calm at first, and suddenly there is a significant change in angry expression such as frowning and glaring. After the behavior analysis device detects this change through the CNN model, the facial image with the angry expression is determined as the core facial image.

[0030] In addition to the CNN model, a rule-based method can also be used to detect core facial images. For example, some rules are pre-set, such as the amplitude and duration of facial expression changes. When a facial image meets these rules, it is determined to be a core facial image. For example, the rule is set that the amplitude of facial expression changes exceeds a certain threshold and the duration reaches a certain length. When the behavior analysis device detects that the amplitude of expression changes in a facial image exceeds the preset threshold and the change lasts for at least 3 frames, the facial image is marked as a core facial image.

[0031] After the core facial images are detected, the distribution position of the facial image sequence is determined based on these core facial images. A facial image sequence refers to a group of continuous or discontinuous facial images associated with specific emotional fluctuations, and the distribution position clarifies the specific range of these images in the video. In order to determine the distribution position, for example, a cluster analysis method is used. The core facial images are clustered according to the similarity of emotional fluctuations, and the core facial images belonging to the same type of emotional fluctuations are classified into one category. Then, based on these clusters, the facial image sequences related to them are searched in the real-time driver facial video. The behavior analysis device can also determine the distribution position of the facial image sequence by the time window method. The behavior analysis device takes the core facial image as the center and sets a fixed-size time window. All facial images within this time window are considered to be facial image sequences related to the core facial image.

[0032] Step S400: segmenting the video feature array based on the distribution position to obtain a target video feature array for each facial image sequence.

[0033] In step S400, the behavior analysis device segments the video feature array based on the distribution position of the facial image sequence to obtain the target video feature array for each facial image sequence. The video feature array is obtained after the behavior analysis device performs feature representation on the real-time driver's facial video. It contains the feature information of the entire video, and the distribution position of the facial image sequence clarifies the specific range of facial images related to different emotional fluctuations in the video. By segmenting the video feature array according to the distribution position, the behavior analysis device can divide the video features according to the different stages of emotional fluctuations, so as to more specifically analyze the emotional state corresponding to each facial image sequence.

[0034] After obtaining the distribution position of the facial image sequence, the behavior analysis device needs to extract the features related to each facial image sequence from the video feature array based on the position information. In order to achieve the segmentation of the video feature array, for example, an indexing method is used. The behavior analysis device converts the distribution position of the facial image sequence into an index range in the video feature array, and then directly extracts features from the video feature array based on this index range. Assuming that the video feature array is F=[f1, f2,..., fn], the index range corresponding to the distribution position of a facial image sequence is from i to j (1≤i≤j≤n), then the target video feature array Ft of the facial image sequence can be expressed as Ft=[fi, fi+1,..., fj]. For example, when n=100, i=10, j=20, Ft=[f10, f11,..., f20].

[0035] In addition to the indexing method, the behavior analysis device can also use the sliding window method for segmentation. With the distribution position of the facial image sequence as the center, set a fixed-size window, slide this window in the video feature array, and extract the features in the window as the target video feature array. Assuming that the window size is k and the index corresponding to the distribution position of the facial image sequence is m, the target video feature array Ft can be expressed as Ft=[fm-(k-1) / 2, fm-(k-1) / 2+1,..., fm+(k-1) / 2] (assuming k is an odd number). For example, when k=5 and m=15, Ft=[f13, f14, f15, f16, f17].

[0036] Step S500: Iterating the initial emotion recognition result according to the target video feature array to obtain the target emotion recognition result corresponding to the real-time driver's facial video.

[0037] In step S500, the behavior analysis device iterates the initial emotion recognition result based on the target video feature array to obtain the target emotion recognition result corresponding to the real-time driver's facial video. The target video feature array is obtained after the behavior analysis device segments the video feature array based on the distribution position of the facial image sequence. Each target video feature array corresponds to a facial image sequence and contains the feature information of the facial images in the sequence. The initial emotion recognition result is an emotion judgment preliminarily generated by the behavior analysis device based on the video feature array of the entire real-time driver's facial video. By iterating the initial emotion recognition result, the behavior analysis device can combine the specific features of each facial image sequence to more accurately identify the driver's emotional state, thereby obtaining a more accurate target emotion recognition result.

[0038] The behavior analysis device determines the emotional information corresponding to each facial image sequence based on the target video feature array. To achieve this, for example, a classification model such as a support vector machine (SVM), a random forest, etc. is used. The behavior analysis device inputs each target video feature array into a pre-trained classification model, and the model determines the emotional state corresponding to the facial image sequence based on the information in the feature array. For example, in a facial image sequence corresponding to a target video feature array, the driver's facial expression shows features such as raised corners of the mouth and squinting eyes. The behavior analysis device inputs the target video feature array into the SVM classification model. After calculation and judgment, the model outputs the emotional state corresponding to the facial image sequence as "happy".

[0039] The behavior analysis device can also use deep learning models, such as recurrent neural networks (RNN) or long short-term memory networks (LSTM), to process the target video feature array. These models are able to process sequence data, taking into account the temporal continuity and correlation in the facial image sequence. The behavior analysis device inputs the target video feature array into the LSTM model in chronological order, and the model learns the changing pattern of features in the sequence to more accurately identify the emotional state. Assuming that a facial image sequence corresponding to a target video feature array records the process of the driver gradually changing from calm to angry, the LSTM model can capture the dynamic changes of this emotion and accurately determine the emotional state at different stages in the sequence.

[0040] After determining the emotional information corresponding to each facial image sequence, the emotional information is integrated to iterate the initial emotion recognition result. For example, a voting method is used to integrate emotional information. The behavior analysis device counts the number of times each emotional state appears in all facial image sequences, and the emotional state with the most occurrences is used as the final iterative result. In addition to the voting method, the behavior analysis device can also use a weighted average method. A weight is assigned to each facial image sequence, and the weight can be determined based on factors such as the length of the sequence and the significance of the features. Then, the behavior analysis device performs a weighted average on the emotional state corresponding to each facial image sequence to obtain the iterated emotional state.

[0041] Step S600: Evaluate the safety of the driver's driving behavior based on the target emotion recognition result.

[0042] In step S600, the behavior analysis device evaluates the driver's driving behavior safety based on the target emotion recognition result. The target emotion recognition result is an accurate judgment of the driver's emotional state presented in the real-time facial video obtained by the behavior analysis device after a series of processing, and the driving behavior safety assessment is based on this emotional result to determine whether the driver's current emotion will have a negative impact on his driving operation, thereby affecting driving safety.

[0043] The behavior analysis device first establishes a correlation model between the target emotion and the safety of driving behavior. Different emotional states have different degrees of influence on the safety of driving behavior. For example, negative emotions such as anger and anxiety may cause the driver to lose concentration, slow down reaction speed, and increase operational errors, thereby increasing the risk of traffic accidents; while positive emotions such as calmness and happiness are relatively more conducive to safe driving.

[0044] In order to establish an association model, machine learning algorithms such as logistic regression and decision trees are used. Taking logistic regression as an example, the behavior analysis device uses the target emotion recognition result as the input feature and whether the driving behavior is safe (0 indicates unsafe and 1 indicates safe) as the output label. By learning a large amount of training data, a logistic regression model is obtained. The formula of this model can be expressed as: P(Y=1|X)=1 / (1+e (-(β0+β1X1+β2X2 +...+βnXn)) ), where P(Y=1|X) represents the probability of safe driving behavior given the input feature X (i.e., the target emotion recognition result), β0, β1,..., β n are the parameters of the model, X1, X2,..., X n are the dimensions of the input features.

[0045] After obtaining the association model, the behavior analysis device inputs the target emotion recognition result into the model to calculate the probability of safe driving behavior. In addition to using probability to evaluate the safety of driving behavior, the behavior analysis device can also adopt a grade classification method. The behavior analysis device divides the driving behavior safety into multiple levels, such as safe, relatively safe, relatively dangerous, dangerous, etc., according to the degree of influence of different emotional states on the safety of driving behavior. For example, when the driver is in a calm mood, the behavior analysis device determines the driving behavior safety level as "safe"; when the driver is in an anxious mood, it is determined to be "relatively safe"; when the driver is in an angry mood, it is determined to be "relatively dangerous"; when the driver is in an extremely angry or fearful mood, it is determined to be "dangerous".

[0046] As an implementation manner, the driver's facial video includes multiple facial images. In step S300, one or more core facial images are detected in the real-time driver's facial video, including:

[0047] Step S310: inferring the prediction confidence of the emotional state corresponding to each facial image based on the video feature array;

[0048] Step S320: Detecting and obtaining node distribution intervals of one or more emotion fluctuation nodes in the real-time driver's facial video based on the prediction confidence;

[0049] Step S330: selecting one or more facial images corresponding to the node distribution interval from the facial images to obtain a core facial image corresponding to each emotion fluctuation node.

[0050] In step S310, the behavior analysis device infers the prediction confidence of the emotional state corresponding to each facial image based on the video feature array. The video feature array is obtained after the behavior analysis device performs feature representation on the target facial video clip, and it contains the feature information of the entire real-time driver facial video. The prediction confidence is a quantitative representation of the reliability of the judgment of the emotional state corresponding to each facial image. The higher the confidence, the more confident the behavior analysis device is in judging the emotional state of the facial image.

[0051] In order to infer the prediction confidence, for example, a classifier model is used. Viable classifiers include support vector machines (SVM), decision trees, neural networks, etc. The relevant concepts have been introduced above and will not be repeated here.

[0052] In step S320, the behavior analysis device detects the node distribution interval of one or more emotion fluctuation nodes in the real-time driver's facial video based on the prediction confidence. The emotion fluctuation node refers to the moment when the driver's emotion changes significantly, and the node distribution interval indicates the location of the specific time period in the video when these emotion fluctuations occur.

[0053] For example, emotion fluctuation nodes can be detected by monitoring changes in prediction confidence. When the prediction confidence of adjacent facial images changes significantly, it can be considered that an emotion fluctuation node has appeared. For example, in consecutive facial images, the prediction confidence of the "happy" emotion corresponding to the previous image is 0.8, while the prediction confidence of the "happy" emotion corresponding to the next image suddenly drops to 0.2, while the prediction confidence of the "angry" emotion increases from 0.1 to 0.7. This significant change indicates that an emotion fluctuation node may have appeared.

[0054] In order to determine the node distribution interval, a sliding window method can be used. Set a window of fixed size, slide the window on the timeline of the video, and calculate the rate of change of the prediction confidence within the window. When the rate of change exceeds a preset threshold, it is considered that there is an emotion fluctuation node in the window, and the start and end positions of the window are recorded as the node distribution interval. Assuming the window size is 10 frames and the threshold is 0.5, when the behavior analysis device calculates that the rate of change of the "happy" emotion prediction confidence in a window is 0.6, the time period corresponding to the window is determined as the node distribution interval of an emotion fluctuation node.

[0055] In addition, cluster analysis methods can also be used to detect emotion fluctuation nodes. The behavior analysis device regards the prediction confidence as data points and clusters them according to the similarity of these data points. When there is obvious discontinuity in the data points within a cluster, it can be considered that there is an emotion fluctuation node in the time period corresponding to the cluster. For example, the behavior analysis device clusters the "happy" emotion prediction confidence of all facial images and finds that one cluster can be divided into two obvious sub-clusters. The transition area between the two sub-clusters may contain emotion fluctuation nodes.

[0056] In step S330, one or more facial images corresponding to the node distribution interval are selected from the facial images to obtain a core facial image corresponding to each emotion fluctuation node. The core facial image is a facial image that can represent the key state of emotion fluctuation. For example, the corresponding image is directly selected from the facial image according to the node distribution interval.

[0057] The behavior analysis device can also combine the feature information of the facial image to screen the core facial image. For example, the facial expression features in the facial image are analyzed, such as the degree of eyebrow raising, the bending angle of the corners of the mouth, etc. For an "angry" emotion fluctuation node, the behavior analysis device selects the facial image with the highest eyebrow raising and the most obvious downward bending of the corners of the mouth as the core facial image. In actual applications, the behavior analysis device may detect multiple emotion fluctuation nodes, each of which has a corresponding core facial image. These core facial images constitute a key image set, for example, based on this set for more in-depth emotion analysis and driving behavior evaluation.

[0058] Through steps S310-S330, the behavior analysis device can accurately detect core facial images in the real-time driver facial video. These core facial images reflect the key moments of the driver's emotions, providing an important basis for subsequent analysis of the driver's emotional state and evaluation of driving behavior safety.

[0059] As an implementation manner, in step S300, based on the core facial image, the distribution position of one or more facial image sequences is detected in the real-time driver facial video, including:

[0060] Step S340: dividing the emotion fluctuation nodes based on the node distribution interval to obtain a plurality of emotion fluctuation node sets, wherein the emotion fluctuation node sets include a set number of emotion fluctuation nodes;

[0061] Step S350: Detecting the image range of the facial image sequence corresponding to each emotion fluctuation node set in the real-time driver facial video according to the core facial image;

[0062] Step S360: augment the image range to obtain a target image range, and use the target image range as the distribution position of the facial image sequence.

[0063] In step S340, the behavior analysis device divides the emotion fluctuation nodes based on the node distribution interval to obtain multiple emotion fluctuation node sets, and the emotion fluctuation node set includes a set number of emotion fluctuation nodes. The node distribution interval is obtained by the behavior analysis device in step S320 based on the prediction confidence detection, which indicates the position of the specific time period in the video when the emotion fluctuation occurs. The emotion fluctuation node is the moment when the driver's emotion changes significantly. By dividing the emotion fluctuation nodes, the related emotion fluctuations can be grouped for subsequent analysis.

[0064] For example, the emotion fluctuation nodes are divided according to the time sequence and the continuity of the node distribution interval. The set number of each emotion fluctuation node set is set to 3. The behavior analysis device selects 3 adjacent emotion fluctuation nodes in the order in which the emotion fluctuation nodes appear in the video to form a set. Assuming that 6 emotion fluctuation nodes are detected in the real-time driver's facial video, and their node distribution intervals are [10-15], [20-25], [30-35], [40-45], [50-55], and [60-65], the behavior analysis device divides the first 3 nodes into one set and the last 3 nodes into another set.

[0065] In order to more reasonably divide the emotion fluctuation node set, the time interval between the node distribution intervals can also be considered. If the time interval between two adjacent node distribution intervals is too long, for example, they are divided into different sets. Assuming that the time interval threshold is set to 10 frames, when the distribution intervals of two adjacent nodes are [10-15] and [30-35] respectively, and the time interval is 15 frames, which exceeds the threshold, the behavior analysis device divides the two nodes into different sets.

[0066] In step S350, based on the core facial image, the image range of the facial image sequence corresponding to each set of emotion fluctuation nodes is detected in the real-time driver's facial video. The facial image sequence is a set of continuous or discontinuous facial images related to specific emotion fluctuations, and the image range specifies the specific number of frames covered by these images in the video.

[0067] For example, the image range is determined by analyzing the features and position of the core facial image. The core facial image represents the key state of emotional fluctuations. For example, the image range is determined by extending a certain number of frames forward and backward with the core facial image as the center. Assuming that the behavior analysis device is set to extend 5 frames forward and backward with the core facial image as the center, and for a core facial image in an emotional fluctuation node set located at the 20th frame, the image range of the facial image sequence corresponding to the core facial image is from the 15th frame to the 25th frame.

[0068] The behavior analysis device can also determine the image range in combination with the intensity of emotional fluctuations. If the emotional fluctuation intensity corresponding to a certain emotional fluctuation node set is large, for example, the image range is appropriately expanded to more comprehensively capture the process of emotional changes. For example, the emotional fluctuation intensity index is used to measure the severity of emotional changes. When the intensity index exceeds a certain threshold, the image range is expanded to 10 frames before and after.

[0069] In step S360, the behavior analysis device expands the image range to obtain a target image range, and uses the target image range as the distribution position of the facial image sequence. The purpose of expanding the image range is to more comprehensively include facial image information related to emotional fluctuations and improve the accuracy of subsequent emotional analysis.

[0070] For example, a method of fixed frame number augmentation is adopted. The behavior analysis device sets a fixed augmented frame number, such as 3 frames each before and after augmentation. For a facial image sequence with an image range of [15-25], the target image range after augmentation is from the 12th frame to the 28th frame.

[0071] In addition, the image can be augmented based on its feature similarity. Images with similar features to the facial images in the image range can be found before and after the image range, and these images can be included in the target image range. Assume that the feature similarity is measured by calculating the cosine similarity between the feature vectors of the images. When the similarity exceeds a certain threshold, the image is included in the target image range.

[0072] After determining the target image range, the behavior analysis device uses it as the distribution position of the facial image sequence. This distribution position clarifies the specific position of each facial image sequence in the real-time driver facial video, providing an important basis for the subsequent segmentation of the video feature array based on the distribution position.

[0073] In step S340, the behavior analysis device may also use a clustering algorithm when dividing the emotion fluctuation node set. The clustering algorithm can cluster similar nodes into one category according to the characteristics of the emotion fluctuation nodes, such as the node distribution interval, prediction confidence, etc., to form an emotion fluctuation node set. For example, using the K-means clustering algorithm, taking the start frame and the end frame of the node distribution interval of the emotion fluctuation node as feature vectors, setting the number of clusters to K, the algorithm will automatically divide the nodes into K sets.

[0074] In step S360, when the behavior analysis device expands the image range, context information can also be introduced. For example, the video content before and after the image range is analyzed to determine whether there is context information related to the current emotional fluctuation. If there is relevant context information, for example, the image containing this information is included in the target image range. For example, before the image range corresponding to an emotional fluctuation node set, event scenes that may cause the emotional fluctuation appear in the video, for example, the images containing these event scenes are included in the target image range, so as to more comprehensively understand the cause of the emotional fluctuation.

[0075] When dividing the set of emotional fluctuation nodes in step S340, the correlation between nodes can also be considered. Some emotional fluctuation nodes may be caused by the same reason, and there is an inherent logical connection between them. For example, by analyzing the emotional state corresponding to the node and the relevant external factors, the nodes with correlation are divided into the same set. For example, the driver has multiple emotional fluctuations due to continuous traffic jams. These fluctuation nodes can be considered to be related and should be divided into the same set. In order to measure the correlation between nodes, for example, a correlation matrix is ​​constructed, and the elements in the matrix represent the degree of correlation between two nodes. The degree of correlation can be calculated by a variety of factors, such as the similarity of emotional states, the closeness of time intervals, etc. Assume that a comprehensive indicator is used to calculate the degree of correlation, which takes into account the cosine similarity of emotional states and the inverse of time intervals. The formula is: correlation degree = α* emotional state cosine similarity + (1-α) * (1 / time interval), where α is a weight coefficient, ranging from 0 to 1.

[0076] As an implementation mode, step S350, based on the core facial image, detects the image range of the facial image sequence corresponding to each emotion fluctuation node set in the real-time driver facial video, including:

[0077] Step S351: based on the node distribution interval, selecting a triggering emotion fluctuation node and a terminating emotion fluctuation node from the emotion fluctuation node set;

[0078] Step S352: Detecting the number of core facial images corresponding to the triggering emotion fluctuation node in the real-time driver's facial video to obtain the number of triggering images;

[0079] Step S353: Detecting the number of core facial images corresponding to the termination emotion fluctuation node in the real-time driver's facial video to obtain the number of termination images;

[0080] Step S354: taking the range formed by the number of trigger images and the number of termination images as the image range of the facial image sequence corresponding to the emotion fluctuation node set.

[0081] In step S351, based on the node distribution interval, the triggering emotion fluctuation node and the terminating emotion fluctuation node are selected from the emotion fluctuation node set. The triggering emotion fluctuation node refers to the node corresponding to the starting moment of a certain emotion fluctuation, and the terminating emotion fluctuation node is the node corresponding to the ending moment of the emotion fluctuation. The behavior analysis device determines these two key nodes through the node distribution interval. The node distribution interval represents the position of the specific time period in the video when the emotion fluctuation occurs. For example, the triggering and terminating nodes are determined according to the starting and ending time of the node distribution interval. Assuming that an emotion fluctuation node set contains three nodes, and their node distribution intervals are [10-15], [20-25], and [30-35] respectively, the behavior analysis device determines the node corresponding to the first node distribution interval (i.e., the node corresponding to [10-15]) as the triggering emotion fluctuation node, and determines the node corresponding to the last node distribution interval (i.e., the node corresponding to [30-35]) as the terminating emotion fluctuation node.

[0082] In practical applications, the prediction confidence can also be combined to more accurately select trigger and termination nodes. If the prediction confidence of a node is high and its node distribution interval is at the beginning or end of the emotional fluctuation, then the node is more likely to be a trigger or termination node. Assuming that in the above example, the prediction confidence of the node corresponding to [10-15] is 0.9, and the prediction confidence of the node corresponding to [30-35] is 0.85, the higher prediction confidence further confirms their rationality as trigger and termination nodes.

[0083] In step S352, the number of core facial images corresponding to the triggering emotion fluctuation nodes is detected in the real-time driver's facial video to obtain the number of triggered images. The core facial image is a facial image that can represent the key state of emotion fluctuations, and the number of triggered images indicates the specific number of core facial images corresponding to the triggering emotion fluctuation nodes in the video. For example, by finding the position of the core facial images corresponding to the triggering emotion fluctuation nodes in the video and counting their number. Assuming that the positions of the core facial images corresponding to the triggering emotion fluctuation nodes in the video are the 12th frame, the 13th frame and the 14th frame, then the number of triggered images is 3.

[0084] In order to more accurately detect the number of triggered images, for example, image feature matching can be combined. For example, the features of the core facial image corresponding to the node that triggers emotional fluctuations are extracted in advance, and then images with similar features are searched in the video. By calculating the similarity between the image feature vectors, when the similarity exceeds a preset threshold, the image is considered to be the core facial image corresponding to the node that triggers emotional fluctuations. Assuming that cosine similarity is used to measure feature similarity, and the threshold is set to 0.8, when the cosine similarity between the feature vector of an image in the video and the core facial image corresponding to the node that triggers emotional fluctuations is 0.85, the image will be counted as a trigger image.

[0085] In step S353, the behavior analysis device detects the number of core facial images corresponding to the termination emotion fluctuation node in the real-time driver's facial video, and obtains the number of termination images. Similar to the number of trigger images, the number of termination images indicates the specific number of core facial images corresponding to the termination emotion fluctuation node in the video. The behavior analysis device can also find the position of the core facial images corresponding to the termination emotion fluctuation node in the video and count their number. Assuming that the core facial images corresponding to the termination emotion fluctuation node are located at the 32nd and 33rd frames in the video, the number of termination images is 2.

[0086] The behavior analysis device can also use the same feature matching method as that used to detect the number of trigger images to detect the number of termination images. The features of the core facial image corresponding to the termination emotion fluctuation node are extracted in advance, and images with similar features are searched in the video, and the number of termination images is determined based on the feature similarity.

[0087] In step S354, the behavior analysis device uses the range formed by the number of trigger images and the number of termination images as the image range of the facial image sequence corresponding to the emotion fluctuation node set. This image range clarifies the specific number of frames covered by the facial images related to the emotion fluctuation node set in the video. Assuming that the image position range corresponding to the number of trigger images is from the 12th frame to the 14th frame, and the image position range corresponding to the number of termination images is from the 32nd frame to the 33rd frame, then the image range of the facial image sequence corresponding to the emotion fluctuation node set is from the 12th frame to the 33rd frame.

[0088] In practical applications, for example, this image range is appropriately adjusted. Considering that emotional fluctuations may have a certain continuity, for example, a certain number of frames are appropriately extended before the start frame corresponding to the number of trigger images and after the end frame corresponding to the number of end images. Assuming that the extended number of frames is set to 2, the adjusted image range is from the 10th frame to the 35th frame.

[0089] Through steps S351-S354, the behavior analysis device can accurately determine the image range of the facial image sequence corresponding to the emotion fluctuation node set in the real-time driver's facial video. These image ranges reflect the specific coverage of facial images related to specific emotion fluctuations, and provide an important basis for the subsequent segmentation of the video feature array based on these image ranges, further analyzing the driver's emotional state and evaluating the safety of driving behavior.

[0090] As an implementation manner, step S360, augmenting the image range to obtain the target image range, includes:

[0091] Step S361: subtract the number of trigger images from the first set number of images to obtain the target number of trigger images;

[0092] Step S362: Add the number of termination images to the second set number of images to obtain the target number of termination images;

[0093] Step S363: taking the range formed by the number of target trigger images and the number of target termination images as the target image range.

[0094] In step S361, the behavior analysis device subtracts the number of starting images corresponding to the number of trigger images from the number of first set images to obtain the number of target trigger images. The number of trigger images is the number of core facial images corresponding to the trigger emotion fluctuation node detected by the behavior analysis device in the previous step, and the first set number of images is a pre-set fixed value used to adjust the number of trigger images. The target number of trigger images is the starting image position corresponding to the adjusted trigger image. Assuming that the starting image corresponding to the number of trigger images is located at the 10th frame of the video, and the first set number of images is 3, the behavior analysis device calculates 10-3=7, and obtains that the starting image position corresponding to the target number of trigger images is the 7th frame. In one embodiment, in order to determine the first set number of images, it can be set according to experimental data and experience. By analyzing the facial videos of different drivers, observing the number of frames involved in some potential feature changes before the start of emotion fluctuations, a suitable value is statistically calculated as the first set number of images. A machine learning method can also be used to train a regression model using historical data, with the input being the relevant features of emotion fluctuations (such as the intensity and type of emotion fluctuations), and the output being the first set number of images. In this way, the number of the first set images can be dynamically adjusted according to different emotional fluctuations.

[0095] In step S362, the behavior analysis device adds the number of ending images corresponding to the number of ending images to the second set number of images to obtain the target number of ending images. The number of ending images is the number of core facial images corresponding to the ending emotional fluctuation node detected by the behavior analysis device, and the second set number of images is also a pre-set fixed value used to adjust the number of ending images. The target number of ending images is the ending image position corresponding to the adjusted ending image. Assuming that the ending image corresponding to the number of ending images is at the 20th frame of the video, and the second set number of images is 2, the behavior analysis device calculates 20+2=22, and obtains that the ending image position corresponding to the target number of ending images is the 22nd frame.

[0096] The second set number of images is determined in a similar manner to the first set number of images. Based on experimental data and experience, the number of frames involved in the emotional continuation features that may exist after the emotional fluctuation ends can be observed to determine a suitable value. A machine learning model can also be used to dynamically predict the second set number of images based on the characteristics of emotional fluctuations.

[0097] In step S363, the behavior analysis device uses the range formed by the target trigger image number and the target termination image number as the target image range. The target image range is the range of the facial image sequence after augmentation, which contains more facial image information related to emotional fluctuations. Assuming that the starting image position corresponding to the target trigger image number is the 7th frame, and the ending image position corresponding to the target termination image number is the 22nd frame, then the target image range is from the 7th frame to the 22nd frame. After obtaining the target image range, it is used as the distribution position of the facial image sequence. This distribution position clarifies the specific position of each facial image sequence in the real-time driver facial video, which provides an important basis for the subsequent segmentation of the video feature array based on the distribution position, further analysis of the driver's emotional state and evaluation of driving behavior safety.

[0098] When determining the number of second set images in step S362, in some embodiments, an emotion recovery model can be introduced. Different people take different times to return to a calm state after experiencing emotional fluctuations, which is related to factors such as personal psychological quality and emotional regulation ability. For example, by collecting a large amount of emotional data and related recovery time information of drivers, an emotion recovery model is established. The model can predict how long it takes for the driver to return to calm after the emotional fluctuation ends based on the driver's personal information (such as age, gender, driving experience, etc.) and the type and intensity of emotional fluctuations, thereby determining the appropriate number of second set images. For example, for young drivers with less driving experience, it may take a long time to return to calm after experiencing angry emotional fluctuations, and the number of second set images can be relatively large; while for older drivers with rich driving experience, the recovery time may be shorter, and the number of second set images can be relatively small.

[0099] In order to further improve the accuracy of the target image range, multi-sensor fusion technology can also be used. In addition to facial video data, the driver's eye movement data, head posture data, etc. can also be combined. Eye movement data can reflect the driver's attention direction and psychological state, and head posture data can assist in judging the driver's emotional changes. For example, these multi-sensor data are fused, and the start and end times of emotional fluctuations are more accurately judged based on the fused information, so as to adjust the first set image number and the second set image number to obtain a more accurate target image range.

[0100] As an implementation mode, the video feature array includes a sub-video feature array corresponding to each facial image. In step S400, based on the distribution position, the video feature array is segmented to obtain a target video feature array for each facial image sequence, including:

[0101] Step S410: selecting one or more facial images corresponding to the target image range from the facial images to obtain a candidate facial image set corresponding to each facial image sequence;

[0102] Step S420: extracting the sub-video feature array corresponding to each facial image in the candidate facial image set from the video feature array, and obtaining a sub-video feature array set corresponding to each facial image sequence;

[0103] Step S430: determining the sub-video feature array set as the target video feature array corresponding to the facial image sequence.

[0104] In step S410, the behavior analysis device selects one or more facial images corresponding to the target image range from the facial image to obtain a candidate facial image set corresponding to each facial image sequence. The target image range is obtained after the behavior analysis device augments the image range in step S360, and the specific position of each facial image sequence in the real-time driver facial video is clarified. The candidate facial image set is a collection of all facial images covered by the facial image sequence.

[0105] In order to accurately select the facial image corresponding to the target image range, for example, the frame index information of the video is used. The video can assign a unique index to each frame image, for example, according to the start and end frame indexes of the target image range, the corresponding facial image is directly extracted from the video. The behavior analysis device can also use image matching technology to verify whether the selected facial image is accurate. The image features of the boundary frame of the target image range are extracted in advance, and then feature matching is performed in the actual selected facial image to ensure that the selected image is consistent with the target image range.

[0106] In step S420, the behavior analysis device extracts the sub-video feature array corresponding to each facial image in the candidate facial image set from the video feature array, and obtains a sub-video feature array set corresponding to each facial image sequence. The video feature array is obtained after the behavior analysis device performs feature representation on the target facial video clip in step S200, and it contains the feature information of the entire real-time driver facial video. The sub-video feature array is the feature representation corresponding to each facial image. Assuming that the candidate facial image set contains 5 facial images, the behavior analysis device finds the sub-video feature arrays corresponding to these 5 images in the video feature array, and combines them into a sub-video feature array set corresponding to the facial image sequence.

[0107] In order to accurately extract the sub-video feature array, for example, a mapping relationship between the facial image and the sub-video feature array is established. When the video feature array is generated in step S200, the position or index of the sub-video feature array corresponding to each facial image is recorded. When it is necessary to extract the sub-video feature array in the candidate facial image set, for example, the corresponding sub-video feature array is directly found according to this mapping relationship. The behavior analysis device can also be verified by a feature matching method. The features of the images in the candidate facial image set are extracted and matched with the sub-video feature array in the video feature array to ensure that the extracted sub-video feature array corresponds to the facial image.

[0108] In step S430, the behavior analysis device determines the sub-video feature array set as the target video feature array corresponding to the facial image sequence. The target video feature array is a feature array that directly corresponds to each facial image sequence after screening and extraction. It contains the key feature information of the facial image sequence and provides an important basis for subsequent emotion recognition and analysis. The behavior analysis device directly uses the sub-video feature array set obtained in step S420 as the target video feature array.

[0109] After determining the target video feature array, it can be further analyzed and processed. The statistical features of the target video feature array, such as the mean, variance, standard deviation, etc., are calculated to understand the feature distribution of the facial image sequence. The behavior analysis device can also input the target video feature array into a classifier for emotion recognition and classification, providing a basis for subsequent iteration of the initial emotion recognition results.

[0110] Steps S410-S430 provide a specific method for the behavior analysis device to extract the target video feature array corresponding to each facial image sequence from the video feature array. By accurately selecting the candidate facial image set, extracting the sub-video feature array set and determining the target video feature array, combined with considering image quality, preprocessing, establishing mapping relationships, verifying data consistency, and associating storage information, the behavior analysis device can more efficiently and accurately complete the task of extracting the target video feature array, laying a solid foundation for subsequent emotion analysis and driving behavior safety assessment.

[0111] As an implementation mode, step S500, iterating the initial emotion recognition result according to the target video feature array to obtain the target emotion recognition result corresponding to the real-time driver's facial video, includes:

[0112] Step S510: determining the target facial emotion corresponding to each facial image sequence according to the target video feature array;

[0113] Step S520: Based on the distribution position, the target facial emotions are fused to obtain fused facial emotions;

[0114] Step S530: The initial emotion recognition result is changed to the fused facial emotion to obtain the target emotion recognition result corresponding to the real-time driver's facial video.

[0115] In step S510, the behavior analysis device determines the target facial emotion corresponding to each facial image sequence based on the target video feature array. The target video feature array is obtained by the behavior analysis device in step S400 by segmenting the video feature array based on the distribution position of the facial image sequence, and contains the key feature information of each facial image sequence. The target facial emotion is an accurate judgment of the emotional state of the driver reflected by each facial image sequence.

[0116] In order to determine the target facial emotion, a classifier model can be used. Feasible classifiers include support vector machine (SVM), decision tree, neural network, etc. Please refer to the above introduction for the relevant principles of the classifier, which will not be repeated here.

[0117] In step S520, the behavior analysis device fuses the target facial emotions based on the distribution position to obtain the fused facial emotions. The distribution position clarifies the specific position of each facial image sequence in the real-time driver facial video. By considering the distribution position, the target facial emotions of different facial image sequences can be reasonably integrated.

[0118] For example, a simple voting method is used for fusion. The number of times each emotional state appears in all facial image sequences is counted, and the emotional state with the most occurrences is used as the fused facial emotion. Alternatively, a weighted average method is used for fusion. A weight is assigned to each facial image sequence, and the weight can be determined based on factors such as the length of the sequence and the significance of the features. Then, the target facial emotions corresponding to each facial image sequence are weighted averaged to obtain the fused facial emotion.

[0119] In step S530, the behavior analysis device changes the initial emotion recognition result to the fused facial emotion, and obtains the target emotion recognition result corresponding to the real-time driver's facial video. The initial emotion recognition result is the emotion judgment initially generated by the behavior analysis device in step S200 based on the video feature array of the entire real-time driver's facial video. By changing it to the fused facial emotion, the emotion recognition result can more accurately reflect the driver's real emotional state in the video.

[0120] As an implementation method, step S200, performing feature representation on the target facial video clip to obtain a video feature array of the real-time driver facial video, includes:

[0121] Step S210: performing feature representation on the target facial video clips based on the video processing network to obtain an initial video feature array corresponding to each target facial video clip;

[0122] Step S220: The initial video feature array is fused to obtain a video feature array corresponding to the real-time driver's facial video.

[0123] In step S210, the behavior analysis device performs feature representation on the target facial video clip based on the video processing network to obtain an initial video feature array corresponding to each target facial video clip. The behavior analysis device inputs the target facial video clip into the trained video processing network, which extracts and converts the input video clip through a series of convolution, pooling, full connection and other operations, and finally outputs the corresponding initial video feature array.

[0124] To illustrate with a simple example, suppose the target facial video clip is a 10-second video with a frame rate of 30 frames per second, then the video clip contains 300 frames of images. The behavior analysis device inputs these 300 frames of images into the video processing network, and the network extracts features from each frame of the image to obtain the feature vector corresponding to each frame of the image. Then, the network combines and transforms these feature vectors, and finally outputs an initial video feature array representing the entire target facial video clip.

[0125] In the implementation process, the video processing network can adopt a variety of different architectures, such as convolutional neural network (CNN), recurrent neural network (RNN), etc. Among them, CNN is suitable for processing image data and can effectively extract the spatial features of the image; RNN is suitable for processing sequence data and can capture the time information in the sequence. For example, choose the appropriate network architecture according to specific needs and scenarios.

[0126] In step S220, the behavior analysis device fuses the initial video feature arrays to obtain a video feature array corresponding to the real-time driver's facial video. In actual operation, since the target facial video segment is obtained by segmenting the real-time driver's facial video, the initial video feature array corresponding to each target facial video segment only represents the local features of the segment. In order to obtain the feature representation of the entire real-time driver's facial video, the behavior analysis device needs to fuse these initial video feature arrays. For example, a variety of different fusion methods are used, such as average fusion, weighted fusion, etc.

[0127] As an implementation manner, step S210, before performing feature representation on the target facial video segments based on the video processing network to obtain an initial video feature array corresponding to each target facial video segment, further includes:

[0128] Step S201: obtaining a plurality of training facial video segments corresponding to the facial video training instance, and performing feature representation on the training facial video segments based on the video processing network to be trained to obtain a training video feature array of the facial video training instance;

[0129] Step S202: inferring the emotional state corresponding to the facial video training instance based on the training video feature array to obtain an inferred emotion recognition result;

[0130] Step S203: Based on the inferred emotion recognition result, the recognition error of the facial video training instance is determined, and according to the recognition error, the parameters of the video processing network to be trained are adjusted to obtain the video processing network.

[0131] Steps S201 to S203 before step S210 are the process of the behavior analysis device training the video processing network.

[0132] In step S201, multiple training facial video clips corresponding to the facial video training instance are obtained, and the training facial video clips are characterized based on the video processing network to be trained to obtain the training video feature array of the facial video training instance. The facial video training instance contains facial videos of different drivers in various driving scenarios. Then, according to the rules of the aforementioned step S100, these facial video training instances are divided into multiple training facial video clips. For example, the segmentation can be based on factors such as the length of the video and emotional fluctuations. Next, the behavior analysis device inputs these training facial video clips into the video processing network to be trained. The network will extract and convert features for each training facial video clip, and finally output the training video feature array of the facial video training instance. For details, please refer to the introduction of the aforementioned step S100.

[0133] In step S202, the behavior analysis device infers the emotional state corresponding to the facial video training instance based on the training video feature array to obtain the inferred emotion recognition result. After obtaining the training video feature array, the behavior analysis device uses the array to infer the emotional state of the driver in the facial video training instance. For example, emotion-related features are extracted from the training video feature array, and then these features are input into a classifier, which judges the driver's emotional state, such as happiness, sadness, anger, etc., based on these features, and finally obtains the inferred emotion recognition result.

[0134] In step S203, the behavior analysis device determines the recognition error of the facial video training instance based on the inferred emotion recognition result, and adjusts the parameters of the video processing network to be trained according to the recognition error to obtain the video processing network. After obtaining the inferred emotion recognition result, the behavior analysis device compares it with the actual emotional state of the facial video training instance and calculates the recognition error. The recognition error can be measured using a variety of indicators, such as mean square error (MSE), cross entropy loss, etc.

[0135] For example, assuming that the true emotional state of the facial video training instance is "sad", and the inference emotion recognition result is "happy", the behavior analysis device calculates the recognition error according to certain rules. If the cross entropy loss is used as a measurement indicator, its formula is: , where y i is the true label, p i is the probability predicted by the network.

[0136] After calculating the recognition error, the behavior analysis device uses the back propagation algorithm to update the parameters of the video processing network to be trained, so that the recognition error continues to decrease. The back propagation algorithm is an optimization algorithm based on gradient descent. It calculates the gradient of the recognition error to the network parameters, and then updates the network parameters according to the direction of the gradient, so that the output of the network is closer to the true label. In the implementation process, for example, different optimization algorithms are used to update the network parameters, such as stochastic gradient descent (SGD), adaptive moment estimation (Adam), etc. Different optimization algorithms have different characteristics and applicable scenarios. The behavior analysis device needs to select a suitable optimization algorithm based on specific needs and data characteristics.

[0137] As an implementation manner, in step S201, before the feature representation of the training facial video segment is performed based on the video processing network to be trained and the training video feature array of the facial video training instance is obtained, a determination process of the video processing network to be trained is also included, including:

[0138] Step S20a: obtaining an initial video processing network, the initial video processing network including a timing processing component;

[0139] Step S20b: if the time series processing component includes an online time series processing component and an offline time series processing component, the initial video processing network is used as the video processing network to be trained;

[0140] Step S20c: If the initial video processing network is a pre-debugging network and the timing processing component is an offline timing processing component, clone the offline timing processing component, determine the cloned offline timing processing component as an online timing processing component, add it to the initial video processing network, and obtain the video processing network to be trained.

[0141] In step S20a, the behavior analysis device obtains an initial video processing network, which includes a timing processing component (CTC). The behavior analysis device obtains a network suitable for processing video data from an existing network model library or a preliminarily constructed network structure as the initial video processing network, which is a timing processing component capable of processing time sequence information in a video sequence. The video is composed of a series of frames arranged in chronological order, and the component can capture the dynamic changes and time dependencies between video frames. For example, the behavior analysis device obtains an initial video processing network based on a combination of a convolutional neural network (CNN) and a recurrent neural network (RNN), in which the RNN part is used as a timing processing component, which can process the video frame sequence and learn the time features between frames. The feasible network structures as timing processing components include long short-term memory networks (LSTM), gated recurrent units (GRU), etc., which can effectively process sequence data.

[0142] In step S20b, the behavior analysis device determines whether the timing processing component includes an online timing processing component and an offline timing processing component. If it does, the initial video processing network is used as the video processing network to be trained. The online timing processing component has instant processing capabilities and short delays, and can quickly process and analyze when receiving video frames; the offline timing processing component needs to be processed uniformly after receiving all facial images, and can use more context information for identification. When the behavior analysis device detects that the timing processing component of the initial video processing network has both types, it means that the network has the ability to work in different processing modes and can be directly used for subsequent training. For example, in the initial video processing network, a part of the LSTM network is configured as an online timing processing component, which can process newly input video frames in real time; another part of the LSTM network is configured as an offline timing processing component, which is used for comprehensive analysis after accumulating a certain number of video frames. At this time, the behavior analysis device determines this initial video processing network as the video processing network to be trained.

[0143] In step S20c, when the initial video processing network is a pre-debugging network (i.e., a pre-trained neural network) and the timing processing component is an offline timing processing component, the behavior analysis device clones the offline timing processing component, and determines the cloned offline timing processing component as an online timing processing component, adds it to the initial video processing network, and obtains a video processing network to be trained. The pre-debugging network is a network that has been pre-trained on other related tasks or large-scale data sets, and it has learned some common features and patterns. When the behavior analysis device finds that the initial video processing network is such a pre-trained network and there is only an offline timing processing component, in order to enable the network to have online processing capabilities, the offline timing processing component is cloned. The cloning operation is to copy a component with the same structure and parameters as the offline timing processing component. For example, the initial video processing network is a GRU-based network pre-trained on a large-scale image data set, in which the GRU is an offline timing processing component. The behavior analysis device copies an identical GRU component, sets it to the working mode of the online timing processing component, and then adds the cloned component to the initial video processing network. In this way, the initial video processing network has both online and offline time series processing capabilities, and the behavior analysis device determines this modified network as the video processing network to be trained.

[0144] As an implementation manner, the facial video training instance includes a plurality of training facial images. In step S201, a plurality of training facial video clips corresponding to the facial video training instance are obtained, including:

[0145] Step S2011: segmenting the facial video training instance based on a preset video window and a preset number of image blocks to obtain a plurality of training facial image sets;

[0146] Step S2012: selecting one or more historical training facial images corresponding to each training facial image set from the training facial images;

[0147] Step S2013: adding the historical training facial images to the corresponding training facial image sets to obtain a plurality of target training facial image sets, and using the target training facial image sets as training facial video clips.

[0148] In step S2011, the behavior analysis device divides the facial video training instance based on the preset video window and the preset number of image blocks to obtain multiple training facial image sets. The preset video window refers to a time range preset by the behavior analysis device for capturing video clips from the facial video training instance; the preset number of image blocks specifies the number of images contained in each training facial image set. The behavior analysis device sequentially captures video clips from the facial video training instance according to the time range of the preset video window, and then divides each video clip according to the preset number of image blocks to obtain multiple training facial image sets. For example, assuming that the duration of the facial video training instance is 60 seconds, the frame rate is 30 frames per second, the preset video window is 10 seconds, and the preset number of image blocks is 300 frames. The behavior analysis device sequentially captures video clips from 0-10 seconds, 10-20 seconds, 20-30 seconds, etc., each clip contains 300 frames of images, and each clip is used as a training facial image set, so that 6 training facial image sets can be obtained. When implementing this step, the time capture and image counting functions in the video processing library can be used to perform segmentation operations according to preset time and quantity.

[0149] In step S2012, the behavior analysis device selects one or more historical training facial images corresponding to each training facial image set from the training facial images. Historical training facial images refer to representative facial images that have been used before, and these images can provide more reference information and context for the current training. The behavior analysis device selects appropriate images from the historical training facial image library according to the characteristics and requirements of each training facial image set. For example, for a training facial image set containing angry facial images, the behavior analysis device selects some images that also have angry emotions and obvious expression features from the historical training facial image library as corresponding historical training facial images. The historical training facial images that meet the requirements can be quickly screened out by establishing image feature indexes and classification labels.

[0150] In step S2013, the behavior analysis device adds the historical training facial images to the corresponding training facial image set to obtain multiple target training facial image sets, and uses the target training facial image sets as training facial video clips. The behavior analysis device adds the historical training facial images selected in step S2012 to the corresponding training facial image set to form a new target training facial image set. These target training facial image sets contain current training facial images and historical training facial images with reference value, which can provide richer information for subsequent training. For example, the selected 5 historical training facial images are added to a training facial image set that originally contains 300 images to obtain a target training facial image set containing 305 images. Finally, the behavior analysis device uses these target training facial image sets as training facial video clips for subsequent feature representation and model training.

[0151] As an implementation mode, step S202, based on the training video feature array, inferring the emotional state corresponding to the facial video training instance to obtain the inferred emotion recognition result, includes:

[0152] Step S2021: inferring the emotional state of the facial video training instance based on the online time series processing component according to the training video feature array, obtaining an online inference emotion recognition result, and detecting and obtaining the training distribution position of one or more training facial image sequences in the facial video training instance;

[0153] Step S2022: segmenting the training video feature array based on the training distribution position to obtain a target training video feature array corresponding to each training facial image sequence;

[0154] Step S2023: Based on the target training video feature array, the emotional state of the facial video training instance is inferred based on the offline timing processing component to obtain the current inferred emotion recognition result, and the online inferred emotion recognition result and the current inferred emotion recognition result are used as the inferred emotion recognition result.

[0155] In step S2021, the behavior analysis device infers the emotional state of the facial video training instance based on the training video feature array and the online timing processing component to obtain the online inference emotion recognition result, and detects the training distribution position of one or more training facial image sequences in the facial video training instance. The online timing processing component can process the input video feature data in real time and has the characteristic of low latency. The behavior analysis device inputs the training video feature array into the online timing processing component, which infers the emotional state of the driver in the video in real time based on the learned patterns and rules, and outputs the online inference emotion recognition result, such as happiness, sadness, anger, etc. At the same time, the behavior analysis device searches for those training facial image sequences with emotion change characteristics in the facial video training instance, and determines their distribution positions in the video. These position information are very important for subsequent analysis. For example, in a 60-second facial video training instance, the online time series processing component determines based on the training video feature array that the driver is calm in the first 20 seconds, angry in 20-40 seconds, and calm again in 40-60 seconds, and at the same time determines the specific distribution position of the training facial image sequence corresponding to the angry emotion in 20-40 seconds in the video. A classifier can be set in the online time series processing component, the emotional state can be classified using a classification algorithm, and the distribution position of the training facial image sequence can be detected using feature matching and threshold judgment methods.

[0156] In step S2022, the behavior analysis device segments the training video feature array based on the training distribution position to obtain the target training video feature array corresponding to each training facial image sequence. The behavior analysis device divides the training video feature array according to the training distribution position of the training facial image sequence obtained in step S2021, extracts the feature data corresponding to each training facial image sequence, and forms an independent target training video feature array. The purpose of doing this is to be able to perform a more detailed analysis on each image sequence with specific emotional changes in the future. For example, if it is determined in step S2021 that there are two training facial image sequences in the video, corresponding to anger and happiness, respectively, and their distribution positions are clear, the behavior analysis device will extract the feature data corresponding to these two sequences from the training video feature array to form target training video feature arrays respectively. The training video feature array can be segmented by indexing and slicing.

[0157] In step S2023, the behavior analysis device infers the emotional state of the facial video training instance based on the target training video feature array and the offline time series processing component to obtain the current inference emotion recognition result, and uses the online inference emotion recognition result and the current inference emotion recognition result as the inference emotion recognition result. The offline time series processing component can perform unified processing after obtaining all relevant data, and can use more context information. The behavior analysis device inputs the target training video feature array obtained in step S2022 into the offline time series processing component, which will comprehensively consider the feature information of the entire sequence, make a more accurate inference of the emotional state, and obtain the current inference emotion recognition result. Finally, the behavior analysis device combines the online inference emotion recognition result and the current inference emotion recognition result to form the final inference emotion recognition result. For example, the online inference emotion recognition result shows that the driver may be in an angry mood in a certain time period, and the offline time series processing component further analyzes the target training video feature array and confirms that the driver is indeed in an angry mood in this time period. At the same time, some emotional details that are not recognized by the online inference may also be found. After integrating these two results, a more comprehensive and accurate inference emotion recognition result is obtained. More complex models and algorithms, such as long short-term memory networks (LSTMs), can be used in offline time series processing components to reason about emotional states.

[0158] As an implementation manner, in step S2021, detecting and obtaining training distribution positions of one or more training facial image sequences in the facial video training instance includes:

[0159] Step S20211: Based on the training video feature array, one or more training core facial images corresponding to training emotion fluctuation nodes are detected in the facial video training instance;

[0160] Step S20212: selecting a target division number corresponding to the training emotion fluctuation node from the set integer set, and dividing the training emotion fluctuation node according to the target division number to obtain one or more training emotion fluctuation node sets;

[0161] Step S20213: Based on the training core facial image, the training distribution position of the training facial image sequence corresponding to each training emotion fluctuation node set is detected in the facial video training instance.

[0162] In step S20211, the behavior analysis device detects and obtains training core facial images corresponding to one or more training emotion fluctuation nodes in the facial video training instance based on the training video feature array. The training video feature array contains the feature information of each frame image in the facial video training instance. The behavior analysis device analyzes these features to find out those nodes that can represent obvious changes in emotions, namely, training emotion fluctuation nodes. The facial images corresponding to these nodes are the training core facial images, which usually have obvious changes in emotional characteristics. For example, in a facial video training instance, the driver originally had a calm expression, but suddenly showed an angry expression. Then the moment of expression change is a training emotion fluctuation node, and the corresponding facial image at this time is the training core facial image. The training emotion fluctuation node can be detected by setting a threshold for feature change. When some feature values ​​in the feature array change by more than the threshold, it is considered that an emotion fluctuation node has appeared.

[0163] In step S20212, the behavior analysis device selects the target number of divisions corresponding to the training emotion fluctuation node from the set integer set, and divides the training emotion fluctuation node according to the target number of divisions to obtain one or more training emotion fluctuation node sets. The set integer set is a set of integers pre-set by the behavior analysis device for selecting a suitable number of divisions. The behavior analysis device selects a suitable integer from the set integer set as the target number of divisions according to the number and distribution of the training emotion fluctuation nodes. Then, the training emotion fluctuation nodes are grouped according to the target number of divisions to form multiple training emotion fluctuation node sets. For example, the integer set is set to {2, 3, 4}, and the behavior analysis device detects that there are 6 training emotion fluctuation nodes in the facial video training instance. After evaluation, the target number of divisions is selected as 3, then the 6 training emotion fluctuation nodes will be divided into 3 training emotion fluctuation node sets, each set containing 2 nodes. The appropriate target number of divisions can be determined by calculating the distance between nodes and cluster analysis.

[0164] In step S20213, the behavior analysis device detects the training distribution position of the training facial image sequence corresponding to each training emotion fluctuation node set in the facial video training instance based on the training core facial image. The training core facial images represent the key nodes of emotion fluctuations. The behavior analysis device uses these images as a reference to determine the range of the facial image sequence corresponding to each training emotion fluctuation node set in the facial video training instance, that is, the training distribution position. For example, for a set containing 3 training emotion fluctuation nodes, the behavior analysis device finds the facial image sequence starting from the training core facial image corresponding to the first node and ending with the training core facial image corresponding to the last node, and determines its specific position in the video. The training distribution position can be determined by the frame number and timestamp of the image.

[0165] In step S20212, when selecting the target number of partitions from the set integer set, the distribution density of the training emotion fluctuation nodes and the characteristics of the data are comprehensively considered. If the nodes are densely distributed, a larger number of partitions can be selected to analyze the emotion changes more carefully; if the nodes are sparsely distributed, a smaller number of partitions can be selected. Cluster analysis can help the behavior analysis device understand the relationship and distribution pattern between nodes, so as to select a suitable target number of partitions. A feasible clustering algorithm is the K-means clustering algorithm, whose basic idea is to divide the data points into K clusters, so that the similarity of the data points within the cluster is high, and the similarity of the data points between clusters is low.

[0166] In step S20213, when determining the training distribution position of the training facial image sequence, the accuracy of the position information must be ensured. The start and end positions of the image sequence can be accurately calculated by accurately positioning and tracking the training core facial image and combining the frame rate and time information of the video. At the same time, the noise and interference factors that may exist in the video must be taken into account, and the position information must be appropriately corrected and adjusted.

[0167] As an implementation mode, step S2023, based on the target training video feature array, inferring the emotional state of the facial video training instance based on the offline time series processing component, and obtaining the current inference emotion recognition result includes:

[0168] Step S20231: extracting an emotion feature array from the target training video feature array;

[0169] Step S20232: According to the emotion feature array, the emotion state of the facial video training instance is inferred based on the offline time series processing component to obtain an offline inference emotion recognition result;

[0170] Step S20233: fusing the emotion feature arrays to obtain the current emotion feature array corresponding to each training facial image sequence, and determining the focus influence coefficient of the current emotion feature array;

[0171] Step S20234: Determine the emotional state corresponding to the current emotional feature array based on the focus influence coefficient, obtain the collaborative reasoning emotion recognition result, and use the collaborative reasoning emotion recognition result and the offline reasoning emotion recognition result as the current reasoning emotion recognition result.

[0172] In step S20231, the behavior analysis device extracts an emotional feature array from the target training video feature array. The target training video feature array contains various feature information of the training facial image sequence, but not all features are directly related to the emotional state. The behavior analysis device needs to filter out features that are closely related to emotions to form an emotional feature array. For example, in facial expression recognition, features such as the degree of opening and closing of the eyes, the angle of the corners of the mouth, etc. are often closely related to the emotional state. The behavior analysis device extracts these features from the target training video feature array to form an emotional feature array. Feature selection algorithms, such as chi-square test, information gain, etc., can be used to screen according to the correlation between features and emotional labels, and select features with higher correlation as elements of the emotional feature array.

[0173] In step S20232, the behavior analysis device infers the emotional state of the facial video training instance based on the emotional feature array and the offline time series processing component to obtain the offline inference emotion recognition result. The offline time series processing component can perform unified processing after obtaining all relevant data, making full use of the context information of the training facial image sequence. The behavior analysis device inputs the emotional feature array into the offline time series processing component, which infers the emotional state of the driver in the video based on the learned patterns and rules, and outputs the offline inference emotion recognition result, such as happiness, sadness, anger, etc. For example, the offline time series processing component can be a trained long short-term memory network (LSTM), which can process sequence data and capture the temporal changes of emotions. LSTM calculates the most likely emotional state based on the input emotional feature array and its own weight parameters.

[0174] In step S20233, the behavior analysis device fuses the emotion feature array to obtain the current emotion feature array corresponding to each training facial image sequence, and determines the focus influence coefficient of the current emotion feature array. The behavior analysis device integrates the various features in the emotion feature array to form a more representative current emotion feature array. At the same time, in order to highlight the influence of certain important features on emotion recognition, the behavior analysis device assigns a focus influence coefficient, also known as attention weight, to each feature in the current emotion feature array. For example, the emotion feature array can be fused by a weighted average method, and each feature is multiplied by its corresponding weight and then added to obtain the current emotion feature array. The focus influence coefficient can be determined by an attention mechanism, which can automatically learn the importance of each feature. For example, in a neural network based on an attention mechanism, the focus influence coefficient of each feature is obtained by calculating the score of each feature and then normalizing the score through a softmax function.

[0175] In step S20234, the behavior analysis device determines the emotional state corresponding to the current emotional feature array based on the focus influence coefficient, obtains the collaborative reasoning emotion recognition result, and uses the collaborative reasoning emotion recognition result and the offline reasoning emotion recognition result as the current reasoning emotion recognition result. The behavior analysis device performs weighted processing on the current emotional feature array according to the focus influence coefficient to highlight the role of important features, and then infers the emotional state based on the weighted feature array to obtain the collaborative reasoning emotion recognition result. Finally, the behavior analysis device combines the collaborative reasoning emotion recognition result and the offline reasoning emotion recognition result to form the current reasoning emotion recognition result. For example, assuming that the offline reasoning emotion recognition result shows that the driver may be in an angry mood, and the collaborative reasoning emotion recognition result also supports this judgment, then the current reasoning emotion recognition result can more accurately determine that the driver is in an angry mood by combining these two results.

[0176] As an implementation mode, step S203, determining the recognition error of the facial video training instance based on the inference emotion recognition result, includes:

[0177] Step S2031: obtaining the emotion supervision mark of the facial video training instance, and performing error calculation between the emotion supervision mark and the online reasoning emotion recognition result to obtain the online reasoning error;

[0178] Step S2032: performing error calculation between the emotion supervision mark and the offline reasoning emotion recognition result to obtain the offline reasoning error, and performing error calculation between the emotion supervision mark and the collaborative reasoning emotion recognition result to obtain the collaborative reasoning error;

[0179] Step S2033: The online reasoning error, the offline reasoning error and the collaborative reasoning error are combined to obtain the recognition error of the facial video training instance.

[0180] In step S2031, the behavior analysis device obtains the emotion supervision label of the facial video training instance, and performs error calculation between the emotion supervision label and the online reasoning emotion recognition result to obtain the online reasoning error. The emotion supervision label is a label of the driver's real emotional state in the facial video training instance, which is usually determined by manual annotation or other reliable methods. The behavior analysis device compares this real emotion label with the online reasoning emotion recognition result obtained by reasoning the online timing processing component, calculates the difference between them, and obtains the online reasoning error. For example, assuming that the emotion supervision label of the facial video training instance shows that the driver is in a "happy" emotion, and the online reasoning emotion recognition result is judged to be "sad", the behavior analysis device calculates the error value between the two according to a certain error calculation method, such as the cross entropy loss function.

[0181] In step S2032, the behavior analysis device calculates the error between the emotion supervision mark and the offline reasoning emotion recognition result to obtain the offline reasoning error, and calculates the error between the emotion supervision mark and the collaborative reasoning emotion recognition result to obtain the collaborative reasoning error. The behavior analysis device also uses a method similar to step S2031 to compare the emotion supervision mark with the offline reasoning emotion recognition result obtained by the offline timing processing component reasoning and the collaborative reasoning emotion recognition result obtained by the focus influence coefficient reasoning, and calculates the error between them. For example, if the emotion supervision mark is "angry", the offline reasoning emotion recognition result is "calm", and the collaborative reasoning emotion recognition result is "slightly angry", the behavior analysis device uses appropriate error calculation methods, such as mean square error (MSE) or cross entropy loss function, to calculate the offline reasoning error and the collaborative reasoning error.

[0182] In step S2033, the behavior analysis device combines the online reasoning error, the offline reasoning error and the collaborative reasoning error to obtain the recognition error of the facial video training instance. The behavior analysis device needs to comprehensively process the three errors calculated above to obtain a recognition error that can fully reflect the difference between the reasoning result and the actual situation. This recognition error will be used to adjust the parameters of the video processing network to be trained later so that the network can be continuously optimized. For example, a weighted average method is used to assign different weights to the three errors according to their importance, and then the weighted errors are added to obtain the final recognition error.

[0183] As an implementation mode, step S2033, combining the online reasoning error, the offline reasoning error and the collaborative reasoning error to obtain the recognition error of the facial video training instance, includes:

[0184] Step S20331: obtaining the error merging influence coefficient, and adjusting the online reasoning error, the offline reasoning error and the collaborative reasoning error respectively according to the error merging influence coefficient;

[0185] Step S20332: Calculate the average of the adjusted online reasoning error and the adjusted offline reasoning error to obtain a target reasoning error;

[0186] Step S20333: Merge the target reasoning error and the adjusted collaborative reasoning error to obtain the recognition error of the facial video training instance.

[0187] In step S20331, the behavior analysis device obtains the error merging influence coefficient, and adjusts the online reasoning error, offline reasoning error and collaborative reasoning error respectively according to the error merging influence coefficient. The error merging influence coefficient is a set of pre-set weight parameters used to measure the proportion of different reasoning errors in the final recognition error. The behavior analysis device obtains this set of coefficients, and then multiplies each reasoning error by the corresponding influence coefficient to achieve error adjustment.

[0188] In step S20332, the behavior analysis device calculates the average of the adjusted online reasoning error and the adjusted offline reasoning error to obtain the target reasoning error. The behavior analysis device adds the adjusted online reasoning error and the adjusted offline reasoning error obtained in step S20331, and then divides them by 2 to obtain the target reasoning error. The purpose of this step is to comprehensively consider the results of online reasoning and offline reasoning to obtain a relatively balanced intermediate error value.

[0189] In step S20333, the behavior analysis device combines the target reasoning error with the adjusted collaborative reasoning error to obtain the recognition error of the facial video training instance. The behavior analysis device adds the target reasoning error obtained in step S20332 and the adjusted collaborative reasoning error obtained in step S20331 to obtain the final recognition error of the facial video training instance.

[0190] In step S20331, the optimal error merging influence coefficient can be determined by a cross-validation method. Specifically, the data set can be divided into multiple subsets, different coefficient combinations can be tried on different subsets, and the coefficient combination that can minimize the recognition error is selected as the final error merging influence coefficient. In addition, a machine learning algorithm can also be used to automatically learn these coefficients. For example, a neural network can be used to learn the error merging influence coefficient, with the online reasoning error, the offline reasoning error, and the collaborative reasoning error as input, and the final recognition error as output. The coefficient is adjusted by training the neural network to minimize the output recognition error.

[0191] The embodiment of the present invention provides a behavior analysis device, such as Figure 2As shown, the behavior analysis device 100 includes: a processor 101 and a memory 103. Among them, the processor 101 and the memory 103 are connected, such as being connected through a bus 102. Optionally, the behavior analysis device 100 may also include a transceiver 104. It should be noted that in actual applications, the transceiver 104 is not limited to one, and the structure of the behavior analysis device 100 does not constitute a limitation on the embodiment of the present invention. The memory 103 is used to store the application code for executing the scheme of the present invention, and is controlled by the processor 101 to execute. The processor 101 is used to execute the application code stored in the memory 103 to implement the content shown in any of the aforementioned method embodiments. That is, the embodiment of the present invention provides a behavior analysis device, and the behavior analysis device in the embodiment of the present invention includes: one or more processors; a memory; one or more computer programs, wherein one or more computer programs are stored in the memory and are configured to be executed by one or more processors, and when one or more programs are executed by the processor, the driving behavior analysis method based on facial emotion recognition provided by the present invention is implemented.

[0192] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.

[0193] The above descriptions are only some embodiments of the present invention. It should be pointed out that, for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A driving behavior analysis method based on facial emotion recognition, characterized in that: include: Acquire a real-time driver's facial video, and segment the real-time driver's facial video to obtain a plurality of target facial video segments; Performing feature representation on the target facial video clip to obtain a video feature array of the real-time driver facial video, and generating an initial emotion recognition result corresponding to the real-time driver facial video based on the video feature array; Detecting a plurality of core facial images in the real-time driver facial video, and detecting distribution positions of a plurality of facial image sequences in the real-time driver facial video based on the core facial images; Based on the distribution positions, segmenting the video feature array to obtain a target video feature array for each facial image sequence; Iterating the initial emotion recognition result according to the target video feature array to obtain a target emotion recognition result corresponding to the real-time driver facial video; Evaluating the safety of the driver's driving behavior based on the target emotion recognition result; The driver's facial video includes a plurality of facial images, and the plurality of core facial images detected in the real-time driver's facial video include: Inferring the prediction confidence of the emotional state corresponding to each facial image based on the video feature array; Based on the prediction confidence, detecting and obtaining node distribution intervals of one or more emotion fluctuation nodes in the real-time driver facial video; Selecting multiple facial images corresponding to the node distribution interval from the facial images to obtain a core facial image corresponding to each emotion fluctuation node; The detecting and obtaining distribution positions of a plurality of facial image sequences in the real-time driver facial video based on the core facial image includes: Dividing the emotion fluctuation nodes based on the node distribution interval to obtain a plurality of emotion fluctuation node sets, wherein the emotion fluctuation node sets include a set number of emotion fluctuation nodes; According to the core facial image, an image range of a facial image sequence corresponding to each set of emotion fluctuation nodes is detected in the real-time driver facial video; The image range is augmented to obtain a target image range, and the target image range is used as a distribution position of the facial image sequence.

2. The driving behavior analysis method based on facial emotion recognition as claimed in claim 1, characterized in that: The detecting, based on the core facial image, in the real-time driver facial video, an image range of a facial image sequence corresponding to each set of emotion fluctuation nodes comprises: Based on the node distribution interval, selecting a triggering emotion fluctuation node and a terminating emotion fluctuation node from the emotion fluctuation node set; Detecting the number of core facial images corresponding to the triggering emotion fluctuation node in the real-time driver's facial video to obtain the number of triggering images; Detecting the number of core facial images corresponding to the termination emotion fluctuation node in the real-time driver's facial video to obtain the number of termination images; Using the range formed by the number of trigger images and the number of termination images as the image range of the facial image sequence corresponding to the emotion fluctuation node set; The step of augmenting the image range to obtain a target image range includes: Subtracting the number of trigger images from the first set number of images to obtain a target number of trigger images; Adding the number of termination images to the second set number of images to obtain a target number of termination images; The range formed by the number of target trigger images and the number of target termination images is used as the target image range; The video feature array includes a sub-video feature array corresponding to each facial image, and the video feature array is segmented based on the distribution position to obtain a target video feature array for each facial image sequence, including: Selecting multiple facial images corresponding to the target image range from the facial images to obtain a candidate facial image set corresponding to each facial image sequence; Extracting a sub-video feature array corresponding to each facial image in the candidate facial image set from the video feature array, and obtaining a sub-video feature array set corresponding to each facial image sequence; The sub-video feature array set is determined as a target video feature array corresponding to the facial image sequence.

3. The driving behavior analysis method based on facial emotion recognition according to claim 1 or 2, characterized in that: The step of iterating the initial emotion recognition result according to the target video feature array to obtain the target emotion recognition result corresponding to the real-time driver facial video includes: Determining a target facial emotion corresponding to each facial image sequence according to the target video feature array; Based on the distribution positions, the target facial emotions are fused to obtain fused facial emotions; Changing the initial emotion recognition result to the fused facial emotion to obtain a target emotion recognition result corresponding to the real-time driver facial video; The step of performing feature representation on the target facial video segment to obtain a video feature array of the real-time driver facial video includes: Performing feature representation on the target facial video clips based on a video processing network to obtain an initial video feature array corresponding to each target facial video clip; The initial video feature arrays are fused to obtain a video feature array corresponding to the real-time driver's facial video.

4. The driving behavior analysis method based on facial emotion recognition as claimed in claim 3, characterized in that: Before the target facial video segments are characterized based on the video processing network to obtain an initial video feature array corresponding to each target facial video segment, the method further includes: Acquire multiple training facial video clips corresponding to the facial video training instance, and perform feature representation on the training facial video clips based on the video processing network to be trained to obtain a training video feature array of the facial video training instance; Inferring the emotional state corresponding to the facial video training instance based on the training video feature array to obtain an inferred emotion recognition result; Based on the inferred emotion recognition result, the recognition error of the facial video training instance is determined, and according to the recognition error, the parameters of the video processing network to be trained are adjusted to obtain the video processing network.

5. The driving behavior analysis method based on facial emotion recognition as claimed in claim 4, characterized in that: Before the feature representation of the training facial video segment based on the video processing network to be trained is performed to obtain the training video feature array of the facial video training instance, a determination process of the video processing network to be trained is also included, including: Obtaining an initial video processing network, wherein the initial video processing network includes a timing processing component; If the time series processing component includes an online time series processing component and an offline time series processing component, the initial video processing network is used as the video processing network to be trained; If the initial video processing network is a pre-debugging network and the timing processing component is the offline timing processing component, the offline timing processing component is cloned, and the cloned offline timing processing component is determined as an online timing processing component, and added to the initial video processing network to obtain a video processing network to be trained; The step of inferring the emotional state corresponding to the facial video training instance based on the training video feature array to obtain the inferred emotion recognition result includes: According to the training video feature array, based on the online time series processing component, the emotional state of the facial video training instance is inferred to obtain an online inference emotion recognition result, and the training distribution position of one or more training facial image sequences is detected in the facial video training instance; Based on the training distribution position, the training video feature array is segmented to obtain a target training video feature array corresponding to each training facial image sequence; According to the target training video feature array, the emotional state of the facial video training instance is inferred based on the offline timing processing component to obtain a current inferred emotion recognition result, and the online inferred emotion recognition result and the current inferred emotion recognition result are used as the inferred emotion recognition result.

6. The driving behavior analysis method based on facial emotion recognition as claimed in claim 5, characterized in that: The detecting and obtaining the training distribution positions of one or more training facial image sequences in the facial video training instance comprises: Based on the training video feature array, detecting and obtaining training core facial images corresponding to one or more training emotion fluctuation nodes in the facial video training instance; Selecting a target number of divisions corresponding to the training emotion fluctuation node from a set integer set, and dividing the training emotion fluctuation node according to the target number of divisions to obtain one or more training emotion fluctuation node sets; Based on the training core facial image, detecting in the facial video training instance the training distribution position of the training facial image sequence corresponding to each training emotion fluctuation node set; The step of inferring the emotional state of the facial video training instance based on the target training video feature array and the offline time series processing component to obtain the current inference emotion recognition result includes: Extracting an emotion feature array from the target training video feature array; According to the emotion feature array, based on the offline time series processing component, the emotional state of the facial video training instance is inferred to obtain an offline inference emotion recognition result; The emotion feature arrays are merged to obtain a current emotion feature array corresponding to each training facial image sequence, and a focus influence coefficient of the current emotion feature array is determined; According to the focus influence coefficient, the emotional state corresponding to the current emotional feature array is determined to obtain a collaborative reasoning emotion recognition result, and the collaborative reasoning emotion recognition result and the offline reasoning emotion recognition result are used as the current reasoning emotion recognition result.

7. The driving behavior analysis method based on facial emotion recognition as claimed in claim 6, characterized in that: The determining, based on the inferred emotion recognition result, a recognition error of the facial video training instance comprises: Obtaining an emotion supervision tag of the facial video training instance, and performing error calculation between the emotion supervision tag and the online reasoning emotion recognition result to obtain an online reasoning error; Performing error calculation between the emotion supervision mark and the offline reasoning emotion recognition result to obtain the offline reasoning error, and performing error calculation between the emotion supervision mark and the collaborative reasoning emotion recognition result to obtain the collaborative reasoning error; Obtaining an error merging influence coefficient, and adjusting the online reasoning error, the offline reasoning error, and the collaborative reasoning error respectively according to the error merging influence coefficient; Calculate the average of the adjusted online reasoning error and the adjusted offline reasoning error to obtain the target reasoning error; The target reasoning error is combined with the adjusted collaborative reasoning error to obtain the recognition error of the facial video training instance.

8. The driving behavior analysis method based on facial emotion recognition as claimed in claim 4, characterized in that: The facial video training instance includes a plurality of training facial images, and the step of obtaining a plurality of training facial video clips corresponding to the facial video training instance includes: Segmenting the facial video training instance based on a preset video window and a preset number of image blocks to obtain a plurality of training facial image sets; Selecting one or more historical training facial images corresponding to each training facial image set from the training facial images; The historical training facial images are added to corresponding training facial image sets to obtain a plurality of target training facial image sets, and the target training facial image sets are used as training facial video clips.

9. A behavior analysis device, characterized in that: include: one or more processors; Memory; one or more computer programs; The one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more computer programs are executed by the processors, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Facial expression analysis method and system and satisfaction analysis method and system

    CN113111690A

  • Driver emotion driving behavior recognition device and method and storage medium

    CN118468079A