Robot target person following method

By constructing a dynamic pseudo-label mechanism and a kernel entropy component analysis framework, and combining facial video and physiological behavior data, the problem of insufficient accuracy and real-time performance in fatigue state recognition in existing technologies has been solved, and efficient tracking of the fatigue state of target personnel has been achieved.

CN120913259AInactive Publication Date: 2025-11-07BOYUAN YOUCHEN TECHNOLOGY (BEIJING) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510998717.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-11-07
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively distinguish and process large amounts of target personnel data, especially during ID switching scenarios, when identifying and tracking fatigue. This leads to tracking loss, and traditional methods rely on standard fatigue data, failing to comprehensively capture facial micro-expressions and physiological behavioral changes.

Method used

A dynamic pseudo-labeling mechanism and kernel entropy component analysis framework are constructed. By collecting facial video and physiological behavior data, expression entropy features and physiological behavior features are extracted. Kernel entropy component analysis is used for nonlinear dimensionality reduction and mutual information filtering. Combined with entropy optimization neural network for fatigue recognition, multi-dimensional feature analysis is achieved.

Benefits of technology

It improves the accuracy and real-time performance of fatigue state tracking, enabling more comprehensive capture of facial micro-expressions and physiological behavioral changes, thereby enhancing the accuracy and real-time performance of fatigue state assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913259A_ABST
    Figure CN120913259A_ABST
Patent Text Reader

Abstract

The invention discloses a robot target person following method, which comprises the steps of collecting to-be-analyzed face video data and working behavior data of a target person, and dividing the to-be-analyzed face video data into image data in different time periods according to a pseudo tag, taking a time period in which to-be-analyzed behavior data and standard behavior data of the working behavior data are different as problem data, and performing nonlinear dimensionality reduction on the time difference feature and the behavior difference feature by adopting kernel entropy component analysis to obtain a principal component and obtain a fatigue correlation feature; and inputting the fatigue correlation characteristics into an entropy optimization neural network for training to obtain a fatigue recognition model, outputting the fatigue level of the target person by using the fatigue recognition model, and tracking and following the recognized fatigue person. According to the method, by constructing a dynamic pseudo-tag mechanism and a kernel entropy component analysis framework, collaborative optimization of facial expression entropy features and physiological behavior features is realized, and the accuracy and real-time performance of fatigue state tracking and following are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of following, and particularly relates to a robot target personnel following method. BACKGROUND

[0002] With the improvement of the automation and intelligentization level of modern industry, the fatigue problem of personnel in a high-intensity and high-repetitive work environment for a long time is increasingly prominent, the invention patent with the publication number CN202310610882 discloses a "fatigue recognition method, device and equipment and storage medium of human face", which constructs a dynamic Bayesian network of a target person based on fatigue environment quantitative data and fatigue observation quantitative data, calculates the conditional fatigue probability of different nodes in the dynamic Bayesian network to the target person, and obtains the fatigue recognition result of the target person based on the conditional fatigue probability, mainly based on environment data and observation data to calculate the fatigue probability of relevant personnel, research shows that the fatigue state will significantly reduce the reaction speed, attention concentration and decision-making ability of personnel, and become an important hidden danger of causing safety accidents and reducing production efficiency, so it is necessary to follow the monitoring of personnel, but the traditional method depends on the obtained standard fatigue data to determine and track the target personnel, and a large amount of target personnel data cannot be distinguished and processed in time, which may lead to tracking loss in the scene of ID switching.

[0003] In view of the above problems, the application provides a robot target personnel following method, which realizes the cooperative optimization of facial expression entropy features and physiological behavior features by constructing a dynamic pseudo-label mechanism and a kernel entropy component analysis framework, and improves the accuracy and real-time performance of fatigue state tracking and following. SUMMARY

[0004] The application aims to provide a robot target personnel following method.

[0005] To achieve the above-mentioned purpose, the application is implemented according to the following technical scheme:

[0006] The application provides a robot target personnel following method in the first aspect, which comprises:

[0007] The target personnel's to-be-analyzed facial video data and work behavior data are collected for preprocessing, and the work behavior data comprises to-be-analyzed behavior data and standard behavior data;

[0008] Expression entropy features are extracted based on the to-be-analyzed facial video data, pseudo labels are obtained by performing density clustering on the expression entropy features, and the to-be-analyzed facial video data is divided into image data of different time periods according to the pseudo labels;

[0009] extracting the work behavior data corresponding to the time period according to the image data, taking the time period in which the to-be-analyzed behavior data and the standard behavior data of the work behavior data are different as problem data;

[0010] extracting time difference features of the eye region of interest and the perioral region of interest according to the problem data, and extracting behavior difference features according to the to-be-analyzed behavior data corresponding to the time difference features;

[0011] performing nonlinear dimension reduction on the time difference features and the behavior difference features by kernel entropy component analysis to obtain principal components, performing mutual information screening using the principal components, and obtaining fatigue correlation features;

[0012] inputting the fatigue correlation features into an entropy optimization neural network for training to obtain a fatigue recognition model, outputting a fatigue level of the target person using the fatigue recognition model, and tracking and following the recognized fatigue person.

[0013] As a further method, the preprocessing method comprises:

[0014] collecting to-be-analyzed facial video data of a target person, performing image enhancement, denoising, uniform image size and posture on the to-be-analyzed facial video data, collecting heart rate data and skin electric response data of the target person as to-be-analyzed behavior data, collecting heart rate data and skin electric response data of the target person in a normal state as standard behavior data, cleaning and deduplicating the to-be-analyzed behavior data and the standard behavior data, and obtaining work behavior data.

[0015] As a further method, the method for extracting expression entropy features based on the to-be-analyzed facial video data comprises:

[0016] obtaining a first sequence by equal-interval division based on the to-be-analyzed facial video data, obtaining a second sequence by equal-interval division based on the first sequence, extracting a facial action unit by a convolutional neural network model based on the second sequence, and generating an action unit data sequence based on an activation state of the facial action unit in an image frame of the second sequence;

[0017] counting the number of occurrences of different action units according to the action unit data sequence, dividing the occurrence frequency by the total number of image frames to obtain a probability value, and summarizing the probability values of the action units of the second sequence to form a probability distribution set, and calculating the first sequence expression entropy features based on the probability distribution set of each second sequence in the first sequence using an expression entropy formula, wherein the expression entropy formula is:

[0018]

[0019] where H is the first sequence index, t is the second sequence index, T is the total number of second sequences, i is the action unit index, n is the total number of action units, p i (t) is the probability of the i-th action unit appearing in the t-th second sequence, is the variance of the static entropy of the second sequence, is the static entropy difference value of adjacent second sequences.

[0020] As a further method, the expression entropy feature is subjected to density clustering to obtain a pseudo-labeling method, comprising:

[0021] The expression entropy feature is subjected to Z-score standardization to obtain an expression entropy feature dataset, and the Euclidean distance between data points is calculated according to the expression entropy feature dataset, and the average of the Euclidean distances of the 7 adjacent data points is taken as the neighborhood radius of hierarchical density clustering;

[0022] The square root integer value of the data amount of the expression entropy feature dataset is used as the minimum neighborhood sample number of hierarchical density clustering, and the expression entropy feature dataset is subjected to hierarchical density clustering based on the neighborhood radius and the minimum neighborhood sample number, and if the number of data points in the neighborhood radius is greater than or equal to the minimum neighborhood sample number, the region is taken as a clustering core, and the data points in the neighborhood radius are included in the same cluster according to the clustering core, until the clustering cluster is obtained, and the data points in the clustering cluster are sequentially assigned numerical labels from more to less according to the number of data points, to obtain a pseudo-labeling.

[0023] As a further method, the method of taking the time period in which the behavior data to be analyzed and the standard behavior data differ as problem data, comprising:

[0024] Based on the behavior data to be analyzed and the standard behavior data, the difference degree is calculated, and the difference degree formula is:

[0025]

[0026] where S is the difference degree of the behavior data to be analyzed and the standard behavior data, i is the summation index variable, j is the summation index variable, n is the number of behavior parameters, m is the number of standard behavior data samples, x i is the behavior data value of the i-th parameter to be analyzed, x ik is the k-th standard behavior data value of the i-th parameter, x i is the mean value of the standard behavior data of the i-th parameter, x jk is the k-th standard behavior data value of the j-th parameter, x i is the behavior data value of the i-th parameter to be analyzed, x i is the standard behavior data value of the i-th parameter, and ∈ is a very small constant to prevent division by zero error;

[0027] The difference mean in the behavior data to be analyzed and the sum of 3 times the difference standard deviation are used as a problem threshold, and the behavior data to be analyzed with a difference greater than the problem threshold in different time periods are taken as problem data.

[0028] As a further method, the method for obtaining the time difference feature comprises:

[0029] According to the image data in the problem data, the eye region of interest and the perioral region of interest are extracted by a convolutional neural network model, the pupil diameter is extracted based on the eye region of interest, and the distance between the left and right corners of the mouth is extracted based on the perioral region of interest.

[0030] According to the eye region of interest and the perioral region of interest, the motion trajectory of the eyelid and the motion trajectory of the lip midpoint within two seconds are obtained by the optical flow method, and the eye opening degree is obtained by using the average distance between the upper eyelid and the lower eyelid divided by the pupil diameter according to the motion trajectory of the eyelid.

[0031] The eye opening degree is integrated in time sequence to obtain an eye opening degree time sequence, and the mouth opening degree is obtained by using the distance between the upper lip midpoint and the lower lip midpoint divided by the distance between the left and right corners of the mouth according to the motion trajectory of the lip midpoint. The mouth opening degree is integrated in time sequence to obtain a mouth opening degree time sequence.

[0032] The difference between two data points and the average change rate within 5 data points are calculated based on the data points of the eye opening degree time sequence in sequence, the difference between two data points and the average change rate within 5 data points are calculated based on the data points of the mouth opening degree time sequence in sequence, and the data point difference of the eye opening degree time sequence, the data point average change rate of the eye opening degree time sequence, the data point difference of the mouth opening degree time sequence, and the data point average change rate of the mouth opening degree time sequence are sequentially spliced to obtain the time difference feature.

[0033] As a further method, the method for obtaining the behavior difference feature comprises:

[0034] The behavior data to be analyzed corresponding to the time point is obtained based on the time difference feature, the difference between two adjacent data points and the average change rate within 5 consecutive data points are calculated in sequence according to the heart rate data of the behavior data to be analyzed, the difference between two adjacent data points and the average change rate within 5 consecutive data points are calculated in sequence according to the galvanic skin response data of the behavior data to be analyzed, and the heart rate data difference, the heart rate data average change rate, the galvanic skin response data difference, and the galvanic skin response data average change rate are sequentially spliced to obtain the behavior difference feature.

[0035] As a further method, the method for obtaining the fatigue correlation feature comprises:

[0036] The time difference feature and the behavior difference feature are spliced to obtain a row vector by time axis alignment based on the time difference feature and the behavior difference feature, the row vector is superimposed to obtain a feature matrix, and a kernel matrix is calculated from the feature matrix by using a Gaussian kernel function;

[0037] The kernel matrix is subjected to centering processing and eigenvalue decomposition to obtain eigenvalues and eigenvectors, the feature matrix is subjected to sparsity analysis to obtain a regularization parameter, the contribution degree is calculated according to the eigenvalues, the eigenvectors and the regularization parameter, and the contribution degree formula is:

[0038]

[0039] wherein S(k ′ ) is a cumulative contribution score of the first k ′ principal components, λ i is the i th eigenvalue of the kernel matrix, e i,k is the k th element of the i th eigenvector, e j,m is the m th element of the j th eigenvector, n is the number of data samples, m is an index, α is a regularization parameter, β is a regularization parameter, and k ′ is the number of principal components currently considered;

[0040] If the cumulative contribution score is greater than or equal to 0.85, the first k ′ principal components are taken as candidate features, the mutual information of the candidate features and the pseudo label is calculated, the candidate features with mutual information higher than the median are retained, and the fatigue correlation features are obtained.

[0041] As a further method, the method for tracking and following the identified fatigue personnel comprises:

[0042] The training set, the validation set and the test set are obtained based on the fatigue correlation features, the training set is input into an entropy optimization neural network, the difference between the predicted results and the pseudo label is calculated by using a cross-entropy loss function during training, the validation set is input into the model, the overfitting state is verified according to the loss function value and the accuracy rate change, the training is ended if the accuracy rate of the validation set is improved by less than 0.7% for 5 consecutive rounds, the difference is fed back to the layers of the model to adjust the parameters by using a back propagation algorithm, the test set is input into the trained entropy optimization neural network, the model performance is evaluated by using the accuracy rate and the confusion matrix, a fatigue recognition model is obtained, and the fatigue level of the target personnel is output based on the fatigue recognition model;

[0043] The identified fatigue personnel is given a tracking identifier, the ReID features of the fatigue personnel are extracted according to the tracking identifier, the cosine similarity is calculated according to the ReID features of different frames in the face video data to be analyzed, the ReID features with a cosine similarity greater than or equal to 0.85 are associated as the same target, and the fatigue personnel is followed and monitored according to the tracking identifier corresponding to the target.

[0044] The second aspect of the present application provides a robot target personnel following system, comprising:

[0045] A data acquisition module is configured to acquire and preprocess face video data of a target personnel to be analyzed and working behavior data, wherein the working behavior data comprises behavior data to be analyzed and standard behavior data.

[0046] A data segmentation module is configured to extract expression entropy features based on the face video data to be analyzed, perform density clustering on the expression entropy features to obtain pseudo labels, and divide the face video data to be analyzed into image data of different time periods according to the pseudo labels.

[0047] A problem data extraction module is configured to extract the working behavior data of a corresponding time period according to the image data, and take a time period in which the behavior data to be analyzed and the standard behavior data of the working behavior data are different as problem data.

[0048] A differential feature extraction module is configured to extract time differential features of eye regions of interest and mouth regions of interest according to the problem data, and extract behavior differential features according to the behavior data to be analyzed at a time point corresponding to the time differential features.

[0049] A fatigue feature extraction module is configured to perform nonlinear dimension reduction on the time differential features and the behavior differential features by kernel entropy component analysis to obtain principal components, perform mutual information screening using the principal components, and obtain fatigue correlation features.

[0050] A fatigue model modeling module is configured to input the fatigue correlation features into an entropy optimization neural network for training, obtain a fatigue recognition model, output a fatigue level of the target personnel using the fatigue recognition model, and track and follow the recognized fatigued personnel.

[0051] Compared with the prior art, the embodiments of the present application have at least the following advantages or beneficial effects:

[0052] The present application can more comprehensively and deeply mine multi-dimensional features in a fatigue state by constructing a dynamic pseudo label mechanism and a kernel entropy component analysis framework, co-optimizing facial expression entropy features and physiological behavior features, capturing the influence of facial micro-expression and physiological behavior changes on the fatigue state, improving the accuracy of fatigue state determination, extracting eye opening degree, mouth opening degree changes based on a convolutional neural network model, and macro behavior pattern feature extraction, multi-dimensionally analyzing the fatigue state of the target personnel, improving the capturing ability and analysis comprehensiveness of fatigue feature changes, and improving the accuracy and real-time performance of fatigue state tracking and following. BRIEF DESCRIPTION OF DRAWINGS

[0053] Figure 1A flowchart of steps of a robot target personnel following method in an embodiment of the present application. DETAILED DESCRIPTION

[0054] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0055] Referring to Figure 1 The robot target personnel following method provided by the present application comprises the following steps.

[0056] The target personnel's to-be-analyzed facial video data and work behavior data are collected and preprocessed, and the work behavior data comprises to-be-analyzed behavior data and standard behavior data.

[0057] In actual evaluation, the target personnel's work scene video is collected by an industrial camera, with a resolution of 1080x720 and a frame rate of 30fps, and is continuously collected for 8 hours; the heart rate and skin electric response are synchronously collected by a wearable device every minute, and there are 28800 data points in total for 8 hours; 30-minute data of the target personnel in a normal non-fatigue working state are collected, and the mean heart rate of 70 times per minute and the mean skin electric response of 2uS are extracted as the reference; the target personnel's work scene video is subjected to image enhancement and Gaussian denoising by OpenCV, and face detection and alignment are performed by using MTCNN, and the unified size is 128x128 pixels; the heart rate and skin electric response data are subjected to deduplication and interpolation to complete missing values, and finally the synchronous timestamp alignment data are obtained.

[0058] Expression entropy features are extracted based on the to-be-analyzed facial video data, pseudo labels are obtained by density clustering on the expression entropy features, and the to-be-analyzed facial video data are divided into image data of different time periods according to the pseudo labels.

[0059] In actual evaluation, 8 hours of face video data to be analyzed are divided into 48 first sequences at equal intervals, each first sequence is divided into 20 second sequences, 12 facial action units AUs including AU4 frown, AU6 cheek raise, AU12 lip corner up, etc. are extracted using the pre-trained VGG-Face to generate an action unit activation state matrix of each second sequence every 30 frames x 12 AUs, the occurrence frequency of each AU in the second sequence is counted, for example, AU4 occurs 15 times in every 30 frames, the probability is 0.5, a probability distribution set is formed, the expression entropy of the first sequence is calculated according to the expression entropy formula, for example, one expression entropy value is 1.2247, wherein the static entropy variance and the adjacent static entropy difference value are calculated by a sliding window, and finally 12-dimensional expression entropy features of each first sequence are obtained; the expression entropy features are subjected to Z-score standardization to obtain 1000 feature data, the Euclidean distance between data points is calculated, the distance average of 7 adjacent points of each point is taken as the neighborhood radius, for example, the neighborhood radius of a data point is about 0.8, the minimum neighborhood sample number is set to the square root of the feature data, about 32, the DBSCAN algorithm is used, the region with a sample number greater than or equal to 32 in the neighborhood is taken as a core cluster, and finally 3 clustering clusters are obtained, pseudo-labels 1, 2 and 3 are assigned according to the number of samples in the cluster from more to less, and the face video data to be analyzed is divided into image data of different time periods according to the pseudo-labels.

[0060] According to the image data, the working behavior data of the corresponding time period is extracted, and the time period in which the working behavior data to be analyzed and the standard behavior data are different is taken as problem data.

[0061] In actual evaluation, according to the image data, the working behavior data of the corresponding time period is extracted, for example, one behavior parameter is heart rate and skin electric response, the sample size of the standard behavior data is 1000, and the difference degree between the working behavior data to be analyzed and the standard behavior data is calculated for each second sequence corresponding to 30 seconds. For example, the heart rate of a certain period is 85 times per minute, the mean is 70, and the standard deviation is 5, the skin electric response is 3 μS, the mean is 2, and the standard deviation is 0.5, the difference degree is 1.8 obtained by substituting into the difference degree formula, the mean of the difference degrees of all time periods is 1.2, and the standard deviation is 0.5, the problem threshold is set to the sum of the mean and 3 times the standard deviation, which is 2.7, and the time period with a difference degree greater than 2.7 is selected as problem data, and a total of 500 problem data segments are obtained.

[0062] According to the problem data, the time difference features of the eye region of interest and the perioral region of interest are extracted, and the behavior difference features are extracted according to the working behavior data to be analyzed corresponding to the time points of the time difference features.

[0063] In actual evaluation, the eye and perioral area are located using the ResNet-50 technology to obtain the pupil diameter and the distance between the left and right corners of the mouth, the eyelid and lip midpoint motion trajectories within 2 seconds are tracked by the optical flow method, the motion trajectories of the eyelid and the lip midpoint within 2 seconds are obtained by the optical flow method according to the eye region of interest and the perioral region of interest, the eye opening degree is obtained by using the average distance between the upper eyelid and the lower eyelid divided by the pupil diameter according to the motion trajectory of the eyelid, the eye opening degree time sequence is obtained by integrating the eye opening degree in time sequence, the mouth opening degree is obtained by using the distance between the upper lip midpoint and the lower lip midpoint divided by the distance between the left and right corners of the mouth according to the motion trajectory of the lip midpoint, the mouth opening degree time sequence is obtained by integrating the mouth opening degree in time sequence, the adjacent frame difference value and the 5-point sliding average change rate are calculated, and the 230-dimensional time difference feature is obtained after splicing; the heart rate and the skin electric response data at the corresponding time points are calculated, the adjacent point difference value and the 5-point average change rate are calculated, and the 200-dimensional behavior difference feature is obtained after splicing.

[0064] The time difference feature and the behavior difference feature are subjected to kernel entropy component analysis for nonlinear dimension reduction to obtain principal components, and the principal components are subjected to mutual information screening to obtain fatigue correlation features.

[0065] In actual evaluation, the 230-dimensional time difference feature and the 200-dimensional behavior difference feature are spliced to obtain a 430-dimensional row vector, a total of 500 question data segments are obtained, a 500*430 feature matrix is formed, a kernel matrix is calculated based on the feature matrix using a Gaussian kernel function with a bandwidth of 1, the kernel matrix is centralized, and then eigenvalue decomposition is started to obtain eigenvalues and eigenvectors, the sparsity of the feature matrix is the total number of elements multiplied by the number of zero elements, and the regularization parameter is adjusted according to the sparsity, for example, if the sparsity is greater than 30%, α is set to 0.1, and if the sparsity is less than or equal to 30%, α is set to 0.01, wherein β is 0.5 times α, the cumulative contribution is calculated using the eigenvalues, eigenvectors and regularization parameter, for example, the cumulative contribution is 0.394, and when the current k ′ cumulative contribution is greater than or equal to 85%, the first 8 principal components are retained as candidate features, the mutual information between the 8 principal components and the pseudo label is calculated, 5 principal components with mutual information higher than the median are retained, and finally 5-dimensional fatigue correlation features are obtained.

[0066] The fatigue correlation features are input into an entropy optimization neural network for training to obtain a fatigue recognition model, the fatigue level of a target person is output using the fatigue recognition model, and the recognized fatigue person is tracked and followed.

[0067] In the actual evaluation, according to the fatigue correlation feature, 500 samples are divided into a training set 300, a validation set 100 and a test set 100 according to 6:2:2, a 3-layer fully connected layer is constructed according to entropy optimization of the neural network, the input layer accepts 5-dimensional input features, the hidden layer contains two layers, the first layer has 128 neurons, and the second layer has 64 neurons; the output layer outputs 3 dimensions, respectively corresponding to the prediction output of low, medium and high fatigue levels, the activation function is ReLU, cross entropy is used as the loss function, the neural network learning rate is set to 0.001, batch_size is 32, the early stopping condition is that the validation set accuracy is less than 0.7% for 5 consecutive rounds, the final training is stopped after 15 rounds, the performance of the test set is input into the model, the accuracy is 92%, the confusion matrix shows that the recognition accuracy of medium and high fatigue levels is greater than or equal to 88%, a fatigue recognition model is obtained, according to the fatigue recognition model, the current fatigue level of the target person is output as high fatigue through the softmax probability distribution, and the tracked identifier is assigned to the recognized fatigue person, for example, the identifier PID-001 is assigned to the high fatigue person, the ReID feature of the fatigue person is extracted according to the tracked identifier, the cosine similarity between different frames of the face video data to be analyzed is calculated using the ReID feature, the ReID features with a cosine similarity greater than or equal to 0.85 are associated as the same target, and the fatigue person is monitored according to the tracked identifier corresponding to the target.

[0068] In the embodiment, the pre-processing method comprises:

[0069] The face video data to be analyzed of the target person is collected, image enhancement, denoising, uniform image size and posture are performed on the face video data to be analyzed, the collected heart rate data and skin electric reaction data of the target person are taken as the behavior data to be analyzed, the heart rate data and skin electric reaction data of the target person in a normal state are taken as the standard behavior data, and the behavior data to be analyzed and the standard behavior data are cleaned and de-duplicated to obtain the working behavior data.

[0070] In the embodiment, the method for extracting expression entropy features based on the face video data to be analyzed comprises:

[0071] The first sequence is obtained by equal-interval division based on the face video data to be analyzed, the second sequence is obtained by equal-interval division based on the first sequence, the facial action unit is extracted by the convolutional neural network model based on the second sequence, and the action unit data sequence is generated based on the activation state of the facial action unit in the image frame of the second sequence.

[0072] According to the action unit data sequence, the number of occurrences of different action units is counted, the frequency of occurrence is divided by the total number of image frames to obtain a probability value, the probability values of the action units of the second sequence are summarized to form a probability distribution set, and a facial expression entropy formula is used to calculate the facial expression entropy feature of the first sequence based on the probability distribution set of each second sequence in the first sequence, and the facial expression entropy formula is:

[0073]

[0074] Where H is the facial expression entropy value of the first sequence, t is the second sequence index, T is the total number of second sequences, i is the action unit index, n is the total number of action units, p i (t) is the probability of the occurrence of the i-th action unit in the t-th second sequence, is the variance of the static entropy in the second sequence, is the difference value of the static entropy of adjacent second sequences.

[0075] In this embodiment, the facial expression entropy feature is subjected to density clustering to obtain a pseudo-labeling method, which comprises:

[0076] The facial expression entropy feature is subjected to Z-score standardization to obtain a facial expression entropy feature dataset, the Euclidean distance between data points is calculated according to the facial expression entropy feature dataset, and the average Euclidean distance of the 7 adjacent data points is taken as the neighborhood radius of hierarchical density clustering;

[0077] The square root integer value of the data amount of the facial expression entropy feature dataset is used as the minimum neighborhood sample number of hierarchical density clustering, the facial expression entropy feature dataset is subjected to hierarchical density clustering based on the neighborhood radius and the minimum neighborhood sample number, the region is taken as a clustering core if the number of data points in the neighborhood radius is greater than or equal to the minimum neighborhood sample number, the data points in the neighborhood radius are included in the same cluster according to the clustering core, and the clustering cluster is obtained, and the pseudo-label is obtained by sequentially assigning numerical labels to the data points in the clustering cluster in the order from more to less according to the number of data points.

[0078] In this embodiment, the method of taking the time period in which the to-be-analyzed behavior data and the standard behavior data differ as problem data comprises:

[0079] The difference degree is calculated based on the to-be-analyzed behavior data and the standard behavior data, and the difference degree formula is:

[0080]

[0081] Where S is the difference degree of the to-be-analyzed behavior data and the standard behavior data, i is the summation index variable, j is the summation index variable, n is the number of behavior parameters, m is the sample number of the standard behavior data, x i is the to-be-analyzed behavior data value of the i-th parameter, x ikis the kth standard behavior data value of the ith parameter, μ i is the mean of the standard behavior data of the ith parameter, x jk is the kth standard behavior data value of the jth parameter, X i is the behavior data value to be analyzed of the ith parameter, Y i is the standard behavior data value of the ith parameter, ∈ is a very small constant to prevent division by zero error;

[0082] The sum of the difference degree mean and 3 times the difference degree standard deviation in the behavior data to be analyzed is used as a problem threshold, and the behavior data to be analyzed with a difference degree greater than the problem threshold in different time periods is taken as problem data.

[0083] In this embodiment, the method for obtaining time difference features comprises:

[0084] According to the image data in the problem data, the eye region of interest and the perioral region of interest are extracted through a convolutional neural network model, the pupil diameter is extracted based on the eye region of interest, and the distance between the left and right corners of the mouth is extracted based on the perioral region of interest.

[0085] According to the eye region of interest and the perioral region of interest, the motion trajectory of the eyelid and the motion trajectory of the lip midpoint within two seconds are obtained by the optical flow method, and the eye opening degree is obtained by using the average distance between the upper eyelid and the lower eyelid divided by the pupil diameter according to the motion trajectory of the eyelid.

[0086] The eye opening degree is integrated in time sequence to obtain an eye opening degree time sequence, and the mouth opening degree is obtained by using the distance between the upper lip midpoint and the lower lip midpoint divided by the distance between the left and right corners of the mouth according to the motion trajectory of the lip midpoint, and the mouth opening degree time sequence is obtained by integrating the mouth opening degree in time sequence.

[0087] The difference between two data points and the average change rate within 5 data points are calculated based on the data points of the eye opening degree time sequence in sequence, the difference between two data points and the average change rate within 5 data points are calculated according to the data points of the mouth opening degree time sequence in sequence, and the data point difference of the eye opening degree time sequence, the data point average change rate of the eye opening degree time sequence, the data point difference of the mouth opening degree time sequence, and the data point average change rate of the mouth opening degree time sequence are sequentially spliced to obtain the time difference features.

[0088] In this embodiment, the method for obtaining behavior difference features comprises:

[0089] The time-difference feature is used to obtain the behavior data at the corresponding time point, and the difference between two adjacent data points and the average change rate in five continuous data points are calculated according to the heart rate data of the behavior data to be analyzed in sequence; the difference between two adjacent data points and the average change rate in five continuous data points are calculated according to the skin conductance response data of the behavior data to be analyzed in sequence, and the heart rate data difference, the heart rate data average change rate, the skin conductance response data difference and the skin conductance response data average change rate are spliced in sequence to obtain the behavior difference feature.

[0090] In the embodiment, the method for obtaining the fatigue correlation feature comprises:

[0091] The time-difference feature and the behavior difference feature are aligned on the time axis, the time-difference feature and the behavior difference feature are spliced to obtain a row vector, the row vector is used for superposition to obtain a feature matrix, a kernel matrix is calculated according to the feature matrix through a Gaussian kernel function;

[0092] The kernel matrix is subjected to centering processing and eigenvalue decomposition to obtain eigenvalues and eigenvectors, the feature matrix is subjected to sparsity analysis to obtain regularization parameters, the contribution degree is calculated according to the eigenvalues, the eigenvectors and the regularization parameters, and the contribution degree formula is:

[0093]

[0094] Wherein S(k ′ ) is the cumulative contribution degree score of the first k ′ principal components, λ i is the i th eigenvalue of the kernel matrix, e i,k is the k th element of the i th eigenvector, e j,m is the m th element of the j th eigenvector, n is the number of data samples, m is the index, α is the regularization parameter, β is the regularization parameter, k ′ is the number of principal components currently considered;

[0095] If the cumulative contribution degree score is greater than or equal to 0.85, the first k ′ principal components are used as candidate features, the mutual information of the candidate features and the pseudo label is calculated, the candidate features with mutual information higher than the median are reserved to obtain the fatigue correlation feature.

[0096] In the embodiment, the method for tracking and following the identified fatigue personnel comprises:

[0097] The training set, the verification set and the test set are obtained based on fatigue correlation feature division, the training set is input into an entropy optimization neural network, during training, the difference between the prediction result and the pseudo label is calculated through a cross entropy loss function, the verification set is input into the model, and the overfitting state is verified according to the loss function value and the accuracy rate change, if the accuracy rate of the verification set is less than 0.7% for 5 consecutive rounds, the training is ended, the difference is fed back to the parameters of each layer of the model using a back propagation algorithm, the test set is input into the trained entropy optimization neural network, the accuracy rate and the confusion matrix are used to evaluate the performance of the model, a fatigue recognition model is obtained, and the fatigue level of the target person is output based on the fatigue recognition model.

[0098] The recognized fatigue personnel is given a tracking identifier, the ReID feature of the fatigue personnel is extracted according to the tracking identifier, the cosine similarity between the ReID features of different frames in the to-be-analyzed face video data is calculated, the ReID features with a cosine similarity greater than or equal to 0.85 are associated as the same target, and the fatigue personnel is monitored according to the tracking identifier corresponding to the target.

[0099] The second aspect of the present application also provides a robot target personnel following system, comprising:

[0100] The data acquisition module is used for collecting and preprocessing the to-be-analyzed face video data and the work behavior data of the target personnel, and the work behavior data includes to-be-analyzed behavior data and standard behavior data.

[0101] The data segmentation module is used for extracting expression entropy features based on the to-be-analyzed face video data, performing density clustering on the expression entropy features to obtain pseudo labels, and dividing the to-be-analyzed face video data into image data of different time periods according to the pseudo labels.

[0102] The problem data extraction module is used for extracting the work behavior data of the corresponding time period according to the image data, and taking the time period in which the to-be-analyzed behavior data and the standard behavior data of the work behavior data are different as problem data.

[0103] The differential feature extraction module is used for extracting time differential features of the eye region of interest and the mouth region of interest according to the problem data, and extracting behavior differential features from the to-be-analyzed behavior data at the corresponding time point according to the time differential features.

[0104] The fatigue feature extraction module is used for performing nonlinear dimension reduction on the time differential features and the behavior differential features by kernel entropy component analysis to obtain principal components, and performing mutual information screening using the principal components to obtain fatigue correlation features.

[0105] The fatigue model modeling module is configured to input the fatigue correlation features into an entropy optimization neural network for training, obtain a fatigue recognition model, output a fatigue level of the target person using the fatigue recognition model, and track and follow the recognized fatigue person.

[0106] The above is only an example and a description of the structure of the present application. Those skilled in the art can make various modifications or supplements to the described specific embodiments or use similar ways to replace, as long as the modifications or supplements do not deviate from the structure of the present application or exceed the scope defined by the present application, and should belong to the protection scope of the present application.

Claims

1. A method of robotic target person following, characterized in that, The method comprises the following steps: Collecting target personnel's to-be-analyzed facial video data and work behavior data for preprocessing, wherein the work behavior data comprises to-be-analyzed behavior data and standard behavior data; Extracting expression entropy features based on the to-be-analyzed facial video data, performing density clustering on the expression entropy features to obtain pseudo labels, and dividing the to-be-analyzed facial video data into image data of different time periods according to the pseudo labels; Extracting the work behavior data corresponding to the time period according to the image data, and taking a time period in which the to-be-analyzed behavior data and the standard behavior data of the work behavior data are different as problem data; Extracting time difference features of eye regions of interest and perioral regions of interest according to the problem data, and extracting behavior difference features from the to-be-analyzed behavior data at corresponding time points according to the time difference features; Performing nonlinear dimension reduction on the time difference features and the behavior difference features by kernel entropy component analysis to obtain principal components, performing mutual information screening using the principal components, and obtaining fatigue correlation features; Inputting the fatigue correlation features into an entropy optimization neural network for training to obtain a fatigue recognition model, outputting a fatigue level of the target personnel using the fatigue recognition model, and tracking and following the recognized fatigued personnel.

2. The method of claim 1, wherein, The preprocessing method comprises: Collecting target personnel's to-be-analyzed facial video data, performing image enhancement, denoising, and unifying image size and posture on the to-be-analyzed facial video data, collecting heart rate data and skin electric response data of the target personnel as to-be-analyzed behavior data, collecting heart rate data and skin electric response data of the target personnel in a normal state as standard behavior data, and performing cleaning and deduplication on the to-be-analyzed behavior data and the standard behavior data to obtain work behavior data.

3. The method of claim 1, wherein, The method for extracting expression entropy features based on the to-be-analyzed facial video data comprises: Dividing the to-be-analyzed facial video data into a first sequence based on equal intervals, dividing a second sequence based on the first sequence according to equal intervals, extracting facial action units based on a convolutional neural network model according to the second sequence, and generating an action unit data sequence based on the activation state of the facial action units in the image frames of the second sequence; Counting the number of occurrences of different action units according to the action unit data sequence, dividing the frequency of occurrence by the total number of image frames to obtain a probability value, summarizing the probability values of the action units of the second sequence to form a probability distribution set, and calculating the first sequence expression entropy features based on the probability distribution set of each second sequence in the first sequence using an expression entropy formula, wherein the expression entropy formula is: where H is the first sequence index, t is the second sequence index, T is the total number of second sequences, i is the action unit index, n is the total number of action units, p i (t) is the probability of the i-th action unit appearing in the t-th second sequence, is the variance of the static entropy in the second sequence, is the static entropy difference value of adjacent second sequences.

4. The method of claim 1, wherein, The method for performing density clustering on the expression entropy features to obtain pseudo labels comprises: Performing Z-score standardization based on the expression entropy features to obtain an expression entropy feature dataset, calculating the Euclidean distance between data points according to the expression entropy feature dataset, and taking the average of the Euclidean distances of the adjacent 7 points of the data points as the neighborhood radius of hierarchical density clustering. The square root of the integer value of the data amount of the expression entropy feature dataset is used as the minimum neighborhood sample number of hierarchical density clustering, hierarchical density clustering is performed on the expression entropy feature dataset based on the neighborhood radius and the minimum neighborhood sample number, if the number of data points in the neighborhood radius is greater than or equal to the minimum neighborhood sample number, the region is taken as a clustering core, the data points in the neighborhood radius are included in the same cluster according to the clustering core, until the clustering cluster is obtained, and the numerical labels are sequentially given according to the order from more to less of the number of data points in the clustering cluster, to obtain pseudo-labels.

5. The method of claim 1, wherein, The method of taking the time period in which the to-be-analyzed behavior data and the standard behavior data differ as problem data comprises: The difference degree is calculated based on the to-be-analyzed behavior data and the standard behavior data, and the difference degree formula is: where S is the difference between the behavior data under analysis and the standard behavior data, i is the summation index variable, j is the summation index variable, n is the number of behavior parameters, m is the number of samples of the standard behavior data, x i is the value of the behavior data under analysis for the i-th parameter, x ik is the k-th value of the standard behavior data for the i-th parameter, x i is the mean of the standard behavior data for the i-th parameter, x jk is the k-th value of the standard behavior data for the j-th parameter, x i is the value of the behavior data under analysis for the i-th parameter, Y i is the value of the standard behavior data for the i-th parameter, and ∈ is a small constant to prevent division by zero errors. The sum of the mean value of the difference degree in the to-be-analyzed behavior data and 3 times the standard deviation of the difference degree is used as a problem threshold, and the to-be-analyzed behavior data with a difference degree greater than the problem threshold in different time periods are taken as problem data.

6. The method of claim 1, wherein, The method of obtaining the time difference feature comprises: The eye region of interest and the perioral region of interest are extracted from the image data in the problem data through a convolutional neural network model, the pupil diameter is extracted based on the eye region of interest, and the distance between the left and right corners of the mouth is extracted based on the perioral region of interest; The motion trajectory of the eyelid and the motion trajectory of the lip midpoint within two seconds are obtained by the optical flow method according to the eye region of interest and the perioral region of interest, and the eye opening degree is obtained by using the average distance between the upper eyelid and the lower eyelid divided by the pupil diameter according to the motion trajectory of the eyelid; The eye opening degree is integrated in time sequence to obtain an eye opening degree time sequence, and the mouth opening degree is obtained by using the distance between the upper lip midpoint and the lower lip midpoint divided by the distance between the left and right corners of the mouth according to the motion trajectory of the lip midpoint, and the mouth opening degree is integrated in time sequence to obtain a mouth opening degree time sequence; The difference between two data points and the average change rate within 5 data points are sequentially calculated based on the data points of the eye opening degree time sequence, the difference between two data points and the average change rate within 5 data points are sequentially calculated based on the data points of the mouth opening degree time sequence, and the data point difference of the eye opening degree time sequence, the average change rate of the data points of the eye opening degree time sequence, the data point difference of the mouth opening degree time sequence, and the average change rate of the data points of the mouth opening degree time sequence are sequentially spliced to obtain the time difference feature.

7. The method of claim 1, wherein, The method of obtaining the behavior difference feature comprises: The to-be-analyzed behavior data at the corresponding time point is obtained based on the time difference feature, the difference between two adjacent data points and the average change rate within 5 consecutive data points are sequentially calculated according to the heart rate data of the to-be-analyzed behavior data, the difference between two adjacent data points and the average change rate within 5 consecutive data points are sequentially calculated according to the galvanic skin response data of the to-be-analyzed behavior data, and the heart rate data difference, the heart rate data average change rate, the galvanic skin response data difference, and the galvanic skin response data average change rate are sequentially spliced to obtain the behavior difference feature.

8. The method of claim 1, wherein, The method of obtaining the fatigue correlation feature comprises: The time difference feature and the behavior difference feature are spliced to obtain a row vector by time axis alignment based on the time difference feature and the behavior difference feature, the row vector is superimposed to obtain a feature matrix, and a kernel matrix is calculated from the feature matrix by using a Gaussian kernel function; The kernel matrix is centralized, eigenvalue decomposition is performed, the eigenvalue and the eigenvector are obtained, the sparsity of the feature matrix is analyzed to obtain a regularization parameter, the contribution degree is calculated according to the eigenvalue, the eigenvector and the regularization parameter, and the contribution degree formula is: where S(k ′ ) is the cumulative contribution score of the first k ′ principal components, λ i is the i-th eigenvalue of the kernel matrix, e i,k is the k-th element of the i-th eigenvector, e j,m is the m-th element of the j-th eigenvector, n is the number of data samples, m is the index, α is a regularization parameter, β is a regularization parameter, and k ′ is the number of principal components currently considered. If the cumulative contribution score is greater than or equal to 0.85, the first k ′ principal components are taken as candidate features, the mutual information of the candidate features and the pseudo labels is calculated, and the candidate features with mutual information higher than the median are retained to obtain the fatigue correlation features.

9. The method of claim 1, wherein, The method for tracking and following the identified fatigue personnel comprises: The training set, the verification set and the test set are obtained based on the fatigue correlation feature division, the training set is input into the entropy optimization neural network, the difference between the prediction result and the pseudo label is calculated by using the cross-entropy loss function during training, the verification set is input into the model, the overfitting state is verified according to the loss function value and the accuracy change, if the accuracy of the verification set is less than 0.7% for five consecutive rounds, the training is ended, the difference is fed back to each layer of the model by using the back propagation algorithm to adjust the parameters, the test set is input into the trained entropy optimization neural network, the model performance is evaluated by using the accuracy and the confusion matrix, and a fatigue recognition model is obtained; and the fatigue level of the target personnel is output based on the fatigue recognition model. The identified fatigue personnel is given a tracking identifier, the ReID feature of the fatigue personnel is extracted according to the tracking identifier, the cosine similarity is calculated according to the ReID features of different frames in the face video data to be analyzed, the ReID features with a cosine similarity greater than or equal to 0.85 are associated as the same target, and the fatigue personnel is followed and monitored according to the tracking identifier corresponding to the target.

10. A robotic target person following system for performing a robotic target person following method according to any one of claims 1 to 9, characterized by The system comprises: A data acquisition module is configured to acquire and pre-process face video data to be analyzed of target personnel and work behavior data, wherein the work behavior data comprises analysis behavior data and standard behavior data; A data segmentation module is configured to extract expression entropy features based on the face video data to be analyzed, perform density clustering on the expression entropy features to obtain pseudo labels, and divide the face video data to be analyzed into image data of different time periods according to the pseudo labels; A problem data extraction module is configured to extract the work behavior data of a corresponding time period from the image data, and regard a time period in which the analysis behavior data and the standard behavior data of the work behavior data are different as problem data; A difference feature extraction module is configured to extract time difference features of eye regions of interest and mouth regions of interest from the problem data, and extract behavior difference features corresponding to time points of the analysis behavior data from the time difference features; A fatigue feature extraction module is configured to perform nonlinear dimension reduction on the time difference features and the behavior difference features by using kernel entropy component analysis to obtain principal components, and perform mutual information screening on the principal components to obtain fatigue correlation features; A fatigue model modeling module is configured to input the fatigue correlation features into an entropy optimization neural network for training, obtain a fatigue recognition model, output a fatigue level of target personnel by using the fatigue recognition model, and track and follow identified fatigue personnel.

Citation Information

Patent Citations

  • Face fatigue recognition method and device, equipment and storage medium

    CN117058731A