A visual-based early fatigue detection method, device and storage medium

CN116246257BActive Publication Date: 2026-09-25ARCSOFT CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211652612.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-21
Publication Date
2026-09-25
Estimated Expiration
2042-12-21

AI Technical Summary

Technical Problem

[0002]疲劳驾驶检测技术能够防止交通意外的发生,目前主要的疲劳驾驶检测包括:基于视觉的方法,比如通过眼睛和嘴巴状态进行疲劳检测,该类方法在光照变化时精度较低;基于生理指标的方法,比如通过分析脑电和肌电信号进行疲劳检测,该类方法通常需要放置与人体接触的采集装置,会影响驾驶员的驾驶体验

Benefits of technology

[0025]在本发明实施例中,通过第一模型对图像序列进行识别,确定当前帧的行为向量,其中,所述当前帧为所述图像序列中的最后一帧;对第一预设时间内的多个所述行为向量进行统计,获得统计结果;利用训练好的第二模型对所述统计结果进行分类,得到早期疲劳检测结果。在本实施例中能够解决无法在早期对疲劳行为提前发出预警的问题,从而给用户更充足的时间进行调整,更多地降低疲劳驾驶风险。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116246257B_ABST
    Figure CN116246257B_ABST
Patent Text Reader

Abstract

The application discloses a kind of early fatigue detection method, device and storage medium based on vision.Therein, the early fatigue detection method includes: by first model, image sequence is identified, and the behavior vector of current frame is determined, wherein current frame is the last frame in image sequence;Multiple behavior vectors in first preset time are counted, and statistical result is obtained;Classify statistical result using the second model trained, and obtain early fatigue detection result.Invention can solve the problem that early fatigue behavior cannot be warned in advance, so as to give user more sufficient time for adjustment, more reduce fatigue driving risk.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This article relates to fatigue behavior detection technology, and more particularly to a vision-based method, device, and storage medium for early fatigue detection. Background Technology

[0002] Fatigue driving detection technology can prevent traffic accidents. Currently, the main methods for fatigue driving detection include: vision-based methods, such as detecting fatigue through eye and mouth movements, which have lower accuracy when lighting changes; and physiological indicator-based methods, such as analyzing electroencephalogram (EEG) and electromyogram (EMG) signals, which typically require the placement of contact devices, affecting the driver's experience. Furthermore, these methods usually only trigger an alarm when the driver is already fatigued, failing to provide early warnings of fatigue behavior and allow sufficient time for adjustment to reduce the risk of fatigued driving.

[0003] There is currently no effective solution to the above problems. Summary of the Invention

[0004] This application provides a vision-based early fatigue warning method, device, and storage medium, which can provide early fatigue warning based on behavioral characteristics.

[0005] According to one aspect of the present invention, a vision-based early fatigue detection method is provided, comprising: identifying an image sequence using a first model to determine a behavior vector of the current frame, wherein the current frame is the last frame in the image sequence; statistically analyzing multiple behavior vectors within a first preset time period to obtain statistical results; and classifying the statistical results using a trained second model to obtain an early fatigue detection result.

[0006] Optionally, obtaining the image sequence includes: filtering valid frames from the input data, wherein the valid frames contain the target object and, and the time interval between two adjacent valid frames is greater than a first interval threshold; taking the valid frame with the latest timestamp as the current frame, and forming the image sequence by combining the current frame and the consecutive valid frames preceding it, wherein the length of the image sequence is a preset length.

[0007] Optionally, the behavior vector of the current frame is determined by recognizing the image sequence through the first model, including: extracting the static feature map sequence and dynamic feature map sequence of the image sequence through the convolutional layer of the first model, wherein the convolutional layer includes a three-dimensional feature layer and / or a pseudo optical flow feature layer; and combining the static feature map sequence and the dynamic feature map sequence to perform classification calculation using the fully connected layer of the first model to obtain the behavior vector of the current frame.

[0008] Optionally, when the convolutional layer includes the pseudo-optical flow feature layer, determining the behavior vector of the current frame includes: extracting a displacement feature sequence based on the static feature map sequence through the pseudo-optical flow feature layer; when the convolutional layer only includes the pseudo-optical flow feature layer, using the displacement feature sequence as the dynamic feature map sequence.

[0009] Optionally, when the convolutional layer includes the three-dimensional feature layer and the pseudo optical flow feature layer, the static image sequence and the displacement feature sequence are concatenated and then input into the three-dimensional feature layer to obtain the dynamic feature image sequence.

[0010] Optionally, the step of extracting the displacement feature sequence through the pseudo optical flow feature layer includes: filtering out a target static feature map from the static feature map sequence; obtaining the displacement vector of each feature point in the target static feature map, determining multiple displacement vectors based on multiple feature points, and assembling all the displacement vectors into a displacement vector map; and using each frame in the static feature map sequence as the corresponding displacement vector map determined by the target static feature map to assemble the displacement vector sequence.

[0011] Optionally, in the target static feature map, the displacement vector of each feature point is obtained, including:

[0012] For the first feature point in the target static feature map, in the previous frame static feature map, find the second feature point with the greatest similarity to the first feature point in the neighborhood of the same position of the first feature point, and calculate the displacement vector based on the position of the second feature point and the position of the first feature point.

[0013] Optionally, when the convolutional layer includes the three-dimensional feature layer, determining the behavior vector of the current frame includes: when the convolutional layer only includes the three-dimensional feature layer, simultaneously extracting the static feature map sequence and the dynamic feature map sequence through the three-dimensional feature layer; when the convolutional layer includes both the two-dimensional feature layer and the three-dimensional feature layer, extracting the static feature map sequence through the two-dimensional feature layer, and extracting the dynamic feature map sequence based on the static feature map sequence through the three-dimensional feature layer.

[0014] Optionally, the behavior vector includes at least one of the following: facial expression vector, body behavior vector, and interaction behavior vector.

[0015] Optionally, before performing statistics on multiple behavior vectors within a first preset time period and obtaining the statistical results, the method includes: obtaining multiple behavior vectors within the first preset time period, including: using a queue structure to cache the behavior vectors of each frame in the latest first preset time period and maintaining the queue, using the queue as the multiple behavior vectors, wherein the behavior vector corresponding to the end of the queue is the behavior vector of the current frame, and the difference between the timestamps of the behavior vector corresponding to the head of the queue and the behavior vector corresponding to the end of the queue is the first preset time period.

[0016] Optionally, the step of statistically analyzing multiple behavior vectors within a first preset time period to obtain statistical results includes: using a sliding window to statistically analyze multiple behavior vectors within the first preset time period to obtain the statistical results, wherein the first preset condition includes: the behavior vector triggering begins when the duty cycle of the behavior vector within the sliding window is greater than a first threshold; and the behavior vector triggering ends when the duty cycle of the behavior vector within the sliding window is less than a second threshold; the statistical results include at least one of the following: number of triggers, average trigger duration, maximum trigger duration, and total trigger duration.

[0017] Optionally, the early fatigue detection results include: no potential early fatigue risk, and potential early fatigue risk.

[0018] Optionally, the training method for the second model includes: constructing a training dataset by collecting sample data from users; training an initial second model using the training dataset, and obtaining the trained second model.

[0019] Optionally, constructing a training dataset from user-collected sample data includes: segmenting the sample data according to the duration of the first preset time to obtain multiple sample data segments, wherein the types of the multiple sample data segments include sample data segments that the user determines are fatigued and sample data segments that the user determines are not fatigued; selecting a target sample data segment from the multiple sample data segments, identifying the target sample data segment using the first model, obtaining a sample behavior vector sequence corresponding to the target sample data segment, statistically analyzing the sample behavior vector sequence to obtain a sample statistical result for the target sample data segment; using each of the multiple sample data segments as the target sample data segment, obtaining the corresponding sample statistical result, and constructing the multiple sample data segments containing the sample statistical result into the training dataset.

[0020] Optionally, the sample behavior vector sequence is composed of sample behavior vectors, wherein the sample behavior vectors include the following: sample facial expression vectors, sample body behavior vectors, and sample interaction behavior vectors.

[0021] Optionally, the sample statistics include the following: number of triggers, average trigger duration, maximum trigger duration, and total trigger duration.

[0022] According to another aspect of the present invention, a vision-based early fatigue detection device is also provided. The device includes: a recognition unit, configured to recognize an image sequence using a first model to determine a behavior vector of the current frame, wherein the current frame is the last frame in the image sequence; a statistics unit, configured to perform statistics on multiple behavior vectors within a first preset time period to obtain statistical results; and a classification unit, configured to classify the statistical results using a trained second model to obtain an early fatigue detection result.

[0023] According to another aspect of the present invention, a computer-readable storage medium is also provided, the computer-readable storage medium storing one or more programs, the one or more programs being executable by one or more processors to implement the early fatigue detection method described in any one of the above embodiments.

[0024] According to another aspect of the present invention, a processing apparatus is also provided, comprising: a memory for storing computer-executable instructions; and a processor for executing the computer-executable instructions to implement the steps of the early fatigue detection method described in any one of the preceding embodiments.

[0025] In this embodiment of the invention, an image sequence is identified using a first model to determine the behavior vector of the current frame, wherein the current frame is the last frame in the image sequence; multiple behavior vectors within a first preset time period are statistically analyzed to obtain statistical results; and a trained second model is used to classify the statistical results to obtain early fatigue detection results. This embodiment solves the problem of not being able to issue early warnings for fatigue behavior, thus giving users more time to adjust and further reducing the risk of fatigued driving.

[0026] Other features and advantages of this application will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the application. Other advantages of this application can be realized and obtained by means of the solutions described in the description and the accompanying drawings. Attached Figure Description

[0027] The accompanying drawings are used to provide an understanding of the technical solutions of this application and constitute a part of the specification. They are used together with the embodiments of this application to explain the technical solutions of this application and do not constitute a limitation on the technical solutions of this application.

[0028] Figure 1 This is a flowchart of an optional vision-based early fatigue detection method according to an embodiment of the present invention;

[0029] Figure 2 This is a flowchart of an optional image sequence acquisition method according to an embodiment of the present invention;

[0030] Figure 3 This is a flowchart of an optional method for determining a behavior vector according to an embodiment of the present invention;

[0031] Figure 4 This is a flowchart of an optional method for constructing training data according to an embodiment of the present invention;

[0032] Figure 5 This is a block diagram of an optional vision-based early fatigue detection device according to an embodiment of the present invention. Detailed Implementation

[0033] This application describes several embodiments, but these descriptions are exemplary and not restrictive, and it will be apparent to those skilled in the art that many more embodiments and implementations are possible within the scope of the embodiments described herein. Although many possible combinations of features are shown in the drawings and discussed in the detailed description, many other combinations of the disclosed features are also possible. Unless specifically limited, any feature or element of any embodiment may be used in combination with, or may replace, any feature or element of any other embodiment.

[0034] This application includes and contemplates combinations of features and elements known to those skilled in the art. The embodiments, features, and elements disclosed in this application may also be combined with any conventional features or elements to form a unique inventive scheme as defined by the claims. Any feature or element of any embodiment may also be combined with features or elements from other inventive schemes to form another unique inventive scheme as defined by the claims. Therefore, it should be understood that any feature shown and / or discussed in this application may be implemented individually or in any suitable combination. Therefore, the embodiments are not limited except by the limitations imposed by the appended claims and their equivalents. Furthermore, various modifications and changes may be made within the scope of the appended claims.

[0035] Furthermore, in describing representative embodiments, the specification may have presented the method and / or process as a specific sequence of steps. However, the method or process should not be limited to the specific order of steps described herein, to the extent that it does not depend on such a specific order. As will be understood by those skilled in the art, other sequences of steps are also possible. Therefore, the specific order of steps set forth in the specification should not be construed as a limitation of the claims. Moreover, the claims relating to the method and / or process should not be limited to the steps performed in the order written, and those skilled in the art will readily understand that these orders can be varied and still remain within the scope of the embodiments of this application.

[0036] This application provides a vision-based early fatigue detection method, which can be applied to various application scenarios such as driving monitoring and industrial production. It features high processing accuracy, fast processing speed, real-time capability, and applicability to most practical application scenarios.

[0037] The present invention will be described below through detailed embodiments.

[0038] This application provides a vision-based method for early fatigue detection. (See reference...) Figure 1 This is a flowchart of a vision-based early fatigue detection method according to an embodiment of this application, such as... Figure 1 As shown, the method includes:

[0039] S100, the image sequence is identified through the first model to determine the behavior vector of the current frame, wherein the current frame is the last frame in the above image sequence;

[0040] S102, Statistical analysis is performed on multiple behavior vectors within the first preset time period to obtain statistical results;

[0041] S104. The trained second model is used to classify the statistical results to obtain the early fatigue detection results.

[0042] In this application, a first model is used to identify the image sequence and determine the behavior vector of the current frame, where the current frame is the last frame in the image sequence. Multiple behavior vectors within a first preset time period are statistically analyzed to obtain statistical results. A trained second model is then used to classify the statistical results to obtain early fatigue detection results. These steps address the problem of not being able to issue early warnings for fatigue behavior, thus giving users more time to adjust and further reducing the risk of fatigued driving.

[0043] The following description is based on the steps of the above embodiments.

[0044] S100, the image sequence is identified through the first model to determine the behavior vector of the current frame, wherein the current frame is the last frame in the above image sequence;

[0045] In this embodiment, a first model is used to identify the input image sequence and determine whether driver behavior corresponding to early fatigue exists. To achieve accurate behavior recognition, the image sequence extracted from the input data needs to contain complete behavioral information. Existing technologies typically use single-frame images for recognition, and the extracted features lack dynamic information, easily leading to misidentification. In this embodiment, a preset length of image sequence containing valid behavioral information is obtained by uniformly sampling the input data.

[0046] In an optional embodiment, obtaining the above image sequence includes: filtering valid frames from input data, wherein the valid frames contain a target object and, and the time interval between two adjacent valid frames is greater than a first interval threshold; taking the valid frame with the latest timestamp as the current frame, and forming the above image sequence by combining the current frame and the consecutive valid frames preceding it, wherein the length of the above image sequence is a preset length.

[0047] In this embodiment, the input data can be pre-stored source data or source data acquired in real time by the image acquisition device. The installation location of the image acquisition device is not limited, but it must be able to capture the target object. Specifically, for in-vehicle applications, the image acquisition device can be fixed to locations such as the A-pillar, steering wheel, and center console. The placement of the device must ensure that the captured image includes the driver's face.

[0048] Since the input data contains redundant information or noise, this embodiment filters out valid frames that meet preset conditions from the input data, and then assembles these valid frames into an image sequence for subsequent behavior recognition. Furthermore, the valid frames must contain the target object, and the time interval between each valid frame in the image sequence must be greater than a first interval threshold.

[0049] Specifically, the embodiments of this application do not limit the method for determining whether a valid frame contains a target object; CNN, YOLO, etc., can be used. The setting of the aforementioned first interval threshold is determined based on the detection accuracy; the smaller the threshold, the smaller the sampling interval, and the higher the detection accuracy. The aforementioned preset length is determined by the conditions set by the first model for the input data. In addition, the image sequence of the preset length contains valid frames with the latest timestamp to ensure real-time performance. The embodiments of this application can ensure that the composed image sequence is updated in real time, contains complete information required for behavior recognition, and achieves uniform sampling of input data, removes redundant information to reduce computation, and efficiently and rationally uses the features of the input data, laying the foundation for subsequent real-time and accurate fatigue detection.

[0050] For example, refer to Figure 2 A flowchart of an optional image sequence acquisition method according to an embodiment of the present invention is provided. For example... Figure 2 As shown, taking the detection of early driver fatigue behavior in an in-vehicle scenario as an example, and setting the first interval threshold to M milliseconds and the preset frame length to L frames, the steps for selecting valid frames from the input data to form an image sequence are as follows:

[0051] Step 1: Set the timestamp of the previous valid frame to tp, allocate a fixed-size memory buffer to store the L-frame image sequence, and set the number of valid frames s in the buffer to 0.

[0052] Step 2: Capture the current frame image (img) and simultaneously obtain the timestamp (t) of that frame image;

[0053] Step 3: Use a face detection algorithm to determine whether there is a face in the frame. If there is no face, the current frame is determined to be an invalid frame and the program jumps back to step 2. If there is a face, jump to step 4.

[0054] Step 4: Determine whether the time interval tp between the current frame and the previous valid frame is greater than the preset value M. If the time interval is less than M, the current frame is determined to be an invalid frame, and the program jumps back to step 2. If the time interval is greater than or equal to M, the current frame is determined to be a valid frame, and the program jumps to step 5.

[0055] Step 5: Determine if the number of valid frames s in the buffer is equal to the preset length of frames L. If it is equal to L, delete the frame with the earliest timestamp in the buffer, and change s to L-1. Then proceed to step 6.

[0056] Step 6: Add the current valid frame to the buffer, increment the number of valid frames s in the buffer by 1, and update the valid timestamp tp of the previous frame to the current timestamp t. Proceed to Step 7.

[0057] Step 7: Determine if the number of valid frames s in the buffer is equal to the preset length frame number L. If it is equal to L, send the image sequence data in the buffer to the next module and then jump to step 2. Otherwise, jump directly to step 2.

[0058] Specifically, the buffer outputs an image sequence of a preset length. If the number of frames in the image sequence does not meet the preset length, initialization is required, such as continuing to capture and cache valid frames from the input data until the number of valid frames reaches the preset length. In this embodiment, by filtering and outputting an image sequence of a preset length from the input data, which includes valid frames with the latest timestamp and consecutive valid frames preceding the latest timestamp, the continuity of data sampling is ensured, redundant information in the input data is removed, the computational load for subsequent recognition is reduced, and complete information containing the information required for behavior recognition is obtained, thus achieving uniform and effective sampling of the input data.

[0059] In this embodiment, a first model is responsible for identifying the input image sequence and determining whether there are driver behaviors corresponding to early fatigue states. Existing technologies detect fatigue behaviors such as drowsiness, yawning, and looking down, which indicate that the driver is already fatigued and cannot provide early warnings to reduce the risk of fatigued driving.

[0060] This application uses behaviors associated with fatigue as the basis for judgment, classifying early fatigue behaviors into three categories: the first category is facial expression behaviors, including at least one of the following: frowning, raising eyebrows, and tightening lips; the second category is physical behaviors, including at least one of the following: adjusting sitting posture, stretching, rubbing face, rubbing nose, rubbing eyes, and twisting neck; and the third category is common human-object interaction behaviors, including at least one of the following: smoking, drinking water, and making phone calls.

[0061] In one alternative embodiment, the behavior vectors include at least one of the following: facial expression vectors, body behavior vectors, and interaction behavior vectors.

[0062] The first model described above takes an image sequence of a preset length as input and outputs a behavior vector. Specifically, the dimension of the behavior vector is determined by the number of types of early fatigue behaviors that the model supports for detection. It can be all or some of the behavior types included in the three types of behaviors mentioned above. The value of each dimension indicates whether the action was not recognized or whether the action was recognized.

[0063] Furthermore, the three types of behaviors can be uniformly identified using the first model. The input is an image sequence of preset length, and the output is a corresponding behavior vector, which consists of at least one of facial expression vectors, body behavior vectors, and interaction behavior vectors. To more accurately identify the corresponding behaviors, the first model can consist of three independent sub-models. The three types of behaviors use corresponding sub-models. The image sequence of preset length is input into the three sub-models respectively, and the three output behavior sub-vectors (facial expression vector, body behavior vector, and interaction behavior vector) are combined to form a single behavior vector. The first model or the three independent sub-models have the same model structure, and the parameters of the three independent sub-models are adaptively adjusted according to the recognition requirements.

[0064] refer to Figure 3 A flowchart for determining an optional behavior vector according to an embodiment of the present invention is provided. The specific implementation steps of each module are as follows:

[0065] S300, the static feature map sequence and dynamic feature map sequence of the above image sequence are extracted through the convolutional layer of the above first model, wherein the above convolutional layer includes a three-dimensional feature layer and / or a pseudo optical flow feature layer.

[0066] S301, combining the above feature map sequence and the above dynamic feature map sequence, the classification calculation is performed using the fully connected layer of the above first model to obtain the behavior vector of the above current frame.

[0067] In this embodiment, static and dynamic feature information of the image sequence is extracted through the convolutional layer of the first model, ensuring complete extraction of image sequence information.

[0068] In this application embodiment, the static feature map sequence can be extracted for each frame of the image sequence; or for at least one frame of the image sequence, where the at least one frame includes the latest valid frame with the latest timestamp; or the static feature map sequence can be extracted by fusing or overlapping the image sequences to form a corresponding fused or overlapping image.

[0069] Extracting a dynamic feature map sequence from an image sequence can be done by obtaining the dynamic features of each frame in the image sequence, or by obtaining a dynamic feature map sequence of at least one frame in the image sequence.

[0070] Different types of convolutional layers can extract different types of features. Three-dimensional feature layers can extract both static and dynamic features, while two-dimensional feature layers can only extract static features. Therefore, dynamic feature information can be extracted based on a sequence of static feature maps, or it can be obtained simultaneously with a sequence of static feature maps, depending on the type of convolutional layer.

[0071] In an optional embodiment, when the convolutional layer includes a three-dimensional feature layer, determining the behavior vector of the current frame includes:

[0072] When the convolutional layer contains only a three-dimensional feature layer, the static feature map sequence and the dynamic feature map sequence are extracted simultaneously through the aforementioned three-dimensional feature layer.

[0073] When a convolutional layer contains two-dimensional feature layers and three-dimensional feature layers, a static feature map sequence is extracted through the two-dimensional feature layer, and a dynamic feature map sequence is extracted through the three-dimensional feature layer based on the static feature map sequence.

[0074] Specifically, when the convolutional layer contains only three-dimensional feature layers, after the image sequence passes through a first set number of three-dimensional feature layers, it can simultaneously extract static and dynamic information, and obtain static feature map sequences and dynamic feature map sequences.

[0075] Since the computational load of the three-dimensional feature layer is enormous, this embodiment also utilizes the advantage of the two-dimensional feature layer in saving computational power, thereby reducing the computational requirements of the entire early fatigue detection system. The image sequence is first processed through a second predetermined number of two-dimensional feature layers to obtain a static feature map sequence, and then the static feature map sequence is input into a third predetermined number of three-dimensional feature layers to obtain a dynamic feature map sequence.

[0076] Convolutional layers use a combination of two-dimensional and three-dimensional feature layers. On the one hand, the two-dimensional feature layers reduce the computational cost of using only three-dimensional feature layers. On the other hand, the three-dimensional features compensate for the lack of dynamic feature extraction capabilities of two-dimensional feature layers alone. The combination of the two can reduce the computational cost of the model while enabling the network to obtain dynamic feature extraction capabilities.

[0077] Since 3D convolution requires high computational power, the 3D feature layer that can be used in the example can be the VGG series. Furthermore, this application does not limit the specific structure of the 2D and 3D feature layers, which consist of several convolutional units and batch normalization units.

[0078] Furthermore, in the process of extracting dynamic feature map sequences based on static feature map sequences, directly passing features from historical frames to the current frame can lead to feature mismatch, as this is because the changing spatial position of objects is not considered. This application also introduces a pseudo-optical flow feature layer to extract dynamic features, improving recognition accuracy while avoiding the enormous computational burden required to calculate optical flow. The pseudo-optical flow information is dynamic information calculated based on the residual between the feature representations of past frames and the feature representation of the current frame.

[0079] In an optional embodiment, when the convolutional layer includes the pseudo-optical flow feature layer, determining the behavior vector of the current frame includes: extracting a displacement feature sequence based on the static feature map sequence using the pseudo-optical flow feature layer; when the convolutional layer only includes the pseudo-optical flow feature layer, using the displacement feature sequence as the dynamic feature map sequence; when the convolutional layer includes the three-dimensional feature layer and the pseudo-optical flow feature layer, concatenating the static map sequence and the displacement feature sequence, and inputting them into the three-dimensional feature layer to obtain the dynamic feature map sequence.

[0080] Specifically, when the convolutional layer contains only pseudo-optical flow feature layers, since the pseudo-optical flow information is used to characterize the change of the spatial position of the target object within adjacent image frames over time, it needs to be based on a static feature map sequence. First, a static feature map sequence is obtained through a fourth predetermined number of two-dimensional feature layers. Then, inter-frame feature displacement information is extracted through a fifth predetermined number of pseudo-optical flow feature layers, and this is used as a dynamic feature map sequence.

[0081] Although the displacement information extracted by the pseudo-optical flow feature layer is a dynamic feature, extracting only displacement information is too one-sided. Therefore, this application embodiment introduces a three-dimensional feature layer to obtain a dynamic feature map sequence, which can make full use of the feature information of the image sequence and improve the effect and accuracy of subsequent recognition processing. For example, when the above convolutional layer includes the above three-dimensional feature layer and the pseudo-optical flow feature layer, the seventh predetermined number of pseudo-optical flow feature layers is located between the eighth predetermined number of two-dimensional feature layers and the sixth predetermined number of three-dimensional feature layers. After the above displacement feature map sequence and the above displacement feature sequence are concatenated, they are input into the above three-dimensional feature layer to obtain the above dynamic feature map sequence.

[0082] In one optional embodiment, the displacement feature sequence is extracted through a pseudo optical flow feature layer, including: filtering out a target static feature map from the static feature map sequence; obtaining the displacement vector of each feature point in the target static feature map, determining multiple displacement vectors based on multiple feature points, and assembling all displacement vectors into a displacement vector map; and assembling a displacement vector sequence by using each frame in the static feature map sequence as the corresponding displacement vector map determined by the target static feature map.

[0083] Specifically, pseudo-optical flow information can be used to represent the displacement between each pixel in the previous frame and the corresponding pixel in the next frame after the latter has moved. Each frame's displacement vector map contains the displacement vector of each feature point in the current image frame relative to the corresponding feature point in the previous image frame. Therefore, each image frame in the image sequence corresponds to one frame's displacement vector map, and all displacement vector maps form a displacement vector sequence. In addition, the initial displacement vector map can be preset.

[0084] In an optional embodiment, in the target static feature map, the displacement vector of each feature point is obtained, and multiple displacement vectors are determined based on multiple feature points, including: for a first feature point in the target static feature map, in the previous frame static feature map, a second feature point with the greatest similarity to the first feature point is found in the neighborhood of the same position of the first feature point, and the displacement vector is calculated based on the position of the second feature point and the position of the first feature point.

[0085] Specifically, the size of the aforementioned neighborhood is not limited and can be determined based on the actual computational accuracy. For example, the process of pseudo-optical flow layer processing is illustrated using a 7x7 neighborhood:

[0086] Define xt (i, j) is the feature vector of the t-th frame of the image sequence with coordinates (i, j). The displacement calculation process relative to the previous frame is as follows:

[0087] 1) Calculate the features within the 7x7 neighborhood of position (i, j) in the previous frame, i.e., frame t-1, and x. t Calculate the vector product of (i, j):

[0088] O t (i, j, m, n) = x t (i, j)*x t-1 (i+m, j+n), -3≤m, n≤3

[0089] 2) Find the values ​​of m and n that maximize the vector product:

[0090] d t (i, j) = argmax O t (i, j, m, n)

[0091] Where d t (i, j) is a two-dimensional vector. The values ​​of m and n, which have the largest vector product, are the coordinate offsets with the highest similarity in the neighborhood. After calculating the displacement for each position in the t-th frame, we finally obtain a displacement vector of h*w*2, where h and w are the height and width of the input feature vector.

[0092] Furthermore, combining the aforementioned feature map sequence and the aforementioned dynamic feature map sequence, the fully connected layer of the first model is used to perform convolutional operations to achieve classification calculations, mapping feature information to the labeled sample space, realizing end-to-end real-time recognition of driving status, and obtaining the behavior vector of the current frame. The aforementioned behavior vector includes at least one of the following: facial expression vector, body behavior vector, and interaction behavior vector. Specifically, the dimension of the behavior vector is determined by the number of types of early fatigue behaviors that the model supports detecting, and can be all or some of the behavior types included in the aforementioned three categories. The value of each dimension indicates whether the action was not recognized or was recognized.

[0093] Furthermore, before each model performs recognition, the above method also includes: multi-frame alignment of the input image. Multi-frame image alignment is performed based on the position of the target object, using the same alignment parameters for different frames to ensure image continuity, remove redundant information from the image sequence, and improve recognition accuracy.

[0094] For example, based on the facial feature point positions output by the face detection algorithm during valid frame discrimination, the center point coordinates and face width of the face in each image of the image sequence are calculated. Then, the average center point coordinates and average face width of all images are calculated. Using the average center point coordinates as the center and the face width as the length and width, a specified area of ​​each frame image is extracted and scaled to a fixed size, ultimately obtaining an image sequence of a preset length and fixed size. This ensures the continuity of the image, removes redundant information in the image sequence, and improves recognition accuracy.

[0095] This application's embodiments improve recognition accuracy by performing multi-frame alignment on image sequences to remove redundant information. A combination of two-dimensional and three-dimensional feature layers is used to extract static and dynamic features, reducing computational load while utilizing complete image feature information. Furthermore, a pseudo-optical flow convolutional layer is introduced, improving recognition accuracy while avoiding the enormous computational burden required to calculate optical flow. By performing behavior recognition on multiple frames, recognition accuracy is significantly improved, fully utilizing the dynamic information of the behavior.

[0096] S102, Statistical analysis is performed on multiple behavior vectors within the first preset time period to obtain statistical results;

[0097] Although a single behavior vector is identified based on a sequence of multiple image frames, judging the state based on a single behavior vector can lead to misjudgments due to the randomness of the target object's actions. In this embodiment, the driver's behavior within a first preset time period is statistically analyzed, and early fatigue detection is performed based on the statistical results, which is beneficial to the stability and accuracy of the detection results.

[0098] In an optional embodiment, before statistically analyzing the multiple behavioral vectors within a first preset time period to obtain the statistical results, the following steps are included:

[0099] Obtaining multiple behavior vectors within the first preset time period includes:

[0100] A queue structure is used to cache the behavior vector of each frame in the latest first preset time and maintain the queue. The queue is used as the multiple behavior vectors. The behavior vector at the end of the queue is the behavior vector of the current frame. The difference between the timestamps of the behavior vector at the head of the queue and the behavior vector at the end of the queue is the first preset time.

[0101] Specifically, the behavior vector output by the first model includes recognition results and timestamps for each dimension. In this embodiment, a queue structure is used to cache and maintain the behavior vectors in real time, retaining them within a first preset time period. That is, the difference between the timestamps of the behavior vectors at the head and tail of the queue is the aforementioned first preset time. If the timestamp span of the cached behavior vectors in the queue is less than the aforementioned first preset time, the latest behavior vector output from the first model is cached. If the timestamp span of the cached behavior vectors in the queue is greater than or equal to the first preset time, the historical data cached before the first preset time is cleared, starting from the latest behavior vector, and the cleared data is statistically analyzed to ensure the real-time performance of the recognition. Furthermore, the first preset time can be adjusted according to the actual detection accuracy.

[0102] In one optional embodiment, statistical analysis is performed on multiple of the aforementioned behavior vectors within a first preset time period to obtain statistical results, including:

[0103] For multiple behavior vectors within the first preset time period, a sliding window is used to count the behavior vectors that meet the first preset condition, and the statistical results are obtained. The first preset condition includes: the behavior vector is triggered when the duty cycle of the behavior vector within the sliding window is greater than a first threshold; and the behavior vector is triggered when the duty cycle of the behavior vector within the sliding window is less than a second threshold.

[0104] In one optional embodiment, the above statistical results include at least one of the following: number of triggers, average trigger duration, maximum trigger duration, and total trigger duration.

[0105] The value of each dimension in the behavior vector indicates whether the action was not detected or was detected. For each type of behavior, this embodiment determines the start and end of the triggering of a specified behavior by the duty cycle of the specified behavior within a sliding window, avoiding misjudgments caused by judging a behavior trigger based on a single dimension value. The sliding window contains a behavior vector of a preset length. For a specified behavior, the triggering of the specified behavior vector begins when the duty cycle of the corresponding dimension value in the preset-length behavior vector within the sliding window is greater than a first threshold; the triggering of the behavior vector ends when the duty cycle of the corresponding dimension value in the preset-length behavior vector within the sliding window is less than a second threshold. The number of triggers, average trigger duration, maximum trigger duration, and total trigger duration for each behavior within a first preset time period are calculated using this method. The statistical results include at least one of the following.

[0106] For example, the preset length of the sliding window is W. For a specific behavior in early fatigue behavior, frowning, statistical analysis is performed. The frowning dimension value is 0 or 1, where 0 indicates the action was not detected and 1 indicates it was detected. For all multiple behavior vectors within a first preset time period, the frowning dimension is statistically analyzed using a sliding window. Specifically, the frowning behavior is triggered when the duty cycle of the frowning dimension value within the sliding window is greater than a preset value R1; the frowning behavior ends when the duty cycle of the frowning dimension value within the sliding window is less than a preset value R2. The span between the start and end times of the frowning trigger behavior is the frowning trigger duration. Furthermore, based on the triggered behavior, the number of frowning triggers, average trigger duration, maximum trigger duration, and total trigger duration within the first preset time period can be statistically analyzed. The statistical methods for other early fatigue behaviors are as described above and will not be repeated.

[0107] In this embodiment of the application, after statistically analyzing the behavior information of the first driver, a statistical result vector of 0xP will be output, where O is the number of early fatigue behaviors included in the statistics, and P is the number of statistical result categories. As mentioned above, the types of early fatigue behaviors include: frowning, raising eyebrows, tightening lips, adjusting posture, stretching, rubbing face, rubbing nose, rubbing eyes, twisting neck, smoking, drinking water, and making phone calls. The total statistical result categories include: number of triggers, average trigger duration, maximum trigger duration, and total trigger duration.

[0108] In this embodiment, by statistically analyzing the driver's behavior within a first preset time period, and by providing information from different dimensions such as facial expressions, body language, and human interaction, the system can make accurate judgments for early fatigue warning, which is beneficial to the stability and accuracy of the detection results.

[0109] S104. The trained second model is used to classify the statistical results to obtain the early fatigue detection results.

[0110] In one alternative embodiment, the early fatigue detection results include: no potential early fatigue risk, and potential early fatigue risk exists.

[0111] The trained second model is used for classification based on statistical results. The input data to the second model is the statistical results, for example, an 0xP statistical result vector as mentioned above. The second model will output the early fatigue detection results within the current first preset time period, including whether there is no potential early fatigue risk or whether there is a potential early fatigue risk. Using the long-term behavioral characteristics statistical results of early fatigue-related behaviors to make decisions on detection results can provide an objective assessment of the driver's driving process and play a good early warning role for unsafe events in traffic scenarios.

[0112] In one optional embodiment, the training method for the second model includes: constructing a training dataset by collecting sample data from users; training an initial second model using the training dataset to obtain the trained second model.

[0113] refer to Figure 4 A flowchart for an optional method of constructing training data according to an embodiment of the present invention is provided. Figure 4 As shown, the training dataset is constructed by collecting sample data from users, including the following steps:

[0114] S400, the above sample data is segmented according to the duration of the first preset time to obtain multiple sample data segments, wherein the types of the multiple sample data segments include the sample data segments that the user determines are fatigued and the sample data segments that the user determines are not fatigued.

[0115] To obtain the logical relationship between behavior and early fatigue, the driving processes of multiple individuals were recorded, acquiring driving process footage (i.e., sample data). The duration of each person's driving process footage was a second preset time. There were no specific limitations on the number of users collected or the duration of the footage; to enrich the sample data, the second preset time for footage collection was longer than the first preset time required for the first model's judgment. The sample data was then segmented, scored, and filtered according to the first preset time duration to obtain data segments containing user-determined fatigue and user-determined non-fatigue data. Labeling the actual collected driving footage with fatigue tags helps to accurately obtain the logical relationship between behavior and early fatigue in subsequent stages.

[0116] For example, the driving processes of 537 people were recorded, with an average driving time of 12 hours per person. The first preset time required for statistical analysis was half an hour. For each person's driving footage, drivers were asked to rate their fatigue level every half hour, with 0 indicating no fatigue, 1 indicating fatigue, and 2 indicating no assessment. The 12 hours of sample data were segmented, with each half-hour video segment serving as a sample data segment. Based on the drivers' fatigue level ratings, the segments where fatigue could not be determined were removed. The remaining segments were divided into two groups: no fatigue and fatigue, which were used as training samples for the second model.

[0117] S401, Select the target sample data segment from the above multiple sample data segments, identify the target sample data segment through the above first model, obtain the sample behavior vector sequence corresponding to the target sample data segment, count the sample behavior vector sequence, and obtain the sample statistical results of the target sample data segment.

[0118] S402, take each of the above multiple sample data segments as the above target sample data segment, obtain the corresponding sample statistical results, and construct the above training dataset from the multiple sample data segments containing the sample statistical results.

[0119] In one optional embodiment, the above-mentioned sample behavior vector sequence is composed of sample behavior vectors, wherein the above-mentioned sample behavior vectors include the following: sample facial expression vectors, sample body behavior vectors, and sample interaction behavior vectors.

[0120] In one optional embodiment, the above sample statistics include the following: number of triggers, average trigger duration, maximum trigger duration, and total trigger duration.

[0121] As shown in S400, for each of the multiple sample data segments, its duration span is a first preset time. For each sample data segment, behavior recognition and statistics are performed using the first model to obtain sample statistical results. The specific steps are the same as those described in S100 and S102 above, and will not be repeated here.

[0122] For example, the first model identifies all early fatigue behaviors in each sample data segment. Each sample data segment corresponds to a sample behavior vector sequence, which contains the identification results of all early fatigue behaviors in all image frames within a first preset time period. The identification results include two types: the action was not identified and the action was identified. Further, each sample behavior vector consists of a sample facial expression vector, a sample limb behavior vector, and a sample interaction behavior vector. Statistical results are obtained for each sample behavior vector sequence. A sample statistical result includes the number of triggers, the average trigger duration, the maximum trigger duration, and the total trigger duration. The definition of trigger is the same as above and will not be repeated. Therefore, a 12×4 statistical result and a corresponding fatigue label can be obtained for each sample data segment. All sample data segments containing the sample statistical results and corresponding fatigue labels are used to construct the above training dataset.

[0123] Furthermore, the second model is not limited to decision trees; it can also use deep learning models such as neural networks, machine learning models such as SVM and random forests, with the corresponding learning strategies adaptively adjusted according to the different models.

[0124] Furthermore, since individual behavioral patterns differ, using the same second model for decision-making could lead to incorrect assessments of early fatigue. This embodiment also includes a personalization unit. When a driver reports an incorrect recognition result, the personalization module dynamically adjusts the model parameters of the second model to achieve personalized decision-making.

[0125] This application embodiment identifies behaviors corresponding to early fatigue, extracts, statistically analyzes, and statistically studies long-term behaviors during driving, and then makes decisions about early fatigue states based on these behavioral characteristics, providing a final prediction result. This application embodiment can provide early warnings when drivers show signs of fatigue, giving drivers more time to adjust and further reducing the risk of fatigued driving.

[0126] This application also provides a vision-based early fatigue detection device, with reference to... Figure 5 This is a block diagram of a vision-based early fatigue detection device according to an embodiment of this application, such as... Figure 5 As shown, the device 50 includes: an identification unit 51, a statistical unit 52, and a classification unit 53, wherein,

[0127] The recognition unit 51 is used to recognize the image sequence through the first model and determine the behavior vector of the current frame, wherein the current frame is the last frame in the image sequence;

[0128] The statistical unit 52 is used to perform statistics on multiple behavior vectors within a first preset time period to obtain statistical results;

[0129] Classification unit 53 is used to classify the statistical results using the trained second model to obtain early fatigue detection results.

[0130] In this implementation, the identification unit 51 is used to identify the image sequence using a first model and determine the behavior vector of the current frame, wherein the current frame is the last frame in the image sequence; the statistics unit 52 is used to perform statistics on multiple behavior vectors within a first preset time period to obtain statistical results; and the classification unit 53 is used to classify the statistical results using a trained second model to obtain early fatigue detection results. In this embodiment, the problem of not being able to issue early warnings for fatigue behavior can be solved, thus giving users more time to adjust and further reducing the risk of fatigued driving.

[0131] In this embodiment, a first model is responsible for identifying the input image sequence and determining whether there are driver behaviors corresponding to early fatigue states. Existing technologies detect fatigue behaviors such as drowsiness, yawning, and looking down, which indicate that the driver is already fatigued and cannot provide early warnings to reduce the risk of fatigued driving.

[0132] In this embodiment, the behavior when fatigue tends to occur is used as the criterion for judgment, and early fatigue behavior is divided into three categories. The first category is facial expression behavior, including at least one of the following: frowning, raising eyebrows, and tightening lips; the second category is physical behavior, including at least one of the following: adjusting sitting posture, stretching, rubbing face, rubbing nose, rubbing eyes, and twisting neck; the third category is common human-object interaction behavior, including at least one of the following: smoking, drinking water, and making phone calls.

[0133] In one alternative embodiment, the behavior vectors include at least one of the following: facial expression vectors, body behavior vectors, and interaction behavior vectors.

[0134] In an optional embodiment, the recognition unit includes: an extraction module, configured to extract a static feature map sequence and a dynamic feature map sequence of the image sequence through a convolutional layer of the first model, wherein the convolutional layer includes a three-dimensional feature layer and / or a pseudo optical flow feature layer; and a classification module, configured to combine the static feature map sequence and the dynamic feature map sequence, and perform classification calculation using a fully connected layer of the first model to obtain the behavior vector of the current frame.

[0135] In an optional embodiment, the extraction module includes: a first extraction submodule, used to extract a displacement feature sequence based on the static feature map sequence through the pseudo optical flow feature layer; a second extraction submodule, used to take the displacement feature sequence as the dynamic feature map sequence when the convolutional layer only contains the pseudo optical flow feature layer; and a third extraction submodule, used to concatenate the static map sequence and the displacement feature sequence and input them into the three-dimensional feature layer to obtain the dynamic feature map sequence when the convolutional layer contains the three-dimensional feature layer and the pseudo optical flow feature layer.

[0136] In an optional embodiment, the extraction module includes: a fourth extraction submodule, used to extract both static feature map sequences and dynamic feature map sequences simultaneously through the three-dimensional feature layer when the convolutional layer contains only a three-dimensional feature layer; and a fifth extraction submodule, used to extract static feature map sequences through the two-dimensional feature layer and extract dynamic feature map sequences based on the static feature map sequences through the three-dimensional feature layer when the convolutional layer contains both a two-dimensional feature layer and a three-dimensional feature layer.

[0137] In an optional embodiment, the statistics unit includes: a statistics submodule, used to perform sliding window statistics on multiple behavior vectors within the first preset time period to obtain the statistical results, wherein the first preset condition includes: the behavior vector triggering starts when the duty cycle of the behavior vector within the sliding window is greater than a first threshold; and the behavior vector triggering ends when the duty cycle of the behavior vector within the sliding window is less than a second threshold; the statistical results include at least one of the following: number of triggers, average trigger duration, maximum trigger duration, and total trigger duration.

[0138] In one optional embodiment, the classification unit includes: a first training module for constructing a training dataset using sample data collected from users; and a second training module for training an initial second model using the training dataset to obtain the trained second model.

[0139] In an optional embodiment, the first training module includes: a first training submodule, configured to segment the sample data according to the duration of the first preset time to obtain multiple sample data segments, wherein the types of the multiple sample data segments include sample data segments that the user determines are fatigued and sample data segments that the user determines are not fatigued; a second training submodule, configured to select target sample data segments from the multiple sample data segments, identify the target sample data segments through the first model, obtain the sample behavior vector sequence corresponding to the target sample data segments, statistically analyze the sample behavior vector sequence, and obtain the sample statistical results of the target sample data segments; and a third training submodule, configured to use each of the multiple sample data segments as the target sample data segment, obtain the corresponding sample statistical results, and construct the multiple sample data segments containing the sample statistical results into the training dataset.

[0140] In an optional embodiment, the device 50 further includes a filtering unit, which includes:

[0141] The sampling module is used to filter valid frames from the input data, wherein the valid frames contain the target object and the time interval between two adjacent valid frames is greater than a first interval threshold.

[0142] The combination module is used to take the latest timestamped valid frame as the current frame, and combine the current frame and the consecutive valid frames before it to form the image sequence, wherein the length of the image sequence is a preset length.

[0143] The embodiments of this application can ensure that the composed image sequence is updated in real time, contains complete information required for behavior recognition, and achieves uniform sampling of input data, removes redundant information to reduce computation, and uses the features of input data efficiently and reasonably, laying the foundation for subsequent real-time and accurate fatigue detection.

[0144] Furthermore, since individual behavioral patterns differ, using the same second model for decision-making could lead to incorrect assessments of early fatigue. This embodiment also includes a personalization unit. When a driver reports an incorrect recognition result, the personalization module dynamically adjusts the model parameters of the second model to achieve personalized decision-making.

[0145] This application embodiment identifies behaviors corresponding to early fatigue, extracts, statistically analyzes, and statistically studies long-term behaviors during driving, and then makes decisions about early fatigue states based on these behavioral characteristics, providing a final prediction result. This application embodiment can provide early warnings when drivers show signs of fatigue, giving drivers more time to adjust and further reducing the risk of fatigued driving.

[0146] According to another aspect of the embodiments of the present invention, the present application also provides a computer-readable storage medium storing one or more programs, which can be executed by one or more processors to implement any of the above-described vision-based early fatigue detection methods.

[0147] According to another aspect of the present invention, this application also provides a processing apparatus, comprising: a memory for storing computer-executable instructions; and a processor for executing the computer-executable instructions to implement the steps of the early fatigue detection method described in any one of the above embodiments.

[0148] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0149] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0150] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

Claims

1. A vision-based method for early fatigue detection, comprising: The first model is used to identify the image sequence and determine the behavior vector of the current frame, wherein the current frame is the last frame in the image sequence, and the behavior vector includes: facial expression vector, body behavior vector and interaction behavior vector; Statistical analysis is performed on multiple behavior vectors within a first preset time period to obtain statistical results; The statistical results are classified using the trained second model to obtain early fatigue detection results; The step of identifying the image sequence using the first model and determining the behavior vector of the current frame includes: The static feature map sequence and dynamic feature map sequence of the image sequence are extracted through the convolutional layer of the first model, wherein the convolutional layer includes a three-dimensional feature layer and / or a pseudo optical flow feature layer; By combining the static feature map sequence and the dynamic feature map sequence, classification calculation is performed using the fully connected layer of the first model to obtain the behavior vector of the current frame; When the convolutional layer includes the pseudo-optical flow feature layer, determining the behavior vector of the current frame includes: Based on the static feature map sequence, a displacement feature sequence is extracted through the pseudo optical flow feature layer; When the convolutional layer contains only the pseudo optical flow feature layer, the displacement feature sequence is used as the dynamic feature map sequence; When the convolutional layer includes the three-dimensional feature layer and the pseudo-optical flow feature layer, the static image sequence and the displacement feature sequence are concatenated and then input into the three-dimensional feature layer to obtain the dynamic feature image sequence.

2. The detection method according to claim 1, characterized in that, Obtaining the image sequence includes: Valid frames are filtered from the input data, wherein the valid frames contain the target object and the time interval between two adjacent valid frames is greater than a first interval threshold. The latest valid frame with the latest timestamp is taken as the current frame, and the current frame and the consecutive valid frames preceding it are combined to form the image sequence, wherein the length of the image sequence is a preset length.

3. The detection method according to claim 1, characterized in that, The step of extracting the displacement feature sequence through the pseudo-optical flow feature layer includes: Target static feature maps are selected from the static feature map sequence; In the target static feature map, the displacement vector of each feature point is obtained, multiple displacement vectors are determined based on multiple feature points, and all displacement vectors are combined to form a displacement vector map. Each frame in the static feature map sequence is used as the corresponding displacement vector map determined by the target static feature map to form the displacement vector sequence.

4. The detection method according to claim 3, characterized in that, In the target static feature map, the displacement vector of each feature point is obtained, including: For the first feature point in the target static feature map, in the previous frame static feature map, find the second feature point with the greatest similarity to the first feature point in the neighborhood of the same position of the first feature point, and calculate the displacement vector based on the position of the second feature point and the position of the first feature point.

5. The detection method according to claim 1, characterized in that, When the convolutional layer contains the three-dimensional feature layer, determining the behavior vector of the current frame includes: When the convolutional layer contains only a three-dimensional feature layer, the static feature map sequence and the dynamic feature map sequence are extracted simultaneously through the three-dimensional feature layer. When the convolutional layer contains a two-dimensional feature layer and a three-dimensional feature layer, the static feature map sequence is extracted through the two-dimensional feature layer, and the dynamic feature map sequence is extracted through the three-dimensional feature layer based on the static feature map sequence.

6. The detection method according to claim 1, characterized in that, Before obtaining the statistical results by statistically analyzing multiple behavior vectors within a first preset time period, the following steps are included: Obtaining multiple behavior vectors within the first preset time period includes: A queue structure is used to cache the behavior vector of each frame in the latest first preset time and maintain the queue. The queue is used as the plurality of behavior vectors, wherein the behavior vector corresponding to the end of the queue is the behavior vector of the current frame, and the difference between the timestamp of the behavior vector corresponding to the head of the queue and the behavior vector corresponding to the end of the queue is the first preset time.

7. The detection method according to claim 1, characterized in that, The step of statistically analyzing multiple behavior vectors within a first preset time period to obtain statistical results includes: For multiple behavior vectors within the first preset time period, a sliding window is used to statistically analyze the behavior vectors that satisfy a first preset condition, and the statistical results are obtained. The first preset condition includes: the behavior vector triggering starts when the duty cycle of the behavior vector within the sliding window is greater than a first threshold; and the behavior vector triggering ends when the duty cycle of the behavior vector within the sliding window is less than a second threshold. The statistical results include at least one of the following: number of triggers, average trigger duration, maximum trigger duration, and total trigger duration.

8. The detection method according to claim 1, characterized in that, The early fatigue detection results include: There is no potential risk of early fatigue, but there is a potential risk of early fatigue.

9. The detection method according to claim 1, characterized in that, The training methods for the second model include: A training dataset is constructed using sample data collected from users; The initial second model is trained using the training dataset, and the trained second model is obtained.

10. The detection method according to claim 9, characterized in that, The construction of the training dataset from user-collected sample data includes: The sample data is segmented according to the duration of the first preset time to obtain multiple sample data segments, wherein the types of the multiple sample data segments include sample data segments that the user determines are fatigued and sample data segments that the user determines are not fatigued. The target sample data segment is selected from the multiple sample data segments, the target sample data segment is identified by the first model, the sample behavior vector sequence corresponding to the target sample data segment is obtained, the sample behavior vector sequence is statistically analyzed, and the sample statistical result of the target sample data segment is obtained. Each of the multiple sample data segments is taken as the target sample data segment, and the corresponding sample statistical results are obtained. The multiple sample data segments containing the sample statistical results are then used to construct the training dataset.

11. The detection method according to claim 10, characterized in that, The sample behavior vector sequence is composed of sample behavior vectors, wherein the sample behavior vectors include the following: sample facial expression vectors, sample body behavior vectors, and sample interaction behavior vectors.

12. The detection method according to claim 10, characterized in that, The sample statistics include the following: number of triggers, average trigger duration, maximum trigger duration, and total trigger duration.

13. A vision-based early fatigue detection device, the device comprising: The recognition unit is used to recognize the image sequence through the first model and determine the behavior vector of the current frame, wherein the current frame is the last frame in the image sequence, and the behavior vector includes: facial expression vector, body behavior vector and interaction behavior vector; The statistical unit is used to perform statistics on multiple behavior vectors within a first preset time period to obtain statistical results; A classification unit is used to classify the statistical results using a trained second model to obtain early fatigue detection results; The recognition unit identifies the image sequence using a first model to determine the behavior vector of the current frame, specifically: The static feature map sequence and dynamic feature map sequence of the image sequence are extracted through the convolutional layer of the first model, wherein the convolutional layer includes a three-dimensional feature layer and / or a pseudo optical flow feature layer; By combining the static feature map sequence and the dynamic feature map sequence, classification calculation is performed using the fully connected layer of the first model to obtain the behavior vector of the current frame; When the convolutional layer includes the pseudo-optical flow feature layer, determining the behavior vector of the current frame includes: Based on the static feature map sequence, a displacement feature sequence is extracted through the pseudo optical flow feature layer; When the convolutional layer contains only the pseudo optical flow feature layer, the displacement feature sequence is used as the dynamic feature map sequence; When the convolutional layer includes the three-dimensional feature layer and the pseudo-optical flow feature layer, the static image sequence and the displacement feature sequence are concatenated and then input into the three-dimensional feature layer to obtain the dynamic feature image sequence.

14. A computer-readable storage medium storing one or more... The method may include one or more programs, which may be executed by one or more processors to implement the early fatigue detection method as described in any one of claims 1 to 12.

15. A processing apparatus, comprising: Memory is used to store executable instructions for a computer; A processor for executing the computer-executable instructions to implement the steps of the early fatigue detection method as described in any one of claims 1 to 12.

Citation Information

Patent Citations

  • Early fatigue detection method and system based on fine eye movement features

    CN112434611A

  • Micro-expression recognition method and system based on deep learning

    CN114882553A