System for measuring fatigue using video analysis

A deep learning-based fatigue measurement system analyzes key frames from a person's face video using multiple neural networks and an ensemble algorithm, addressing the limitations of existing methods by offering rapid and accurate fatigue assessment.

WO2025183241A1PCT designated stage Publication Date: 2025-09-04VAIV CO INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/002620
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-29
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Existing fatigue measurement methods are subjective, time-consuming, and unsuitable for rapid assessment in high-stress environments, lacking objective and accurate measures for fatigue levels in individuals, particularly for professionals requiring high concentration and quick decision-making.

Method used

A fatigue measurement system using deep learning algorithms to analyze a video of a person's face, employing multiple neural network models and an ensemble algorithm to classify and synthesize fatigue levels, focusing on key frames for efficient analysis.

Benefits of technology

Provides quick and accurate objective fatigue measurement by analyzing key frames, reducing computational resources and time, while ensuring precise determination of fatigue levels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024002620_04092025_PF_FP_ABST
    Figure KR2024002620_04092025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a system for measuring fatigue using video analysis. More specifically, the present invention relates to an information system for determining a fatigue measurement level on the basis of video analysis which is capable of determining a fatigue level by analyzing, with a deep learning algorithm, a video obtained by imaging the face of a person, and to a method for determining the fatigue level of the person by analyzing 'the video obtained by imaging the face of the person for a certain time'. The present invention allows a video obtained by imaging a subject to be decomposed into frame units, and then fatigue levels to be classified by neural networks using three different models such as an object feature analysis model, an object-CNN model, and a CNN model. Then, the present invention allows an ensemble algorithm to be applied to each classification result to combine same, and the result to be determined as a fatigue level for the subject.
Need to check novelty before this filing date? Find Prior Art

Description

Image analysis-based fatigue measurement system

[0001] The present invention relates to a fatigue measurement system based on image analysis. More specifically, it is an information system for determining a fatigue measurement level based on image analysis, which can determine a fatigue level by analyzing a video of a person's face using a deep learning algorithm, and relates to a method for determining a person's fatigue level by analyzing a 'video of a person's face captured over a certain period of time'. The present invention decomposes a video of a subject into frames, and then classifies the fatigue level using a neural network using three different models, including an object feature analysis model, an object-CNN model, and a CNN model. Then, an ensemble algorithm is applied to each classification result to synthesize it, and the result is determined as the fatigue level of the subject.

[0002] People living in modern society interact with each other in both real and virtual spaces, leading to busy, uninterrupted days without sufficient rest. Consequently, fatigue is inevitable for modern people, and its causes and forms are as diverse as the environments and interactions people experience. In particular, modern people's lifestyles and work patterns have a significant impact. These factors include physiological characteristics, sleep disorders, individual lifestyles, stress, living environments, and various activities for maintaining health. Among these, sleep disorders are the most significant cause of fatigue. Sleep duration impacts performance abilities such as concentration, decision-making, and reaction time, and a decline in these abilities can manifest as fatigue.

[0003] Meanwhile, for professionals who require a high level of concentration when performing their duties, such as pilots of aircraft or fighter jets, or operators of precision weapon systems, or for workers in high-risk occupations such as crew members of ships or long-distance drivers who must work for long hours while also making quick and accurate judgments in the event of an emergency, if they are deployed to work or combat while highly fatigued, their cognitive functions will decline, which will increase the likelihood of safety accidents. In addition, in the case of combat personnel, this will have a significant impact on reducing combat capabilities. Therefore, it is necessary to strictly operate a system that allows people with high fatigue to be identified in advance and excluded from work.

[0004] Therefore, a method or system capable of objectively measuring fatigue in individual workers or combat personnel is needed. However, the concept of fatigue is not clearly defined for humans, and even individuals themselves cannot easily determine whether they are experiencing high levels of fatigue. Therefore, objectively identifying individuals with high levels of fatigue is challenging.

[0005] The most objective conceptual distinction regarding fatigue is between subjective fatigue, which can be considered central or mental fatigue, and peripheral fatigue, which is defined as task-related fatigue based on maintaining an appropriate level of physical arousal for performing a task. Various versions of self-report assessment tools, often used as clinical scales, have been developed to measure subjective fatigue. Furthermore, existing methods for assessing fatigue levels include questionnaires, biochemical and physiological tests, and reaction time tests.

[0006] The questionnaire method has shown statistically significant results in the development of fatigue measurement tools based on the contents of the Multidimensional Fatigue Inventory (MFI). There are also methods such as the measurement and evaluation of circadian rhythm disturbances using actigraphy, the ECG (ElectrocardioGraphy) method that measures using indicators such as heart rate and pulse, the Ecological Momentary Assessment (EMA) method that uses indicators such as skin temperature, and physiological methods such as collecting blood, saliva, etc.

[0007] However, since these conventional measurement methods are primarily aimed at assessing the degree of fatigue from a disease perspective, while they are useful for identifying the causes and extent of pathological fatigue, they are limited in measuring the degree of fatigue that manifests in various causes and forms. Furthermore, fatigue measurement methods rely on subjective judgments and are subject to variation depending on individual conditions and circumstances, making them difficult to use as an absolute standard for assessing fatigue. Furthermore, survey-based fatigue measurement methods are time-consuming and expensive, and many factors, such as diet, exercise, and infection, can affect fatigue measurements.

[0008] Meanwhile, advancements in machine learning and artificial intelligence (AI) technologies are driving ongoing efforts to analyze vast amounts of data and utilize them in diverse fields. In particular, advancements in image analysis technology in the field of computer vision are yielding analysis results that reach the level of human judgment in certain fields. Image analysis technologies, such as behavioral analysis, measure human states and behaviors, effectively analyzing subsequent actions and emotions.

[0009] Recent research is gathering fatigue-related data and training deep learning models to measure fatigue. Representative examples include measuring fatigue by collecting body movement data through sensors, or studying changes in pupil and iris morphology, such as blinking, by continuously filming drivers for extended periods of time. However, these methods rely on video recording of subjects over extended periods of time, making them unsuitable for applications such as pre-commissioning fatigue assessments for fighter pilots.

[0010] Therefore, the present invention proposes a system capable of measuring the fatigue state of a subject by analyzing a video of the subject's face taken in a relatively short time using a model based on a deep learning-based image analysis technique, and also provides methods for extracting and using only key frames having meaningful images among the frame images constituting the video in order to efficiently utilize computer resources and shorten analysis time according to video analysis.

[0011] The present invention, which was created to solve the above-described problems, aims to provide a system that quickly and accurately determines a person's fatigue level by analyzing a video of a person's face using a deep learning algorithm.

[0012] Another object of the present invention is to provide a system for measuring fatigue level using a video of a subject's face taken for a relatively short period of time, rather than using a video of the subject taken for a long period of time.

[0013] Another object of the present invention is to provide a system for measuring fatigue levels more accurately by classifying fatigue levels using various neural network models and then applying an ensemble algorithm to the classification results.

[0014] Another object of the present invention is to provide a system for measuring fatigue by extracting only meaningful frame images that can be used for fatigue measurement from among many frame images included in a video, in order to efficiently utilize computer resources and shorten analysis time according to video analysis.

[0015] The technical problems to be solved by the present invention are not limited to the technical problems mentioned above, and other technical problems not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present invention belongs from the description below.

[0016] In order to achieve the above object, the present invention provides a fatigue measurement system based on image analysis, which is installed in an information system and can determine a fatigue level by analyzing a video of a person's face using a deep learning algorithm, the system comprising: a preprocessing module that receives a video including a subject's face as an analysis target video, decomposes the analysis target video into frame units, and then selects an analysis target frame from among the frames; a first classification module that classifies the fatigue level for each of the analysis target frames into a first fatigue level using an object feature analysis model; a second classification module that classifies the fatigue level for each of the analysis target frames into a second fatigue level using an object-CNN model; and a third classification module that classifies the fatigue level for each of the analysis target frames into a CNN model and classifies it into a third fatigue level; And an ensemble module that calculates a final fatigue level for each of the analysis target frames by applying an ensemble algorithm to the first fatigue level, the second fatigue level, and the third fatigue level; wherein the ensemble module determines the fatigue level that is most 'appearing' (high?) ​​among the average of the final fatigue levels for all of the analysis target frames or the final fatigue level for each of the analysis target frames as the fatigue level for the subject, it is preferable that the image analysis-based fatigue measurement system be configured.

[0017] The present invention also preferably comprises an image analysis-based fatigue measurement system characterized in that, in addition to the above-described features, the first analysis module: - detects a face region of the subject in the analysis target frame using an object recognition model, - extracts a predetermined number of feature point coordinates (x, y) from the face region using a feature point coordinate detection model, - classifies the fatigue level for each of the analysis target frames by comparing the feature point coordinates with a true fatigue value using a classification algorithm-based fatigue measurement model, and classifies this as the first fatigue level.

[0018] The present invention is also preferably an image analysis-based fatigue measurement system characterized in that, in addition to the above-described features, the second analysis module - detects the face region of the subject in the analysis target frame using an object recognition model, - extracts a feature element in vector format from the face region using a CNN, - classifies the fatigue level for each of the analysis target frames by sending the feature element to a fully connected layer within the CNN and classifies it as the second fatigue level.

[0019] The present invention is also preferably an image analysis-based fatigue measurement system characterized in that, in addition to the above-described features, the third analysis module extracts a feature element in vector format from each of the analysis target frames using CNN, and sends the feature element to a fully connected layer within the CNN to classify the fatigue level for each of the analysis target frames and classifies it as the third fatigue level.

[0020] The present invention also provides, in addition to the above-described features, the object recognition model is a model that learns faces with the WIDER FACE dataset for YOLO, the feature point coordinate detection model is an Ensemble of Regression Trees model of the Dlib library, the feature point coordinates (x, y) are coordinates for 68 feature points including facial contours, eyes, nose, and mouth among the face area, the fatigue measurement model is composed of three fully connected layers and two BN (Batch Normalization) layers, and receives 136 coordinates constituting the 68 feature points, and each neuron of the three fully connected layers is composed of 16, 8, and 2, and a softmax function is applied in the last layer to classify into one fatigue level, and when learning the fatigue measurement model, the correct answer (y) is i ) and predicted values It is also desirable to use the cross entropy loss function used in the above learning as an image analysis-based fatigue measurement system, which is characterized by the following, so that learning can be done in the direction of minimizing the loss function by multiplying.

[0021]

[0022] The present invention also implements, in addition to the above-described features, the fully connected layer is connected to a Global Average Pooling (GAP) layer, the CNN uses one frame size (W*H*C) as an input, has one convolution 2D layer and 16 filters, and has a kernel size of (3, 3) to extract the feature elements from the analysis target frame, and after making a multidimensional array into one dimension with the GAP layer, classifies fatigue with one fully connected layer, and when learning the CNN, the correct answer (y) i ) and predicted values It is also desirable to use an image analysis-based fatigue measurement system, which is characterized by the following loss function as cross entropy, so that learning can be done in the direction of minimizing the loss function by multiplying.

[0023]

[0024] The present invention also preferably provides an image analysis-based fatigue measurement system, characterized in that, in addition to the above-described features, the ensemble algorithm assigns a contribution (a) to each model according to the learning result of each model, such that the sum of the contributions (a) to each model is 1, assigns a weight (b) for each level of fatigue level to each model according to the learning result of each model, such that the weight (b) for each level is assigned in the range of 0 to 2, and the final fatigue level for each of the analysis target frames is an average of a value obtained by applying the contribution (a) for each model and the weight (b) for each score to the fatigue level classified by each model.

[0025] The present invention also provides that, in addition to the above-described features, the selection of the analysis target frame in the preprocessing module is performed by: - ​​M frame images (f) decomposed into frame units in the second step M ) generates a color histogram for each of them, and M color histograms (h) are generated in time order. M ), then repeat the process below until m = 1 and m = M -- and create the M color histograms (h M ) mth color histogram (h m ) is the m+nth (n starts from 1) color histogram (h m+n ) and the similarity compared to the m-th frame image (f m ) for similarity (S m ) as the first step; -- the above similarity (S m ) is greater than the threshold, the m+nth frame (f m+n) as a similar frame, and then repeat the above process from step 2-1 with n+1 as n; -- The above similarity (S m ) is less than the threshold value, the m-th frame image (f m ) as a key frame (kf m ) and set n as the number of similar frames, then set m + n as m and set n again as 1 and repeat the first process again; - The key frame (kf) selected in the first process to the third process m ) is provided as the analysis target frame, and it is also desirable to use an image analysis-based fatigue measurement system.

[0026] The present invention also provides that, in addition to the above-described features, the selection of the analysis target frame in the preprocessing module is performed by: - ​​decomposing the frame image (f) into frames in the second step m ) For each, the object is removed and the background is extracted using the Gaussian Mixture Model (GMM) - the frame image (f m ) After extracting the object-centered frame by subtracting the pixel value of the background from each original frame - the frame image (f m ) In each object-centered frame, all pixels are summed to obtain the frame image (f m ) behavioral information (MI) for each m ) is defined, - the frame image (f m ) behavioral information (MI) for each m ) are grouped into a certain number of units to create a window, and then - the above behavioral information (MI) included in the window m ) has the largest local maxima among the bundles of action information (MI) m ) corresponding to the mth frame image is a key frame (kf m) and select the key frame (kf) within the window. m ) exists, the back frame is the key frame (kf m ) as a similar frame, and - the key frame (kf) within the above window m ) exists, the preceding frame is the key frame (kf m ) previous key frame (kf m-1 ) as a similar frame, and - the above key frame (kf m ) is provided as the analysis target frame, and it is also desirable to have an image analysis-based fatigue measurement system.

[0027] The present invention also provides that, in addition to the above-described features, the ensemble module calculates the average of the fatigue levels or the most frequently occurring fatigue levels by calculating the key frame (kf) m ) It is also desirable to use a fatigue measurement system based on image analysis, characterized in that it calculates by reflecting the number (n) of similar frames above.

[0028] As described above, the present invention provides a system capable of determining a person's fatigue level by analyzing a video of a person's face taken for a relatively short period of time using a deep learning model that applies an ensemble algorithm, thereby having the effect of quickly and accurately measuring a subject's fatigue level through an information system.

[0029] The present invention also has the effect of providing objective and accurate fatigue measurement results because the information system measures the fatigue level of a subject using an ensemble algorithm including an object feature analysis model, an object-CNN model, and a CNN model.

[0030] In addition, it is a system that uses all of the models, including a model that detects and analyzes the facial area and then classifies the fatigue level, a model that classifies the fatigue level by extracting and comparing the coordinates of feature points in the facial area and analyzing them, and a model that classifies the fatigue level by analyzing parts other than the face, and applies an ensemble algorithm that reflects the result value according to the learning result, so it has the effect of providing fatigue measurement results objectively and accurately.

[0031] The present invention also extracts frame images at points in time when there is a change in facial movement or posture, etc., as key frames among numerous frames of an analysis target video, and analyzes only the key frames, and does not analyze images that are not key frames, that is, frame images that have the same content as the key frames because there is no change, so that computer resources can be efficiently utilized and fatigue can be analyzed quickly.

[0032] The present invention also selects only images that are meaningful for analysis from among the analysis target images and analyzes them as key frames, so that even if only some frames are analyzed, the same effect as analyzing all frame images can be achieved.

[0033] The present invention also has the effect of more accurately calculating the fatigue average because, even if only the key frame is extracted from the analysis target image and used as the analysis target, the number of identical or similar images following the key frame is also determined and provided, and this is reflected when calculating the fatigue average.

[0034] Figure 1 is a configuration diagram of a fatigue measurement system based on image analysis according to the present invention.

[0035] Figure 2 is a procedure diagram illustrating the execution process of the present invention.

[0036] Figure 3 is a conceptual diagram for explaining the basic concept of the object feature analysis model included in the present invention.

[0037] Figure 4 illustrates a framework in which an object feature analysis model included in the present invention is performed.

[0038] Figure 5 illustrates a concept of extracting coordinates of 68 facial key feature points from one facial image in an object recognition model included in the present invention.

[0039] Figure 6 illustrates a classifier based on a fully connected layer used in an object feature analysis model included in the present invention.

[0040] Figure 7 is a conceptual diagram for explaining the basic concept of the object-CNN model included in the present invention.

[0041] Figure 8 illustrates a framework in which the object-CNN model included in the present invention is performed.

[0042] Figure 9 illustrates a classifier based on a CNN and a fully connected layer used in the object-CNN model and CNN model included in the present invention.

[0043] Figure 10 is a conceptual diagram for explaining the basic concept of the CNN model included in the present invention.

[0044] Figure 11 illustrates a framework in which a CNN model included in the present invention is performed.

[0045] Figure 12 illustrates a framework in which a key frame is generated using a color histogram in the present invention.

[0046] Figure 13 is a flowchart of a procedure for generating a key frame using a color histogram in the present invention.

[0047] Figure 14 illustrates a framework in which key frames are generated using a Gaussian mixture model in the present invention.

[0048] Figure 15 is a flowchart of a procedure for generating key frames using a Gaussian mixture model in the present invention.

[0049] Figure 16 illustrates a classifier model used for learning in an experimental example of an object feature analysis model of the present invention.

[0050] Figure 17 illustrates a classifier model used for learning in an experimental example of the CNN model of the present invention.

[0051] Figure 18 illustrates changes in accuracy and loss in the learning and verification stages in an experimental example of the present invention.

[0052] The present invention relates to a fatigue measurement system based on image analysis, which is installed in an information system and can determine a fatigue level by analyzing a video of a person's face using a deep learning algorithm, the system comprising: a preprocessing module that receives a video including a subject's face as an analysis target video, decomposes the analysis target video into frame units, and then selects an analysis target frame from among the frames; a first classification module that classifies the fatigue level for each of the analysis target frames into a first fatigue level using an object feature analysis model; a second classification module that classifies the fatigue level for each of the analysis target frames into a second fatigue level using an object-CNN model; and a third classification module that classifies the fatigue level for each of the analysis target frames into a CNN model and classifies it into a third fatigue level; And an ensemble module that calculates a final fatigue level for each of the analysis target frames by applying an ensemble algorithm to the first fatigue level, the second fatigue level, and the third fatigue level; wherein the ensemble module determines the fatigue level that is most 'appearing' (high?) ​​among the average of the final fatigue levels for all of the analysis target frames or the final fatigue level for each of the analysis target frames as the fatigue level for the subject, it is preferable that the image analysis-based fatigue measurement system be configured.

[0053] The present invention will be described in detail below to clarify the above-described purposes and features, thereby enabling those skilled in the art to readily implement the technical concepts of the present invention. Furthermore, in describing the present invention, if a detailed description of a prior art related to the present invention is already well known in the art, and if it is determined that a detailed description of such prior art may unnecessarily obscure the gist of the present invention, such a detailed description will be omitted.

[0054] In addition, the terms used in the present invention are selected from widely used general terms as much as possible. However, in certain cases, there are terms arbitrarily selected by the applicant, and in such cases, the meanings thereof are described in detail in the description of the relevant invention. Therefore, it is to be noted that the present invention should be understood based on the meaning of the terms, not simply the names of the terms. The terms used in the description of the embodiments are used only to describe specific embodiments and are not intended to limit the embodiments. The singular expression includes the plural expression unless the context clearly indicates otherwise.

[0055] The embodiments may be modified in various ways and may have various additional embodiments. Specific embodiments are illustrated in the drawings and described in detail herein. However, this is not intended to limit the embodiments to a specific form, but rather to encompass all modifications, equivalents, or alternatives that fall within the spirit and technical scope of the embodiments.

[0056] In the descriptions of various embodiments, expressions such as "first," "second," "first," or "second" may describe various components of the embodiments, but do not limit those components. For example, the expressions do not limit the order and / or importance of the components. The expressions may be used to distinguish one component from another.

[0057] The present invention will be described below with reference to the attached drawings and other documents. Technologies for analyzing video data based on deep learning have been studied in various fields. Since video data is typically composed of multiple frame images combined, various video analysis techniques implemented in various ways have been developed and applied for analysis. The representative deep learning model used for video analysis is the Convolutional Neural Network (CNN), but since video data contains time-series information, the Recurrent Neural Network (RNN) model is also applied.

[0058] CNN models that can be utilized in image analysis include those presented in the prior art literature above, such as VGG, GoogleNet Inception, ResNet, Xception, MobileNet, DenseNet, EfficientNet, ViT, and CoAtNet, which are known and used as ideal models. Models such as ViT and CoAtNet, which are recently evaluated as the models with the best performance in image classification problems, are the application of Transformer, which is used in the field of natural language processing, to the problem of image classification.

[0059] Meanwhile, video data, unlike general data, has spatial elements as one of its characteristics. Therefore, CNN models are suitable for video analysis because they extract data features by including these. Furthermore, RNN models can be constructed using hidden layers such as LSTM and GRU. RNN models use sequences as input and output, and are primarily used for data containing time-series information, such as handwriting recognition and speech recognition. In the present invention, we combined known image analysis algorithms into a method applicable to fatigue measurement, and developed models that can be used in practice through experiments and verification.

[0060] Existing research on identifying frames in videos that match specific meanings primarily focuses on image-based action recognition. Initially, research focused on facial feature classification, beginning with the existing HOG method that utilizes frontal images of the face, and now pursuing analysis using natural images. The UCF101 dataset contains 101 action classes, 13,000 clips, and 27 hours of video data. Research is being conducted on techniques for recognizing human actions using this dataset. A representative example is a technique utilizing convolutional 3D.

[0061] However, the present invention does not aim to find a frame corresponding to a specific meaning in a video, but aims to find features in frames that make up the entire video and utilize them to measure fatigue. However, since there is no need to analyze all identical or similar frames, rather than targeting all frames included in a video, a method is provided to measure fatigue by selecting only key frame images that are meaningful for analysis, such as frames where there is a change of a certain size or more in facial movement or posture, thereby minimizing the consumption of resources and time due to repeated analysis of identical or similar frames.

[0062] Meanwhile, assuming that information related to fatigue is provided in the image data, the fatigue information suitability judgment model through image analysis becomes a type of classification model. Figure 1 illustrates the basic configuration of the image analysis-based fatigue measurement system according to the present invention. As illustrated in Figure 1, the present invention,

[0063] An information system for measuring fatigue based on image analysis that can determine the level of fatigue by analyzing a video of a person's face using a deep learning algorithm, the information system preferably includes a preprocessing module (100) that receives a video containing the subject's face as an analysis target video, decomposes the analysis target video into frame units, and then selects an analysis target frame from among the frames; a classification module (200) that classifies the fatigue level for each of the analysis target frames using various classification methods; and an ensemble module (300) that applies an ensemble algorithm to the results classified by the classification module (200) to calculate a final fatigue level for each of the analysis target frames.

[0064] In addition, it is preferable that the analysis module (200) includes a first classification module (210) that classifies the fatigue level for each of the analysis target frames into a first fatigue level by using an object feature analysis model, a second classification module (220) that classifies the fatigue level for each of the analysis target frames into a second fatigue level by using an object-CNN model, and a third classification module (230) that classifies the fatigue level for each of the analysis target frames into a CNN model by using a third fatigue level.

[0065] And the ensemble module (300) applies an ensemble algorithm to the first fatigue level, the second fatigue level, and the third fatigue level to calculate the final fatigue level for each of the analysis target frames, but it is preferable to determine the fatigue level that appears most frequently among the average of the final fatigue levels for all of the analysis target frames or the final fatigue levels for each of the analysis target frames as the fatigue level for the subject.

[0066] Methods available for decomposing a video into frame-by-frame images include random extraction, time-based random extraction, and key frame extraction. These separated frames are then analyzed using video analysis techniques.

[0067] FIG. 2 illustrates a flowchart of a procedure for classifying the fatigue level of a subject by a fatigue measurement system based on image analysis according to the present invention. As shown in FIG. 2, it is preferable that the preprocessing module (100) included in the present invention first performs a step (s10) of receiving a video including the subject's face as an analysis target image. Next, it is preferable that the preprocessing module (100) performs a step (s20) of decomposing the analysis target image into frames, and then performs a step (s30) of selecting an analysis target frame from among the decomposed frames.

[0068] And then, it is preferable that each of the classification modules (200) classify the fatigue level for each of the analysis target frames (s41 to s43). In this process, it is preferable that each of the first classification module (210), the second classification module (220), and the third classification module (230) classify the fatigue level for each of the analysis target frames using different models. The first classification module (210) performs the fourth step (s41) of classifying the fatigue level for each of the analysis target frames using an object feature analysis model and classifying it as a first fatigue level, the second classification module (220) classifies the fatigue level for each of the analysis target frames using an object-CNN model and classifies it as a second fatigue level, and the third classification module (230) classifies the fatigue level for each of the analysis target frames using a CNN model and classifies it as a third fatigue level.

[0069] Next, the ensemble module (300) performs a seventh step (s50) of calculating the final fatigue level for each of the analysis target frames by applying an ensemble algorithm to the first fatigue level, the second fatigue level, and the third fatigue level. Here, the ensemble algorithm assigns a contribution (a) to each model according to the learning result of each model, such that the sum of the contributions (a) to each model is 1, assigns a level-specific weight (b) of the fatigue level to each model according to the learning result of each model, such that the level-specific weight (b) is assigned in the range of 0 to 2, and the final fatigue level for each of the analysis target frames is preferably an average value obtained by applying the contribution (a) to each model and the weight (b) to each score to the fatigue level classified by each model.

[0070] For example, if the learning results for each model are 52% for the object feature analysis model, 59% for the object-CNN model, and 67% for the CNN model, the contribution (a1) of the object feature analysis model will be 52 / (52+59+67) = 29%, the contribution (a2) of the object-CNN model will be 59 / (52+59+67) = 33%, and the contribution (a3) ​​of the CNN model will be 67 / (52+59+67) = 37%.

[0071] And the level-by-level weight (b) of the fatigue level for each model is to apply differently according to the classification result of the fatigue level within each model, since the fatigue level evaluated by each model may be accurately judged at some score range but not at others. For example, when the fatigue level is classified into 5 levels from 1 to 5, the object feature analysis model accurately classifies fatigue levels 4 to 5, but if the accuracy is not high for fatigue levels 1 to 2, the level-by-level weight (b1) can be assigned 1.5 for 4 to 5, 1 for 3, and 0.5 for 1 to 2 according to the accuracy. In this way, the object-CNN model and the CNN model also have level-by-level weights (b) of fatigue levels.

[0072] Therefore, the final fatigue level for each of the above analysis target frames is the value obtained by dividing (fatigue level classified by object feature analysis model x contribution of object feature analysis model (a1) x level-by-level weight of fatigue level of object feature analysis model (b1)) + (fatigue level classified by object-CNN model x contribution of object-CNN model (a2) x level-by-level weight of fatigue level of object-CNN model (b2)) + (fatigue level classified by CNN model x contribution of CNN model (a2) x level-by-level weight of fatigue level of CNN model (b2)) by 3, which is the final fatigue level.

[0073] Meanwhile, after the final fatigue level is calculated using the ensemble algorithm, it is preferable to perform an 8th step (s60) of determining the fatigue level for the subject as the average of the final fatigue levels for all of the analysis target frames or the fatigue level that appears most (frequently appears) among the final fatigue levels for each of the analysis target frames.

[0074] Hereinafter, the process of classifying the fatigue level for each model will be described. First, FIGS. 3 to 6 illustrate diagrams for explaining the object feature analysis model used in the first classification module (210) included in the present invention. FIG. 3 illustrates a conceptual diagram for explaining the basic concept of the object feature analysis model, and FIG. 4 illustrates a framework in which the object feature analysis model is performed. FIG. 5 illustrates a concept for extracting coordinates of 68 facial key feature points from one facial image in an object recognition model, and FIG. 6 illustrates a classifier based on a fully connected layer used in the object feature analysis model.

[0075] To summarize the outline of the object feature analysis model among the fatigue level classification models included in the present invention, this model first decomposes the analysis target image into frame units, then selects the analysis target frame from the decomposed frames, detects the face area of ​​the subject from the analysis target frame selected in this way, recognizes the face, extracts the elements that make up the face, and then calculates the fatigue level using this.

[0076] To this end, the system according to the present invention preferably, in step (s41) performed after the step (s10) of inputting a video including the face of a subject as an analysis target video, the step (s20) of decomposing the analysis target video into frames, and the step (s30) of selecting an analysis target frame from among the frames, detects the face area of ​​the subject in the analysis target frame using an object recognition model, extracts a predetermined number of feature point coordinates (x, y) from the face area using a feature point coordinate detection model, and classifies the fatigue level for each of the analysis target frames by comparing the feature point coordinates with a true fatigue value using a classification algorithm-based fatigue measurement model, thereby determining the fatigue level as the first fatigue level.

[0077] That is, for video analysis, the image is decomposed into frames, and then face detection is performed to remove noise such as background and bowed heads from the decomposed frames. In order to extract the locations of the eyes, nose, mouth, etc. from the detected face, landmarks named as feature point coordinates (x,y) are extracted in a fixed number, and then a machine learning / deep learning model is used to determine whether the extracted coordinates are suitable for the fatigue true value.

[0078] Hereinafter, the process of each step of the model will be described in detail with reference to FIGS. 3 and 4. First, when the first step, in which a video containing the subject's face is input as the analysis target video, is performed, and then the second step, in which the analysis target video is decomposed into frame units, is performed, a single video of the subject is decomposed into the following set of images.

[0079]

[0080] Here, there is no significant difference in image analysis efficiency whether black and white or color is selected. However, assuming that color can generally improve performance, it is recommended to set the channel to 3 (color). In addition, it is desirable to use the frame size directly from the original video size as input to the face detection model. However, since the size of 1280*720 is too large to store in memory, it is also desirable to perform image decomposition and face region detection simultaneously to reduce the frame size and improve resource efficiency.

[0081] And it is preferable to perform a third step of selecting a "frame to be analyzed" from among the above frames. It is also preferable that the method of selecting the frame to be analyzed be a time-based random extraction method in which a random frame is extracted at regular intervals. That is, the maximum number of frames and a regular interval are set, and the frame extraction is performed after setting a random number of frames to be extracted between the regular intervals.

[0082] Meanwhile, after the third step of selecting a target frame for analysis from among the above frames is performed, it is preferable to perform the fourth step of detecting a face area of ​​the subject in the target frame for analysis using an object recognition model. To this end, a face area is detected in each of the target frames for analysis. The face area is extracted within the resolution of the target frame for analysis, and representative object recognition models used for this include R-CNN, YOLO, and SSD. The face area detected using the object recognition model can be expressed as follows.

[0083]

[0084] Next, it is desirable to extract a fixed number (68) of feature point coordinates (x, y) from the face region using a feature point coordinate detection model. That is, a list of key coordinates to be used for analysis is extracted from the feature point coordinates (x, y) in the detected face region. It is also desirable to use the Ensemble of Regression Trees model of the Dlib library as a method for detecting the feature point coordinates included in the face region (Face landmark detection). The key coordinates of the face can be expressed as follows. At this time, as shown in Fig. 5, since |L| = 68, it is desirable to extract 68 key facial feature point coordinates from one face image. /

[0085]

[0086] Next, it is desirable to classify the first fatigue level for each of the analysis target frames by comparing the feature point coordinates (x, y) with the true fatigue value using a classification algorithm-based fatigue measurement model. At this time, the extracted elements are fitted to a given fatigue level through a deep learning algorithm, and these are accumulated and connected to a classifier that ultimately determines the fatigue level. There are various ways to build this analysis system, but due to the nature of the data that uses a list of major facial coordinates as input, it is desirable to use a classifier using a fully connected layer (FCL), and its basic structure is as shown in Fig. 6. In the post-training test process, the image is decomposed into frames, and then the probability for each fatigue level is derived for each frame through the learned model, and the value with the largest number of derived fatigue levels is returned as the final fatigue level.

[0087] Meanwhile, FIGS. 7 to 9 illustrate diagrams for explaining the object-CNN model used in the second classification module (220) among the fatigue level classification models included in the present invention. FIG. 7 illustrates a conceptual diagram for explaining the basic concept of the object-CNN model, and FIG. 8 illustrates its framework. In addition, FIG. 9 illustrates a classifier based on CNN and a fully connected layer used in the object-CNN model.

[0088] To summarize the outline of the object-CNN model included in the present invention, the object-CNN model included in the present invention first decomposes the analysis target image into frames, then selects the analysis target frame from the decomposed frames, detects the face area of ​​the subject from the selected analysis target frame, recognizes the face, extracts the elements that constitute the face, and then calculates the fatigue level using this.

[0089] To this end, the analysis system according to the present invention preferably comprises the steps of: receiving a video including the face of a subject as an analysis target video (s10); decomposing the analysis target video into frames (s20); and selecting an analysis target frame from among the frames (s30); then, using an object recognition model, detecting the face region of the subject in the analysis target frame; extracting feature elements in vector format from the face region using a CNN; and sending the feature elements to a fully connected layer within the CNN to classify the fatigue level for each of the analysis target frames and determining it as the second fatigue level. Here, one image of the subject is decomposed into a set of images as described in the object feature analysis model, and the face region is also detected using the same method as the object feature analysis model.

[0090] That is, in the object-CNN model included in the present invention, the face region is detected in the same manner as the object feature analysis model, and then the CNN is used to extract the feature elements for each face region for each analysis target frame, and these are sent to a fully connected layer to match the given fatigue level and classify and determine the second fatigue level. In addition, a CNN model is used to extract the feature elements of the face region, and according to the structure of the model, the feature elements are extracted in vector format. This vector is directly input into a classifier composed of a fully connected layer to match the given fatigue level.

[0091] It's desirable to implement fully-connected layers in CNNs using Global Average Pooling (GAP) layers. This prevents overfitting due to the explosive increase in parameter count that occurs when using traditional fully-connected layers, such as in object-feature models. Regardless of the size of the extracted feature map, the GAP layer replaces each color channel with the average of the values ​​contained within it. This prevents overfitting and improves classification performance by reflecting spatial information.

[0092] Meanwhile, FIGS. 10 and 11 illustrate diagrams for explaining the CNN model used in the third classification module (230) included in the present invention. FIG. 10 illustrates a conceptual diagram for explaining the basic concept of the CNN model included in the present invention, and FIG. 11 illustrates its framework. In addition, the CNN and fully connected layer-based classifier used in the CNN model included in the present invention are identical to FIG. 9.

[0093] To summarize the outline of the CNN model included in the present invention, although the CNN model included in the present invention decomposes a video into frame-by-frame images, there is no preprocessing step to recognize specific parts of the decomposed frames as objects, and no separate step is performed to extract feature elements of major parts existing in the corresponding images (feature extraction). That is, as shown in FIGS. 10 and 11, object recognition is not used, and feature element extraction is processed within the CNN model rather than through a separate method.

[0094] To this end, the analysis system according to the present invention preferably comprises the steps of: receiving a video including the face of a subject as an analysis target video (s10); decomposing the analysis target video into frame units (s20); and selecting an analysis target frame from among the frames (s30); then extracting a feature element in vector format from each of the analysis target frames using CNN; and sending the feature element to a fully connected layer within the CNN to classify the fatigue level for each of the analysis target frames and determining it as the third fatigue level.

[0095] That is, using a CNN, feature elements are extracted for each frame and sent to a fully connected layer to match the given fatigue level to determine the third fatigue level. A CNN-based model is used to extract frame features, and according to the model's structure, the features are extracted in vector format. This vector is then directly input into a classifier (Figure 9) consisting of a fully connected layer to match the given fatigue level.

[0096] It is desirable to use a Global Average Pooling (GAP) layer to implement a fully connected layer in a CNN model. This prevents overfitting due to the explosive increase in the number of parameters that occurs when using traditional fully connected layers such as object-feature models. Since the GAP layer replaces each color channel with the average of the values ​​included, regardless of the size of the extracted feature map, it is expected to improve classification performance by preventing overfitting and reflecting spatial information. In the post-training testing process, the image is decomposed into frames, and the probability of each fatigue level is derived using the trained model. The value with the highest number of derived fatigue levels is returned as the final fatigue level.

[0097] Meanwhile, in the third step of determining the analysis target frame of the present invention, it is also preferable to use a 'method using key frames' rather than a 'method of extracting random frames at regular intervals'. Applying a method of using key frames when determining the analysis target frame allows for more efficient utilization of computer resources for video analysis, shortens analysis time, and enables more accurate measurement of fatigue levels across the entire video.

[0098] In the present invention, the method of using key frames excludes from analysis any frames that are the same or similar to the previous frame, and selects and analyzes only frames that are meaningful for analysis because they are different from the previous frame, while also determining and providing the number of identical or similar frames located after the selected key frame, thereby enabling a more accurate value to be calculated when calculating the fatigue average from the entire video.

[0099] That is, since each frame captured while the subject has no change in expression or movement is the same or similar frame, the facial area will be detected the same, and the coordinates of the feature points will all be detected the same, so when measuring fatigue, they will all come out as the same value. Therefore, in the present invention, among the frame images, a frame that is meaningful for measuring fatigue, which appears different from the previous frame due to movement or change in expression, is made into a key frame, and only the image for the key frame is extracted to measure fatigue, and the number of identical or similar frames repeated after the key frame is also determined and reflected in the calculation of the fatigue average, thereby increasing the accuracy of the calculation of the fatigue average.

[0100] To this end, the present invention provides a method using a Gaussian Mixture Model (GMM) along with a method using a color histogram.

[0101] First, the key frame generation method using a color histogram is to examine frames using a color histogram color distribution diagram as shown in Fig. 12, select a frame with a large difference in similarity with the previous frame as a key frame, and provide the number of identical or similar frames for the selected key frames. The color distribution diagram is used to indicate the frequency of occurrence of each value for all pixel values ​​that pixels within an image can have. This color distribution diagram can be used to check the rough distribution of pixels within an image and can be used when comparing with other images.

[0102] In order to find a key frame using a color histogram, a color distribution corresponding to a frame of the video is first created. The video data, which is the flow of frames, can be identified by the flow of the color distribution. The point in time when a significant change occurs in this flow is extracted as a key frame. It is desirable to calculate the similarity of the color distributions of the preceding and following frames using cosine similarity, and determine whether the change in the frame is meaningful based on a set threshold. In this way, the present invention utilizes a color distribution on a frame-by-frame basis, so the processing speed is fast, and since it does not use a function or preprocessing method that depends on specific data in advance, it is easy to implement for utilizing multiple datasets.

[0103] To achieve this, the input video is divided into individual frames, and a color distribution map is generated for each frame image. The generated color distribution maps are then arranged in chronological order within the video, and each color distribution map is compared with the previous one to measure similarity. If the measured similarity is lower than a set threshold or if the frame has low similarity compared to the previous image, the frame is extracted as a key frame.

[0104] Fig. 13 illustrates a specific procedure flow for performing a key frame generation method using a color histogram included in the present invention. The following will be described with reference to Fig. 13. First, a step (s110) is performed to input a video including the subject's face as an analysis target image, and then a step (s120) is performed to decompose the analysis target image into frames, and then M frame images (f) decomposed into frame units in the second step are generated. M ) generates a color histogram for each of them, and M color histograms (h) are generated in time order. M ) is to be performed in the process of creating (s130).

[0105] And it is desirable for the above information system to perform the process (s140 to s148) of repeating the process below from m = 1 until m = M.

[0106] That is, the above M color histograms (h M ) mth color histogram (h m ) is the m+nth (n starts from 1) color histogram (h m+n ) and the similarity compared to the m-th frame image (f m ) for similarity (S m ) is performed, and then the similarity (S m ) is greater than the threshold value (s142), the m+nth frame (f m+n) is set as a similar frame, it is desirable to perform the second process by repeating the process from the above 2-1 process by setting n+1 as n (s143).

[0107] However, the above similarity (S m ) is less than the threshold value (s142), the m-th frame image (f m ) as a key frame (kf m ) and set n as the number of similar frames (s144), then set m + n as m (s146), set n again as 1 (s147), and perform the third process of repeating from the first process, but it is preferable to perform the process until m becomes M (s148).

[0108] And, the key frames (kf) selected in the above processes m ) is preferably performed to provide the analysis target frame (s150). Here, the similarity (S m ) is preferably measured using cosine similarity as shown in the equation below.

[0109]

[0110] Meanwhile, when calculating the 'average of fatigue levels (p)' or 'most frequently occurring fatigue levels', the key frame (kf) m ) is preferably calculated by reflecting the number (n) of the similar frames for each key frame. For example, when calculating the 'average of fatigue level (p)', each key frame (kf m ) considering the fatigue level (p) of each key frame (kf) and the number of similar frames (n). m ) and the fatigue level (p) for the corresponding similar frames ((n+1) xp), and then all key frames (kf m) is the sum of the above fatigue levels (p) considering the number (n) of the similar frames for each frame, and the value divided by the number of all frames (M) in the video is taken as the 'average of fatigue levels (p)'. In addition, when obtaining the 'most frequently appearing fatigue level', each key frame (kf) is calculated in the same way. m ) is calculated by taking into account the number of similar frames (n) for each fatigue level (p), and then calculating the number of fatigue levels (p) that appear most frequently.

[0111] As described above, another method for selecting the target frame for analysis is to utilize a Gaussian Mixture Model (GMM). This will be described with reference to FIGS. 14 and 15. This method also extracts only frame images at points in time when there is a change of a certain magnitude or greater in facial movement from video information including a person's face and provides them as information for analysis. After analyzing the flow of frames using GMM, a part where the difference between the front and back of the frames appears sharply is discovered and the corresponding frame is extracted as a key frame.

[0112] Figure 14 illustrates the concept. First, the input image is divided into frames, and then GMM is used to remove objects from each frame and extract the background. By subtracting the pixel values ​​of the extracted background from the original frame, an object-centered frame with the background removed can be extracted. Then, by summing all pixels in each frame containing only the object, the motion information of the frame is defined. Here, a window of frames is used to sequentially check the motion information of frames of a predetermined size, and the largest local maxima in the window is selected. In this way, the frame with the maximum value in each window is extracted as a keyframe. In keyframe extraction using conventional techniques, it is sometimes difficult to detect significant changes in the flow of frames. However, the present invention utilizes motion information to compensate for this.

[0113] Figure 15 illustrates the procedure flow for performing a method for generating key frames for image information analysis using a Gaussian mixture model. First, a step (s210) is performed to input a video containing the subject's face as an analysis target image, and a step (s220) is performed to decompose the analysis target image into frames, and then the frame image (f) decomposed into frames is generated. m ) It is desirable to perform a process of removing objects and extracting backgrounds using Gaussian Mixture Model (GMM) for each (s230).

[0114] And, the above frame image (f m ) A process (s240) is performed to extract an object-centered frame by subtracting the pixel value of the background from each original frame, and then the frame image (f m) In each object-centered frame, all pixels are summed to obtain the frame image (f m ) behavioral information (MI) for each m ) is to be performed in the process of defining (s250).

[0115] And, the above frame image (f m ) behavioral information (MI) for each m ) is grouped into a certain number (N) units to create windows (s261), and then the behavioral information (MI) included in the window is created. m ) has the largest local maxima among the bundles of action information (MI) m ) corresponding to the mth frame image is a key frame (kf m ) It is desirable to perform the selection process (s262).

[0116] After performing the above s262 process, the key frame (kf) within the above window m ) exists, the back frame is the key frame (kf m ) as a similar frame (s263) and the key frame (kf) within the window m ) exists, the preceding frame is the key frame (kf m ) previous key frame (kf m-1 ) It is desirable to perform the process (s264) of making it into a similar frame.

[0117] And finally, the selected key frame (kf m) is provided as the analysis target frame, and then it is preferable to repeat the process s261 to s265 for the next N frames. In addition, when calculating the average fatigue level or the most frequently occurring fatigue level, as described in the method of using the histogram, the key frame (kf m ) It is desirable to calculate it by reflecting the number (n) of similar frames above.

[0118] By providing key frames selected using a color histogram or Gaussian mixture model as the frames to be analyzed, analysis of identical frames without movement can be omitted and fatigue can be measured only for meaningful frame images, allowing for efficient use of computer resources and quick analysis. On the other hand, since the number of identical or similar frames repeated after the key frame can be identified and reflected in the calculation of the fatigue average, the accuracy of the calculation of the fatigue average can be improved.

[0119] The present invention will be described in more detail below through experimental examples. The experimental examples described below are provided to aid understanding of the present invention, and it should be understood that the present invention can be implemented with various modifications that differ from the experimental examples described herein. It will be apparent to those skilled in the art that various modifications and variations are possible within the scope and technical spirit of the present invention, and it is also natural that such modifications and variations fall within the scope of the appended claims.

[0120] <Experimental Example>

[0121] 1. Dataset

[0122] The data used for the experiment was collected through a fatigue-related video and biometric data collection system. The fatigue level was estimated by analyzing the subject's video data and fatigue level. Each data set was a one-minute video of the subject according to a specific scenario and script, generally filmed from the subject's front. Fatigue levels were assessed based on the subject's subjective judgment and were rated as best (1), good (2), average (3), bad (4), and worst (5).

[0123] For the present invention, experiments were conducted using images corresponding to fatigue levels 1 and 5 from the entire dataset. A total of 533 images were divided into 80%, 10%, and 10%, respectively, and used as learning, validation, and test datasets. The validation dataset was used to adjust the hyperparameters for each model to find the most appropriate model, and finally, the test dataset was used to verify the performance of the proposed fatigue measurement models.

[0124] Structured random sampling (SRS) was used to extract frames from a video. The entire video duration was divided into n-second units, and a random frame was extracted from each unit. In other words, all videos in the experiment were preprocessed into sets of n frames (S = 60).

[0125] 2. Algorithms for Experiments

[0126] The object feature analysis model used the YOLO v3 object recognition model for the face detection model, trained faces with the WIDER FACE dataset (Non-patent document 28), and utilized it for face detection. For facial feature point extraction, the Ensemble of Regression Trees (Non-patent document 29) model of the Dlib library was utilized. A total of 68 face, eye, nose, and mouth coordinates (x, y) were extracted from the frame and used as input for the classifier. The classifier used for learning is as shown in Fig. 16.

[0127] 68 coordinates were used as input for 136 vectors, and with 3 fully connected layers and 2 BN (Batch Normalization) layers, each neuron was composed of 16, 8, and 2, and the softmax function was applied in the last layer to classify it into a given fatigue level.

[0128] The CNN model is a form in which a fully connected layer is combined in CNN, as shown in Figure 17. One frame (H, W, 3) size is used as input, and one convolution 2D layer, 16 filters, and (3, 3) kernel size are used to extract the frame's features, and after converting the multidimensional array into one dimension with a GAP (GlobalAveragePooling) layer, fatigue is classified with one fully connected layer.

[0129] The loss function used during learning was cross entropy, and the method is as follows. Correct answer (y i ) and predicted values We trained in the direction of minimizing the loss function by multiplying .

[0130]

[0131] The model's training and testing were performed in a Python 3.7-based TensorFlow-Keras environment on a single workstation equipped with a Linux-based GPU. Precision and recall for each fatigue level of the experimental subjects and accuracy for the entire fatigue level were used as measures of performance evaluation.

[0132] 3. Experimental Results and Discussions

[0133] The object feature analysis model has a verification accuracy exceeding 60% as learning progresses, but the verification loss has a large variability and the loss value is 1 to 3.5, so it was observed that the learned model has limitations in verifying the effectiveness using the verification dataset.

[0134] The CNN model exhibited overfitting, with validation loss decreasing and then rapidly increasing as training progressed. Validation loss was calculated to be 0.65–0.85, which was relatively lower than that of the object feature analysis model. While the CNN model exhibited significant variability in validation accuracy and loss, suggesting model instability, its performance on the experimental dataset predicted that a CNN model with a high maximum validation accuracy and a low minimum validation loss would exhibit high classification performance.

[0135] The changes in accuracy and loss during the training and validation stages are shown in Figure 18. Experimental results show that validation accuracy converges rapidly before 30 epochs, while training loss continuously decreases and validation loss increases, indicating overfitting. As training progresses, validation accuracy exceeds 60%, but validation loss fluctuates significantly, with loss values ​​ranging from 1 to 3.5, limiting the effectiveness of the trained model on the validation dataset. The performance results obtained from experiments using test data are shown in Table 1 below.

[0136] ModelClassPrecisionRecallAccuracyObjectFeatureAnalysisModel10.61540.50000.5185ObjectFeatureAnalysisModel50.42860.54550.5185CNN10.65910.90620.6667CNN50.70000.31820.6667

[0137] While the present invention has been described with numerous examples, it is not necessarily limited to these examples, and various modifications may be made without departing from the scope of the present invention. Therefore, the examples disclosed herein are intended to illustrate, rather than limit, the technical concept of the present invention, and the scope of the present invention is not limited by these examples. The scope of protection of the present invention should be interpreted according to the following claims, and all technical ideas within the scope equivalent thereto should be interpreted as being included within the scope of the present invention.

[0138] The present invention can provide a system that quickly and accurately determines a person's fatigue level by analyzing a video of a person's face using a deep learning algorithm.

Claims

1. An information system that measures fatigue based on image analysis that can determine the level of fatigue by analyzing a video of a person's face using a deep learning algorithm. A preprocessing module that receives a video containing the subject's face as an analysis target video, decomposes the analysis target video into frames, and then selects analysis target frames from among the frames; A first classification module that classifies the fatigue level for each of the above analysis target frames into a first fatigue level by using an object feature analysis model; A second classification module that classifies the fatigue level for each of the above analysis target frames using an object-CNN model and classifies it into a second fatigue level; A third classification module that classifies the fatigue level for each of the above analysis target frames using a CNN model and classifies it into a third fatigue level; and An ensemble module that applies an ensemble algorithm to the first fatigue level, the second fatigue level, and the third fatigue level to calculate a final fatigue level for each of the analysis target frames; The above ensemble module is characterized in that it determines the fatigue level that appears the most among the average of the final fatigue levels for all of the analysis target frames or the final fatigue levels for each of the analysis target frames as the fatigue level for the subject, based on image analysis.

2. In paragraph 1, The above first analysis module, - Using the object recognition model, the face area of ​​the subject is detected in the analysis target frame, - Using the feature point coordinate detection model, a set number of feature point coordinates (x, y) are extracted from the face area, - A fatigue measurement system based on image analysis, characterized in that it classifies the fatigue level for each of the analysis target frames by comparing the coordinates of the feature points with the true fatigue value using a fatigue measurement model based on a classification algorithm and sets it as the first fatigue level.

3. In paragraph 1, The above second analysis module, - Using the object recognition model, the face area of ​​the subject is detected in the analysis target frame, - Extract feature elements in vector format from the above facial area using CNN, - A fatigue measurement system based on image analysis, characterized in that the feature elements are sent to a fully connected layer within the CNN to classify the fatigue level for each of the analysis target frames and determine it as the second fatigue level.

4. In paragraph 1, The above third analysis module, - Using CNN, feature elements are extracted in vector format from each of the above analysis target frames, - A fatigue measurement system based on image analysis, characterized in that the feature elements are sent to a fully connected layer within the CNN to classify the fatigue level for each of the analysis target frames and determine it as the third fatigue level.

5. In paragraph 2 The above object recognition model is trained on faces using the WIDER FACE dataset for YOLO. The above feature point coordinate detection model is the Ensemble of Regression Trees model of the Dlib library. The above feature point coordinates (x, y) are coordinates for 68 feature points including the facial contour, eyes, nose, and mouth among the above facial areas. The above fatigue measurement model consists of three fully connected layers and two BN (Batch Normalization) layers, and receives 136 coordinates that constitute the 68 feature points. Each neuron of the three fully connected layers above consists of 16, 8, and 2, and the softmax function is applied in the last layer to classify them into one fatigue level. When learning the above fatigue measurement model, the correct answer (y i ) and predicted values The cross entropy loss function used in the above learning is characterized by the following, so that learning can be done in the direction of minimizing the loss function by multiplying: a fatigue measurement system based on image analysis 6. In paragraph 3 or 4 The above fully connected layer is implemented by connecting to the Global Average Pooling (GAP) layer. The above CNN uses one frame size (W*H*C) as input, has one convolution 2D layer and 16 filters, and has a kernel size of (3, 3) to extract the feature elements from the frame to be analyzed, and classifies fatigue with one fully connected layer after making the multidimensional array into one dimension with the GAP layer. When training the above CNN, the correct answer (y i ) and predicted values The loss function used in the above learning is cross entropy, which is characterized by the following image analysis-based fatigue measurement system so that learning can be done in the direction of minimizing the loss function by multiplying 7. In paragraph 1 The above ensemble algorithm is, A contribution (a) is assigned to each model based on the learning results of each model, but the sum of the contributions (a) of each model is assigned to be 1. A weight (b) for each level of fatigue is assigned to each model according to the learning results for each model, and the weight (b) for each level is assigned in the range of 0 to 2. The final fatigue level for each of the above analysis target frames is an image analysis-based fatigue measurement system characterized in that the average of the values ​​obtained by applying the contribution (a) for each model and the weight (b) for each score to the fatigue level classified by each model.

8. In paragraph 1 In the above preprocessing module, the selection of the analysis target frame is: - M frame images (f) decomposed into frame units in the above second step M ) generates a color histogram for each of them, and M color histograms (h) are generated in time order. M ) after creating it, - Repeat the process below from m = 1 until m = M. -- The above M color histograms (h M ) mth color histogram (h m ) is the m+nth (n starts from 1) color histogram (h m+n ) and the similarity compared to the m-th frame image (f m ) for similarity (S m ) as the first step; -- The above similarity (S) m ) is greater than the threshold, the m+nth frame (f m+n ) is set as a similar frame, and then the second process is repeated from the above 2-1 process with n+1 as n; -- The above similarity (S) m ) is less than the threshold value, the m-th frame image (f m ) as a key frame (kf m ) and selects the similar frames, and then sets n as the number of similar frames, and then sets m + n as m and sets n again as 1, and repeats the first process again; - The key frame (kf) selected in the first to third processes m ) is provided as the analysis target frame, and is characterized by providing the image analysis-based fatigue measurement system.

9. In paragraph 1 In the above preprocessing module, the selection of the analysis target frame is: - The frame image (f) decomposed into frame units in the second step m ) For each, the object is removed and the background is extracted using the Gaussian Mixture Model (GMM). - The above frame image (f m ) After extracting the object-centered frame by subtracting the pixel value of the background from each original frame, - The above frame image (f m ) In each object-centered frame, all pixels are summed to obtain the frame image (f m ) behavioral information (MI) for each m ) is defined, - The above frame image (f m ) behavioral information (MI) for each m ) are grouped into a certain number of units to create a window, - The above behavioral information (MI) included in the above window m ) has the largest local maxima among the bundles of action information (MI) m ) corresponding to the mth frame image is a key frame (kf m ) and select them, - The key frame (kf) within the above window m ) exists, the back frame is the key frame (kf m ) as a similar frame, - The key frame (kf) within the above window m ) exists, the preceding frame is the key frame (kf m ) previous key frame (kf m-1 ) as a similar frame, - The above key frame (kf m ) is provided as the analysis target frame, and is characterized by providing the image analysis-based fatigue measurement system.

10. In paragraph 8 or 9, The above ensemble module calculates the average of the fatigue levels or the most frequently occurring fatigue levels based on the key frame (kf m ) A fatigue measurement system based on image analysis, characterized in that it calculates by reflecting the number (n) of similar frames above.

Citation Information

Patent Citations

  • Quick violent-terrorism video recognition method based on comparison

    CN108734106A

  • Apparatus and method for treating substrate

    KR1020250001035A

  • Seismic apparatus for switch board using air cylinder structure

    KR1020250166592A