Miner working pressure identification method based on facial expression

Through the improved MiniXception network model and feature processing technology, the problem of low accuracy in facial expression recognition of miners is solved, and efficient real-time pressure recognition is achieved in complex environments.

CN120452039APending Publication Date: 2025-08-08XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510511294.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The prior art miner facial expression recognition accuracy in complex environments is low and cannot meet the real-time pressure recognition needs.

Method used

The improved MiniXception network model is adopted, combining the depth-separable convolution and dynamic adaptive feature fusion module to perform facial feature extraction, and the image is processed through weighted average, bilinear interpolation and Gaussian filtering to improve recognition accuracy.

Benefits of technology

In complex environments, the accuracy of expression recognition is improved, the data processing delay is reduced, and the need for real-time pressure recognition is met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452039A_ABST
    Figure CN120452039A_ABST
Patent Text Reader

Abstract

The invention discloses a miner working pressure identification method based on facial expressions, which comprises the following steps: continuously acquiring a first image of the face of an operator, performing data caching, and outputting the first image based on a caching threshold value; performing graying processing on the first image to obtain a grayscale image; performing normalization processing on the grayscale image by using a bilinear interpolation algorithm and a histogram equalization algorithm to obtain a normalized image; performing noise reduction processing on the normalized image by using a Gaussian filtering algorithm to obtain a noise-reduced image; and performing feature extraction on the noise-reduced image by using an expression recognition model, determining an expression type, and transmitting the expression type to a pressure early warning module to determine whether to trigger early warning. According to the expression recognition model, the deep separable convolution and dynamic adaptive feature fusion module is adopted to efficiently extract facial features, the accuracy of image recognition is improved, and the problem of low expression recognition accuracy in a complex environment in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pressure recognition, and in particular to a method for recognizing the working pressure of miners based on facial expressions. Background Art

[0002] With the development of the mining industry, the psychological stress and risks faced by miners at work are gradually gaining attention. The high-intensity working environment, heavy workload, and potential safety hazards can place a significant psychological burden on miners and even lead to safety accidents. Therefore, timely identification of miners' psychological states, especially their work-related stress, is crucial. Traditional stress measurement relies primarily on physiological characteristics (such as heart rate and galvanic skin response). While effective, these methods typically require specialized monitoring equipment and cannot provide real-time data. Furthermore, the assessment of psychological stress often requires subjective feedback, which can be subject to significant individual variation and subjectivity.

[0003] Existing technologies for recognizing miners' facial expressions have low recognition accuracy due to the complex working environment (such as lighting changes, the presence of obstructions, subtle differences in expressions, etc.), which affects the accurate judgment of the miners' work stress status. In addition, existing technologies have delays in processing and analyzing facial expression data, which cannot meet the needs of real-time recognition of miners' work stress. Therefore, it is necessary to design a stress recognition method with high recognition accuracy and low data processing delay. Summary of the Invention

[0004] The purpose of the present invention is to overcome the deficiencies in the above-mentioned prior art and provide a method for identifying work stress in miners based on facial expressions. The expression recognition model of the present invention adopts an improved MiniXception network model structure. The model uses a deep separable convolution and a dynamic adaptive feature fusion module to efficiently extract facial features, improve the accuracy of image recognition, and solve the problem of low accuracy of expression recognition in complex environments in the prior art.

[0005] To achieve the above-mentioned purpose, the technical solution adopted by the present invention is: a method for identifying work stress of miners based on facial expressions, comprising: continuously collecting a first image of the operator's face and caching the data, and outputting the first image based on a cache threshold; graying the first image using a weighted average method to obtain a grayscale image; normalizing the grayscale image using a bilinear interpolation algorithm and a histogram equalization algorithm to obtain a normalized image; denoising the normalized image using a Gaussian filtering algorithm to obtain a denoised image; extracting features from the denoised image using an expression recognition model to determine the expression type, and transmitting the expression type to a pressure warning module to determine whether to trigger a warning.

[0006] Preferably, the weighted average method converts the pixel values of the color image into grayscale values according to preset weights based on the sensitivity of the human eye to different colors, as shown in the following formula:

[0007] Gray=d1*R+d2*G+d3*B (1)

[0008] Where Gray represents the grayscale value of the pixel, d1 represents the red weight, R represents the red pixel value of the pixel, d2 represents the green weight, G represents the green pixel value of the pixel, d3 represents the blue weight, and B represents the blue pixel value of the pixel.

[0009] Preferably, the normalization includes size normalization and brightness normalization. The size normalization uses a bilinear interpolation algorithm to adjust the size of the grayscale image to 100*100 pixels; the brightness normalization uses a histogram equalization algorithm to map the brightness value of the grayscale image to the range of [0,1], as shown in the following formula:

[0010]

[0011] Where: (x, y) is the pixel coordinate, CDF max is the maximum value of the cumulative distribution function, CDF min is the minimum value of the cumulative distribution function.

[0012] Preferably, the bilinear interpolation algorithm comprises the following steps:

[0013] Step 1: Calculate the corresponding coordinates (x, y) of the point in the first image in the target image. d ,y d );

[0014] Step 2: Assuming the original image size is W*H, the coordinate transformation formula is expressed as follows:

[0015]

[0016] Step 3: Determine the four nearest neighboring pixels around (x′, y′), and set their coordinates to (i, j), (i+1, j), (i, j+1), and (i+1, j+1). The corresponding grayscale values are f(i, j), f(i+1, j), f(i, j+1), and f(i+1, j+1).

[0017] Step 4: Calculate the gray value f(xd, yd) of the point in the target image with coordinates (xd, yd) d ,y d ):

[0018] f(x d, y d)=(1-a)(1-b)f(i,j)+a(1-b)f(i+1,j)+(1-a)bf(i,j+1)+abf(i+1,j+1)(5)

[0019] Where: a = x'-i, b = y'-j.

[0020] Preferably, the noise reduction process uses a Gaussian filtering algorithm, and the Gaussian filtering algorithm uses a convolution kernel of size 3*3 to perform a convolution operation on the image.

[0021] Preferably, the expression recognition model adopts an improved MiniXception network model structure, including:

[0022] Convolution module: uses depth-wise separable convolution to extract features from the input image multiple times and output the first feature map;

[0023] Dynamic adaptive feature fusion module: performs multi-feature fusion on the first feature map output by the convolution module and outputs a fused feature map;

[0024] Pooling module: uses the maximum pooling and average pooling alternating method to process the fused feature map, reduce the resolution of the fused feature map, reduce the amount of data, and output the second feature map;

[0025] Fully connected module: Integrates the image features of the second feature map and outputs 7 nodes, corresponding to 7 expression categories, and calculates the probability of each expression through the softmax function.

[0026] Preferably, the dynamic adaptive module includes: multi-scale feature extraction branch: inputting the first feature map into multiple branches with different convolution kernel sizes in parallel; dynamic weight allocation: performing spatial dimension compression and channel number compression on each branch to obtain an attention map, and mapping the attention map to the range of [0, 1]; feature fusion and output: performing element-wise multiplication of the attention map and the first feature map, weighting the first feature map of each branch, and adding the weighted feature maps to obtain a fused feature map.

[0027] Preferably, before using the expression recognition model, the expression recognition model is trained first, the model parameters are optimized using the stochastic gradient descent algorithm, and the parameter update process is further optimized using the Adam optimizer. During the model training process, the cross entropy loss function is used to measure the difference between the model prediction and the true label, as shown in the following formula:

[0028]

[0029] Where: m is the number of samples, y is the true label of the sample, h θ (x i) is the model prediction probability for the i-th sample under the θ parameter.

[0030] Preferably, the pressure warning module includes a pressure identification unit and a safety warning unit. The pressure identification unit determines the miner's pressure state based on a predefined correspondence between expression and pressure, and the safety warning unit determines whether to trigger a warning according to the miner's pressure state.

[0031] Preferably, the expressions in the predefined expression-stress correspondence relationship include positive expressions, negative expressions and neutral expressions, the positive expression is happiness; the negative expressions include anger, sadness, disgust and fear; and the neutral expressions include normal and surprise.

[0032] Compared with the prior art, the present invention has the following advantages:

[0033] 1. The facial expression recognition model of the present invention adopts an improved MiniXception network model structure. The model uses deep separable convolution and dynamic adaptive feature fusion modules to efficiently extract facial features, improve the accuracy of image recognition, and solve the problem of low facial expression recognition accuracy in complex environments in the existing technology.

[0034] 2. The improved MiniXception model adopted in this invention is a lightweight network architecture design, which reduces the amount of computation while maintaining feature extraction capabilities, thereby increasing the image recognition rate and reducing data processing delays.

[0035] The present invention is further described in detail below through the accompanying drawings and examples. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a schematic diagram of the process of the present invention;

[0037] Figure 2 Schematic diagram of the corresponding relationship between facial expression and pressure of the present invention. DETAILED DESCRIPTION

[0038] like Figure 1 As shown, the present invention discloses a method for identifying work stress of miners based on facial expressions, comprising: continuously collecting a first image of the operator's face and caching the data, and outputting the first image based on a cache threshold; gray-scaling the first image using a weighted average method to obtain a gray-scale image; normalizing the gray-scale image using a bilinear interpolation algorithm and a histogram equalization algorithm to obtain a normalized image; denoising the normalized image using a Gaussian filtering algorithm to obtain a denoised image; extracting features from the denoised image using an expression recognition model to determine the expression type, and transmitting the expression type to a pressure warning module to determine whether to trigger a warning.

[0039] An explosion-proof camera is used to collect facial image data of miners to ensure image quality and continuity. The explosion-proof camera shoots at a frame rate of 25 frames per second and caches the collected image data to the local memory. When the memory storage data reaches a certain threshold (such as 50MB, 100MB, etc.) or at a certain time interval (such as 5 minutes, 10 minutes, etc.), the cached data is output as the first image; the resolution of the explosion-proof camera is not less than 1080P, and it has good low-light performance, autofocus and anti-shake functions; the explosion-proof camera is installed in key positions in the working area, such as near the cab of tunnel boring machines, scrapers and pickup trucks, and is debugged from multiple angles to ensure that the facial expressions of miners can be clearly captured; the camera uses wired transmission for data transmission to ensure the reliability and real-time performance of data transmission.

[0040] like Figure 2 As shown, the pressure warning module includes a pressure recognition unit and a safety warning unit. The pressure recognition unit determines the miner's pressure status based on the predefined correspondence between expression and pressure, and the safety warning unit determines whether to trigger a warning based on the miner's pressure status.

[0041] The data preprocessing module includes grayscale, normalization and noise reduction processing for the data. After receiving the first image, the data preprocessing module performs integrity check on the first image to check whether the data is lost or damaged. If there is a problem, it requests retransmission, and then grayscales the first image that passes the check, and calculates the grayscale value according to the weighted average method; then the bilinear interpolation algorithm is used to adjust the image size and perform size normalization; then the histogram equalization algorithm is used to process the image brightness and perform brightness normalization; finally, the Gaussian filter algorithm is used to remove noise and perform noise reduction processing; the processed image is cached locally and the expression recognition model is notified to process it; the expression recognition model receives the preprocessed image data After the notification, the image data is read and input into the trained expression recognition model for feature extraction and expression classification calculation. The expression recognition model outputs the probability value of each expression, selects the expression with the highest probability as the recognition result, and transmits the recognition result to the pressure warning module. After receiving the expression recognition result, the pressure warning module performs data verification again to ensure the accuracy of the result. According to the predefined correspondence between expression and pressure, the miner's stress status is judged; the stress status result is added with timestamp, miner number and other information, and transmitted to the mine management center server. After receiving the result, the server stores it in the database and pushes it to the management personnel terminal device in real time. If the miner is under stress, an alarm is triggered.

[0042] The weighted average method converts the pixel values of a color image into grayscale values according to preset weights based on the human eye's sensitivity to different colors, as shown in the following formula:

[0043] Gray=d1*R+d2*G+d3*B (1)

[0044] Where Gray represents the grayscale value of the pixel, d1 represents the red weight, R represents the red pixel value of the pixel, d2 represents the green weight, G represents the green pixel value of the pixel, d3 represents the blue weight, and B represents the blue pixel value of the pixel.

[0045] In this embodiment, the weights of the RGB values are determined according to the sensitivity of the human eye to different colors, such as d1 = 0.299, d2 = 0.587, and d3 = 0.114, and the color image is converted into a grayscale image, effectively reducing the amount of data while retaining information related to the expression features.

[0046] The normalization includes size normalization and brightness normalization, and the size normalization uses a bilinear interpolation algorithm to adjust the size of the grayscale image to 100*100 pixels;

[0047] The brightness normalization uses the histogram equalization algorithm to map the brightness value of the grayscale image to the range of [0, 1], as shown in the following formula:

[0048]

[0049] Where: (x, y) is the pixel coordinate, CDF max is the maximum value of the cumulative distribution function, CDF min is the minimum value of the cumulative distribution function.

[0050] Size normalization adjusts the image to a uniform size of 100*100 pixels, so that images of different sizes are consistent when input into the expression recognition model, making it easier for the expression recognition model to accurately extract features.

[0051] In formula (2), first, the number of pixels at each gray level (0-255) in the image is counted, and the cumulative distribution function is calculated to determine the minimum and maximum values of the cumulative distribution function, and finally brightness normalization is performed; the image brightness value is mapped to the range of [0,1] through the histogram equalization technology, which effectively eliminates the influence of illumination changes in the deep well environment on the image and enhances the contrast and recognizability of the image.

[0052] The bilinear interpolation algorithm comprises the following steps:

[0053] Step 1: Calculate the corresponding coordinates (x, y) of the point in the first image in the target image. d ,y d );

[0054] The target image is an image of 100*100 pixels.

[0055] Step 2: Assuming the original image size is W*H, the coordinate transformation formula is expressed as follows:

[0056]

[0057] Step 3: Determine the four nearest neighboring pixels around (x′, y′), and set their coordinates to (i, j), (i+1, j), (i, j+1), and (i+1, j+1). The corresponding grayscale values are f(i, j), f(i+1, j), f(i, j+1), and f(i+1, j+1).

[0058] Step 4: Calculate the gray value f(xd, yd) of the point in the target image with coordinates (xd, yd) d ,y d ):

[0059] f(x d, y d )=(1-a)(1-b)f(i,j)+a(1-b)f(i+1,j)+(1-a)bf(i,j+1)+abf(i+1,j+1)(5)

[0060] Where: a = x'-i, b = y'-j.

[0061] The grayscale value of a point in the target image is estimated through linear interpolation in the horizontal and vertical directions; the grayscale values of the four adjacent pixels are weighted averaged, taking into account the spatial position relationship and grayscale continuity between the pixels to avoid discontinuous pixel values.

[0062] The noise reduction process uses a Gaussian filtering algorithm, and the Gaussian filtering algorithm uses a 3*3 convolution kernel to perform a convolution operation on the image.

[0063] In this embodiment, a 3*3 convolution kernel is used as the convolution kernel. The 3*3 convolution kernel has a small amount of computation and low requirements for hardware equipment. It is more advantageous when resources are limited. The smaller 3*3 convolution kernel is faster when processing images and can output expression recognition results faster, meeting the needs of real-time monitoring of the miners' stress status. The image is convolved through the Gaussian filtering algorithm to smooth the image, remove noise, and retain expression features.

[0064] The expression recognition model adopts an improved MiniXception network model structure, including:

[0065] Convolution module: uses depth-wise separable convolution to extract features from the input image multiple times and output the first feature map;

[0066] In this embodiment, depthwise separable convolution includes depthwise convolution and pointwise convolution. The depthwise convolution performs convolution operations on each channel of the input layer independently, using a 3*3 convolution kernel, a step size of 1, and a padding of 1. The number of convolution kernels is consistent with the number of input channels. This setting can effectively capture local details in the input image, especially subtle features of facial expressions, such as changes in muscles around the eyes and mouth. Subsequently, the pointwise convolution uses a 1*1 convolution kernel to perform a weighted combination of the results obtained by the depthwise convolution in the depth direction to generate a first feature map. By acting in sequence on multiple such convolution modules, it is convenient to extract different levels of more abstract expression features from the input image.

[0067] Dynamic adaptive feature fusion module: performs multi-feature fusion on the first feature map output by the convolution module and outputs a fused feature map;

[0068] After each set of convolution modules, a dynamic adaptive feature fusion module is introduced. This module can extract facial expression features from multiple scales and dynamically fuse these multi-scale features, which helps to fully capture the subtle features, local details and macroscopic morphology of facial expressions, providing rich feature information for subsequent expression classification.

[0069] There are four depth-wise separable convolutions, and the distribution of the four depth-wise separable convolutions and dynamic adaptive feature fusion modules is as follows: depth-wise separable convolution-dynamic adaptive feature fusion-depth-wise separable convolution-dynamic adaptive feature fusion-depth-wise separable convolution-depth-wise separable convolution. The improved MiniXception network model adds dynamic adaptive feature fusion in the shallow layer to more effectively fuse local details and enhance multi-scale expression capabilities. In the deep stage of the network, features do not need to rely on dynamic adaptive feature fusion to integrate multi-scale information. At the same time, if dynamic adaptive feature fusion is continued to be used in the deep layer, it is easy to lead to excessive model complexity, increase the amount of calculation and the risk of overfitting. Therefore, using convolution stacking in the deep layer and continuing to use two depth-wise separable convolutions will help improve generalization and enhance the nonlinear expression capabilities of feature mapping.

[0070] The dynamic adaptive module includes:

[0071] Multi-scale feature extraction branch: the first feature map is input into multiple branches with different convolution kernel sizes in parallel;

[0072] Multiple different convolution kernels use 3*3, 5*5 and 7*7 convolution kernels respectively. The 3*3 convolution kernel is good at capturing fine details of facial expressions; the 5*5 convolution kernel is used to obtain medium-scale local features; the 7*7 convolution kernel focuses on more macroscopic facial expression forms. The first feature map enters the 3*3, 5*5 and 7*7 convolution kernels respectively for multi-scale feature extraction, thereby improving the accuracy of miners' expression recognition.

[0073] Dynamic weight allocation: After compressing the spatial dimension and the number of channels of each branch, the attention map is obtained and mapped to the range of [0, 1].

[0074] Perform global average pooling on the 3*3 convolution kernel, compress the spatial dimension of the first feature map to 1*1, and obtain the global statistical information of the 3*3 convolution kernel; then use the 1*1 convolution kernel to compress the number of channels to 1, and obtain the channel-level attention map of the 3*3 convolution kernel;

[0075] Perform global average pooling on the 5*5 convolution kernel, compress the spatial dimension of the first feature map to 1*1, and obtain the global statistical information of the 5*5 convolution kernel; then use the 1*1 convolution kernel to compress the number of channels to 1, and obtain the channel-level attention map of the 5*5 convolution kernel;

[0076] Perform global average pooling on the 7*7 convolution kernel, compress the spatial dimension of the first feature map to 1*1, and obtain the global statistical information of the 7*7 convolution kernel; then use the 1*1 convolution kernel to compress the number of channels to 1, and obtain the channel-level attention map of the 7*7 convolution kernel;

[0077] Finally, the sigmoid activation function is used to map the attention maps of the 3*3, 5*5, and 7*7 convolution kernels to the range of [0, 1] to obtain the importance weight of each channel.

[0078] Feature fusion and output: Perform element-wise multiplication of the attention map and the first feature map, weight the first feature map of each branch, and add the weighted feature maps to obtain the fused feature map.

[0079] The attention maps of the 3*3, 5*5 and 7*7 convolution kernels are respectively element-wise multiplied with the first feature map, and the first feature maps after element-wise multiplication are weighted to obtain the weighted feature maps; the weighted first feature maps of the 3*3, 5*5 and 7*7 convolution kernels are added together to achieve the fusion of features of different scales; the fused feature map is then passed through a 1*1 convolution layer to adjust the number of channels, and then output to the pooling module for the next step of processing. This dynamic and adaptive feature fusion method can flexibly integrate features of different scales according to the characteristics of the input image, significantly improving the network's ability to recognize complex expressions.

[0080] Pooling module: uses the maximum pooling and average pooling alternating method to process the fused feature map, reduce the resolution of the fused feature map, reduce the amount of data, and output the second feature map;

[0081] In this embodiment, the pooling kernel size is 2*2 and the step size is 2. By alternating between maximum pooling and average pooling, the key local features in the fusion feature map are highlighted while retaining the overall information.

[0082] Fully connected module: Integrates the image features of the second feature map and outputs 7 nodes, corresponding to 7 expression categories, and calculates the probability of each expression through the softmax function.

[0083] The integration of the fully connected module is to flatten the 3D vector of the second feature map into a 1D vector, then map it to the 7-dimensional output space, and use the softmax function to normalize the probability of the 7 nodes.

[0084] The probability of classifying x into class j in the softmax function is:

[0085]

[0086] Where, P(y (i) =j|x (i) ; θ) is the probability that image x corresponds to each pressure classification j, θ is the parameter to be fitted, k is the number of samples, and i represents the index of the current data in the dataset.

[0087] Before using the expression recognition model, the expression recognition model is trained first. The stochastic gradient descent algorithm is used to optimize the model parameters, and the Adam optimizer is used to further optimize the parameter update process. During the model training process, the cross entropy loss function is used to measure the difference between the model prediction and the true label, as shown in the following formula:

[0088]

[0089] Where: m is the number of samples, y is the true label of the sample, h θ (x i ) is the model prediction probability for the i-th sample under the θ parameter.

[0090] During model training, we first constructed a dataset by collecting facial expression images of miners from field experiments. These images covered different work scenarios, such as tunnel boring machines drilling tunnels, scrapers shoveling ore, and pickup truck drivers transporting ore. These images also covered different time periods, such as before the morning shift, during the morning shift, after the morning shift, before the middle shift, during the middle shift, after the middle shift, before the night shift, during the night shift, and after the night shift. The model took into account the changes in miners' emotions and stress at different work stages. The collected images were screened and annotated to ensure data accuracy. Furthermore, we introduced the Fer2013 facial expression dataset for data expansion, enriching data diversity and enhancing the model's generalization capabilities.

[0091] Then the model is trained and the stochastic gradient descent algorithm is used to optimize the model parameters. The learning rate is initially set to 0.001. As the training progresses, the learning rate decays to 0.1 times the original value every 30 epochs. The Adam optimizer is used to further optimize the parameter update process. The parameter update formula of Adam is:

[0092]

[0093] Where: t is the number of iterations, θ is the parameter to be updated, η is the learning rate, is the mean value of the gradient at the first moment, is the variance at the second moment, β1=0.9, β2=0.999, ε=le-8.

[0094] The cross-entropy loss function is used to measure the difference between the model prediction and the true label. The training is carried out for 120 epochs. After each epoch, the model performance is evaluated on the validation set. If the performance does not improve for 50 consecutive epochs, the training is terminated early to prevent overfitting.

[0095] The pressure warning module includes a pressure recognition unit and a safety warning unit. The pressure recognition unit determines the miner's pressure state based on the predefined correspondence between expression and pressure, and the safety warning unit determines whether to trigger a warning based on the miner's pressure state.

[0096] The expressions in the predefined expression-stress correspondence relationship include positive expressions, negative expressions and neutral expressions. The positive expression is happiness; the negative expression includes anger, sadness, disgust and fear; and the neutral expression includes normal and surprised.

[0097] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any way. Any simple modification, change and equivalent structural transformation made to the above embodiment based on the technical essence of the present invention shall still fall within the scope of protection of the technical solution of the present invention.

Claims

1. A method for identifying work stress of miners based on facial expressions, characterized in that: include: Continuously collecting a first image of the operator's face and caching the data, and outputting the first image based on a cache threshold; grayscale the first image using a weighted average method to obtain a grayscale image; Normalizing the grayscale image using a bilinear interpolation algorithm and a histogram equalization algorithm to obtain a normalized image; Performing noise reduction processing on the normalized image using a Gaussian filtering algorithm to obtain a noise-reduced image; The expression recognition model is used to extract features from the noise reduction image to determine the expression type, and the expression type is transmitted to the stress warning module to determine whether to trigger an early warning.

2. A method for identifying miners' work stress based on facial expressions according to claim 1, characterized in that: The weighted average method converts the pixel values of a color image into grayscale values according to preset weights based on the human eye's sensitivity to different colors, as shown in the following formula: Gray=d1*R+d2*G+d3*B (1) Where Gray represents the grayscale value of the pixel, d1 represents the red weight, R represents the red pixel value of the pixel, d2 represents the green weight, G represents the green pixel value of the pixel, d3 represents the blue weight, and B represents the blue pixel value of the pixel.

3. A method for identifying miners' work stress based on facial expressions according to claim 1, characterized in that: The normalization includes size normalization and brightness normalization, and the size normalization uses a bilinear interpolation algorithm to adjust the size of the grayscale image to 100*100 pixels; The brightness normalization uses the histogram equalization algorithm to map the brightness value of the grayscale image to the range of [0, 1], as shown in the following formula: Where: (x, y) is the pixel coordinate, CDF max is the maximum value of the cumulative distribution function, CDF min is the minimum value of the cumulative distribution function.

4. A method for identifying miners' work stress based on facial expressions according to claim 1, characterized in that: The bilinear interpolation algorithm comprises the following steps: Step 1: Calculate the corresponding coordinates (x, y) of the point in the first image in the target image. d ,y d ); Step 2: Assuming the original image size is W*H, the coordinate transformation formula is expressed as follows: Step 3: Determine the four nearest neighboring pixels around (x′, y′), and set their coordinates to (i, j), (i+1, j), (i, j+1), and (i+1, j+1). The corresponding grayscale values are f(i, j), f(i+1, j), f(i, j+1), and f(i+1, j+1). Step 4: Calculate the gray value f(xd, yd) of the point in the target image with coordinates (xd, yd) d ,y d ): f(x d, y d )=(1-a)(1-b)f(i,j)+a(1-b)f(i+1,j)+(1-a)bf(i,j+1)+abf(i+1,j+1) (5) Where: a = x'-i, b = y'-j.

5. A method for identifying miners' work stress based on facial expressions according to claim 1, characterized in that: The noise reduction process uses a Gaussian filtering algorithm, and the Gaussian filtering algorithm uses a 3*3 convolution kernel to perform a convolution operation on the image.

6. A method for identifying miners' work stress based on facial expressions according to claim 1, characterized in that: The expression recognition model adopts an improved MiniXception network model structure, including: Convolution module: uses depth-wise separable convolution to extract features from the input image multiple times and output the first feature map; Dynamic adaptive feature fusion module: performs multi-feature fusion on the first feature map output by the convolution module and outputs a fused feature map; Pooling module: uses the maximum pooling and average pooling alternating method to process the fused feature map, reduce the resolution of the fused feature map, reduce the amount of data, and output the second feature map; Fully connected module: Integrates the image features of the second feature map and outputs 7 nodes, corresponding to 7 expression categories, and calculates the probability of each expression through the softmax function.

7. A method for identifying miners' work stress based on facial expressions according to claim 6, characterized in that: The dynamic adaptive module includes: Multi-scale feature extraction branch: the first feature map is input into multiple branches with different convolution kernel sizes in parallel; Dynamic weight allocation: After compressing the spatial dimension and the number of channels of each branch, the attention map is obtained and mapped to the range of [0, 1]. Feature fusion and output: Perform element-wise multiplication of the attention map and the first feature map, weight the first feature map of each branch, and add the weighted feature maps to obtain the fused feature map.

8. A method for identifying miners' work stress based on facial expressions according to claim 1, characterized in that: Before using the expression recognition model, the expression recognition model is trained first. The stochastic gradient descent algorithm is used to optimize the model parameters, and the Adam optimizer is used to further optimize the parameter update process. During the model training process, the cross entropy loss function is used to measure the difference between the model prediction and the true label, as shown in the following formula: Where: m is the number of samples, y is the true label of the sample, h θ (x i ) is the model prediction probability for the i-th sample under the θ parameter.

9. A method for identifying miners' work stress based on facial expressions according to claim 1, characterized in that: The pressure warning module includes a pressure recognition unit and a safety warning unit. The pressure recognition unit determines the miner's pressure state based on a predefined correspondence between expression and pressure, and the safety warning unit determines whether to trigger a warning based on the miner's pressure state.

10. A method for identifying miners' work stress based on facial expressions according to claim 9, characterized in that: The expressions in the predefined expression-stress correspondence relationship include positive expressions, negative expressions and neutral expressions. The positive expression is happiness; Said negative expressions include anger, sadness, disgust and fear; The neutral expressions include normal and surprised.