Mine operator state identification method, equipment, medium and product

Through multimodal data preprocessing and multi-view Transformer+ fusion model, combining physiological signals and face images, the problem of large and time-consuming and labor-intensive state recognition errors among mining operators is solved, and efficient and accurate state recognition is achieved.

CN120495829AInactive Publication Date: 2025-08-15XIAN UNIV OF SCI & TECH

Patent Information

Application Number
CN202510985388.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-08-15
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing mining operator status identification methods have large errors and are time-consuming and labor-intensive, making it difficult to accurately identify negative emotions and psychological fatigue states.

Method used

Multimodal data preprocessing, feature extraction and multi-view Transformer+ fusion model are used to automatically identify the emotional state and fatigue state of mine operators in combination with physiological signals and face images.

Benefits of technology

It improves recognition efficiency, reduces errors, realizes objective and accurate status recognition, and shortens recognition time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495829A_ABST
    Figure CN120495829A_ABST
Patent Text Reader

Abstract

The invention discloses a mine worker state recognition method and device, a medium and a product, and relates to the field of emotion recognition, and the method comprises the steps: carrying out the preprocessing of the multi-modal data of a mine worker, and determining the preprocessed multi-modal data; the multi-modal data comprises a physiological signal and a face image; performing feature extraction on the preprocessed multi-modal data, and determining a physiological feature set and a facial feature set; inputting the physiological feature set and the facial feature set into a multi-view Transform + fusion model, and determining an emotional state and a fatigue state of the mine operator; the emotional state comprises happiness, sadness, fear and neutrality; the fatigue states comprise normal, mild, moderate and severe states, and the method can reduce the state recognition error of the mine operation personnel and improve the recognition efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of emotion recognition, and in particular to a method, device, medium and product for identifying the status of mining personnel. Background Art

[0002] The unsafe conditions experienced by workers in typical mining scenarios can be primarily categorized into two aspects: negative emotions and mental fatigue. Negative emotions, including tension, depression, anger, fatigue, and confusion, represent vulnerable, intrinsic human emotions. Miners, for example, are particularly susceptible to experiencing these negative emotions. This is primarily because they not only require complex technical knowledge but also operate in extreme environments, including high temperatures, high humidity, and high dust levels. Furthermore, in the context of smart mine development, workers' work in typical mining scenarios is gradually shifting from simple physical labor to more complex and diverse mental activities. This process requires workers to engage in cognitive activities for extended periods, leading to brain overload and mental fatigue. When workers experience mental fatigue, it negatively impacts their attention, reaction time, and judgment, increasing the risk of operational errors.

[0003] In typical real-world mining scenarios, the status of mine workers is primarily determined through subjective questionnaires and physical examination reports. However, these methods suffer from significant bias and are time-consuming and labor-intensive. Summary of the Invention

[0004] The purpose of this application is to provide a method, equipment, medium and product for identifying the status of mine workers, so as to solve the problem that the existing methods for identifying the status of mine workers have large errors and are time-consuming and labor-intensive.

[0005] To achieve the above objectives, this application provides the following solutions.

[0006] In a first aspect, the present application provides a method for identifying the status of mining personnel, comprising the following steps.

[0007] Preprocessing is performed on multimodal data of mine workers to determine preprocessed multimodal data; the multimodal data includes physiological signals and facial images.

[0008] Feature extraction is performed on the preprocessed multimodal data to determine a physiological feature set and a facial feature set.

[0009] The physiological feature set and the facial feature set are input into a multi-view Transformer+ fusion model to determine the emotional state and fatigue state of the mine worker; the emotional state includes happiness, sadness, fear, and neutral; the fatigue state includes normal, mild, moderate, and severe.

[0010] In a second aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned method for identifying the status of mining personnel.

[0011] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the above-mentioned method for identifying the status of mining personnel when executed by a processor.

[0012] In a fourth aspect, the present application provides a computer program product, including a computer program, which implements the above-mentioned method for identifying the status of mine workers when executed by a processor.

[0013] According to the specific embodiments provided in this application, this application has the following technical effects: Based on the technical solution of this application, during the recognition process, as long as the multimodal data of the mine workers is input into the multi-view Transformer+ fusion model, the emotional state and fatigue state of the mine workers can be automatically identified. Compared with questionnaires and physical examination reports, this application objectively identifies the emotional state and fatigue state, and the recognition result error is smaller, which greatly shortens the recognition time and is more efficient. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0015] Figure 1 This is an application environment diagram of a mine operator status identification method in one embodiment of the present application.

[0016] Figure 2 A flowchart of a method for identifying the status of mine workers provided in one embodiment of the present application.

[0017] Figure 3 This is a block diagram of a method for identifying the status of mine workers in one embodiment of the present application.

[0018] Figure 4 A schematic diagram of the GRU structure provided in one embodiment of the present application.

[0019] Figure 5 A schematic diagram of the multi-view Transformer+ fusion model structure provided in one embodiment of the present application.

[0020] Figure 6 A schematic diagram of the hardware device framework provided in one embodiment of the present application.

[0021] Figure 7 A schematic diagram of a cross-entropy loss function curve based on physiological characteristics provided in one embodiment of the present application.

[0022] Figure 8 A schematic diagram of the cross entropy loss function curve for facial features provided in one embodiment of the present application.

[0023] Figure 9 A schematic diagram of the cross-entropy loss function curve under physiological features + facial features provided in one embodiment of the present application.

[0024] Figure 10 A schematic diagram of the online operation interface of a mine operator status identification system corresponding to the mine operator status identification method provided in one embodiment of the present application.

[0025] Figure 11 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0026] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0027] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0028] The method for identifying the status of mine workers provided in the embodiment of the present application can be applied to Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the video to be processed to the server 104. After the server 104 receives the video to be processed, the server 104 pre-processes the multimodal data of the mine workers for the video to be processed to determine the pre-processed multimodal data; the multimodal data includes physiological signals and facial images; feature extraction is performed on the pre-processed multimodal data to determine the physiological feature set and the facial feature set; the physiological feature set and the facial feature set are input into the multi-view Transformer+ fusion model to determine the emotional state and fatigue state of the mine workers; the emotional state includes happiness, sadness, fear and neutral; the fatigue state includes normal, mild, moderate and severe. The server 104 can feed back the obtained multimodal data to the terminal 102. In addition, in some embodiments, the method for identifying the status of mine workers can also be implemented independently by the server 104 or the terminal 102. For example, the terminal 102 can directly perform mine worker status identification on the multimodal data, or the server 104 can obtain the multimodal data to be identified from the data storage system and perform mine worker status identification on the multimodal data to be identified.

[0029] Terminal 102 may include, but is not limited to, various desktop computers, laptops, smartphones, tablet computers, IoT devices, and portable wearable devices. IoT devices may include smart speakers, smart TVs, smart air conditioners, and smart car devices. Portable wearable devices may include smart watches, smart bracelets, and head-mounted devices. Server 104 may be implemented as a standalone server or a server cluster consisting of multiple servers, or may be a cloud server.

[0030] In an exemplary embodiment, Figure 2 As shown, a method for identifying the status of mine workers is provided. The method is executed by a computer device, specifically a computer device such as a terminal or a server, or a terminal and a server. In the embodiment of the present application, the method is applied to Figure 1 The server 104 in the example is used for explanation, including the following S1-S3.

[0031] S1: Preprocessing multimodal data of mine workers to determine preprocessed multimodal data; the multimodal data includes physiological signals and facial images.

[0032] S2: Perform feature extraction on the preprocessed multimodal data to determine a physiological feature set and a facial feature set.

[0033] S3: Input the physiological feature set and the facial feature set into a multi-view Transformer+ fusion model to determine the emotional state and fatigue state of the mine worker; the emotional state includes happiness, sadness, fear, and neutral; the fatigue state includes normal, mild, moderate, and severe.

[0034] In practical applications, the objective, reliable and accurate state recognition is achieved by integrating the physiological signals (electrocardiogram, skin conduction, blood oxygen saturation, body temperature and blood pressure) and facial expression (face image) of mining workers. The main contents of the method include data preprocessing, feature extraction and feature fusion. The framework diagram of the mining worker state recognition method of this application is shown in the figure below. Figure 3 shown.

[0035] In an exemplary embodiment, S1 may be replaced by the following steps, wherein the preprocessed multimodal data includes preprocessed physiological signals and preprocessed facial images.

[0036] S11: intercepting valid data of the physiological signals; the physiological signals include electrocardiogram signals, skin electrical signals, blood oxygen saturation signals, body temperature and blood pressure signals.

[0037] S12: reducing the sampling frequency of the effective data, and using a wavelet packet decomposition method to filter and reduce noise on the effective data collected at the sampling frequency, to determine a preprocessed physiological signal.

[0038] S13: Using a Lanczos interpolation method to adjust the size of the facial image, and using a maximum-minimum normalization method to normalize the pixels of the facial image to determine a preprocessed facial image.

[0039] In practical applications, physiological data is often the reflection of electrical activity on the body's surface. These signals are weak, contain some invalid data, and are susceptible to interference from power frequency and noise. Therefore, to accurately analyze the desired phenomena, it is necessary to preprocess the raw physiological data.

[0040] The physiological signal preprocessing process is as follows: First, valid data is captured. To ensure data quality, missing or abnormal data must be eliminated, retaining valid data. Second, the sampling frequency is reduced. To address data consistency issues and improve overall efficiency, it is necessary to reduce the initial sampling frequency based on the Shannon sampling theorem. Finally, noise is filtered. ECG and skin conductance data are waveform signals and are more susceptible to external environmental factors and data acquisition equipment. Therefore, wavelet packet decomposition is used for filtering and noise reduction.

[0041] In practical applications, to ensure the effectiveness of feature extraction models, facial image preprocessing primarily involves pixel resizing and normalization. Facial data preprocessing is performed as follows: First, to prevent deformation and improve model accuracy, resizing is performed to adjust all facial images to a fixed size. According to the ResNeXt-101 requirements, the input facial image size should be 224×224 dpi, and resizing is performed using the Lanczos interpolation method. Second, pixel normalization is performed using the maximum-minimum normalization method, which linearly scales pixels to a range of 0 to 1. This method enables comparison between different facial images and reduces the impact of pixel distribution differences on model training.

[0042] In an exemplary embodiment, S2 may be replaced by the following steps.

[0043] S21: performing time domain analysis and frequency domain analysis on the preprocessed physiological signal to determine the time domain characteristics and frequency domain characteristics of the preprocessed physiological signal.

[0044] S22: Select the time domain features and frequency domain features using a maximum correlation minimum redundancy and recursive feature elimination algorithm to construct a physiological feature set.

[0045] S23: Using a convolutional neural network to extract spatial information of the preprocessed facial image.

[0046] S24: Based on the spatial information, a temporal feature extraction network is used to extract temporal features of the preprocessed facial image.

[0047] S25: Constructing a facial feature set according to the temporal features.

[0048] In practical applications, feature extraction includes the following steps.

[0049] (1) Physiological feature extraction and selection.

[0050] Physiological signal analysis methods are primarily categorized into time domain analysis and frequency domain analysis. Specifically, time domain waveforms contain complete temporal information, allowing for more accurate observation of changes in both abnormal and regular waveforms. Frequency domain analysis converts physiological data from the time domain to the frequency domain, revealing its distribution at different frequencies. Therefore, physiological feature extraction primarily utilizes time and frequency domain analysis methods. Furthermore, feature selection is essential for optimizing the physiological feature set, taking into account correlation, redundancy, irrelevance, and interactivity between features.

[0051] The time domain characteristics of the physiological signal include: maximum value, minimum value, average value, peak-to-peak value, average value of absolute value, variance, standard deviation, tilt, skewness, root mean square, form factor, peak factor, pulse factor and margin factor.

[0052] The frequency domain characteristics of the physiological signal include: center of gravity frequency, frequency mean, frequency standard deviation, and frequency root mean square.

[0053] Feature selection: The maximum relevance minimum redundancy (MRMV) and recursive feature elimination (RFE) algorithms are used to select and retain physiological features of high importance. Specifically, the MRMV algorithm minimizes the redundancy of a feature subset, maximizes the correlation between the feature subset and the response variable, and identifies several features in the feature space that have the highest correlation and the lowest redundancy.

[0054] Therefore, the maximum relevance minimum redundancy algorithm was used to sort the original physiological features and select the most important features as the initial feature set. The recursive feature elimination algorithm aims to select the optimal feature subset by recursively training the model, evaluating feature importance, and removing features. Therefore, the recursive feature elimination algorithm was used to iteratively remove the least important features from the initial feature set, ultimately obtaining the optimal physiological feature set. The discrete choice model (Logit) was selected as the base model due to its strong interpretability, high stability, and low complexity to ensure accurate and efficient feature selection.

[0055] (2) Facial feature extraction.

[0056] Convolutional neural networks (CNNs) have proven to be an effective method for face recognition. Their advantage lies in their ability to automatically extract facial features and perform efficient and accurate classification and recognition. However, miners' emotional and fatigue states fluctuate over time. Recurrent neural networks possess memory and contextual association capabilities, enabling them to establish connections between consecutive frames. Therefore, this method uses the ResNeXt-101 with multiple convolutional layers to extract rich facial features, while employing a simple gated recurrent unit (GRU) to capture the dependencies between these features.

[0057] ResNeXt-101 combines residual networks with multi-channel convolution operations and introduces parallel connections, ensuring that each convolutional layer can extract features from different perspectives, better capturing complex details and contextual information in the image. Furthermore, ResNeXt-101 boasts more powerful nonlinear modeling capabilities and higher efficiency than convolutional neural networks. It constructs a deep network by stacking multiple residual blocks, reducing the number of parameters and computational complexity while maintaining accuracy. The ResNeXt-101 architecture is shown in Table 1. Conv2, conv3, conv4, and conv5 consist of 3, 4, 23, and 3 residual blocks, respectively.

[0058] Table 1 ResNeXt-101 structure

[0059] GRU is a structure of recurrent neural network that controls the flow of information by introducing update gate and reset gate, such as Figure 4 As shown in the figure, z is the update gate; σ is the sigmoid activation function; r is the reset gate; h is the hidden gate; multiplication is multiplication; sum is addition; splice is concatenation; Tanh is the hyperbolic tangent function; and X is the input feature. The update gate determines the weight between the input of the current time step and the hidden state of the previous time step, affecting whether information is retained or forgotten. On the other hand, the reset gate selectively combines the previous hidden state with the current input to better capture long-term dependencies in the sequence.

[0060] In an exemplary embodiment, S3 may be replaced by the following steps, wherein the multi-view Transformer+ fusion model includes a multi-view embedding module and a Transformer+ model.

[0061] S31: Based on the multi-view embedding module, the physiological features in the physiological feature set and the facial features in the facial feature set are fused, and the fused features are converted into multiple different views to generate a fused feature embedding sequence.

[0062] S32: Based on the Transformer+ model, the fused feature embedding sequence is processed using the self-attention mechanism, and according to the processed embedding sequence, the emotional state decoder and the fatigue state decoder decode the physiological features and the facial features to output the emotional state and the fatigue state.

[0063] In an exemplary embodiment, S31 may be replaced by the following steps, wherein the multi-view embedding module includes a linear layer and a batch normalization layer.

[0064] S311: Utilize the linear layer to convert the physiological features and the facial features into an embedding sequence; the embedding sequence includes a physiological feature embedding sequence and a facial feature embedding sequence.

[0065] S312: Processing the facial feature embedding sequence using an activation function to determine a processed facial feature embedding sequence.

[0066] S313: Reconstruct and stack the physiological feature embedding sequence and the processed facial feature embedding sequence to generate a fused feature.

[0067] S314: Based on the batch normalization processing layer, batch normalization is performed on the fused features, converted into multiple different views, and a fused feature embedding sequence is generated.

[0068] In practical applications, the fusion method based on feature connection ignores the correlation and complementarity between different features, resulting in the loss of effective information. In order to make full use of all key information, this application introduces a multi-view Transformer+ fusion model after fully connecting physiological features and facial features. Figure 5 As shown in Figure 3, the multi-view Transformer+ fusion model consists of a multi-view embedding module and a Transformer+ model.

[0069] The multi-view embedding module converts input physiological and facial features into multiple embeddings, encouraging the model to focus on different views of the feature. First, a linear layer converts the input features into an embedding sequence. Then, an activation function is applied to retain useful information in the embedding sequence while removing useless and redundant information. Finally, the embedding sequence is reconstructed, stacked, and normalized to generate a fused feature embedding sequence. In this way, a single input feature is converted into a series of tokens from different views that can be further processed by the subsequent fusion model.

[0070] The core component of the Transformer+ model is multi-head self-attention. First, three linear layers transform the embedding sequence into the query, key, and value vectors of the self-attention mechanism. The concatenation of these self-attention mechanisms, known as multi-head self-attention, improves the model's expressiveness and performance. Then, to address fatigue and emotion, the traditional Transformer model, consisting of one encoder and one decoder, is modified to one encoder and two decoders (corresponding to the two states). Finally, residual connections and layer normalization are introduced to further enhance the model's expressiveness.

[0071] The encoder is used to encode the fused feature embedding sequence and determine the processed embedding sequence.

[0072] The two decoders include an emotional state decoder and a fatigue state decoder, which are used to combine the processed embedding sequences to generate different emotional states and fatigue states. The specific hierarchical results are as follows.

[0073] Multi-head attention: collects information about how each feature in the embedding sequence is related to other features.

[0074] Feedforward fully connected network: performs linear transformation and nonlinear activation on the embedded sequence to further extract useful features.

[0075] Concatenation layer with normalization: concatenates the embedding sequences and normalizes the output of all embedding sequences.

[0076] Linear layer: maps the output of the decoder to the final output dimension.

[0077] Softmax: Converts the output of the linear layer into a probability distribution for final prediction.

[0078] In an exemplary embodiment, before S32, the following steps are further included: Different emotion labels and fatigue labels are assigned to historical physiological features and historical facial features.

[0079] The emotional state decoder is trained based on the historical physiological features, the historical facial features, and the corresponding emotional labels.

[0080] The fatigue state decoder is trained according to the historical physiological features, the historical facial features, and the corresponding fatigue labels.

[0081] In an exemplary embodiment, S32 may be replaced by the following steps.

[0082] S321: Convert the fused feature embedding sequence into a query vector, a key vector, and a value vector in a self-attention mechanism through a linear layer.

[0083] S322: Based on the connection layer, the normalization layer, and the feed-forward fully connected layer, determine a normalization processing sequence according to the query vector, the key vector, and the value vector.

[0084] S323: Decoding the physiological features and the facial features using the emotional state decoder in combination with the normalization processing sequence, and outputting the emotional state.

[0085] S324: Decoding the physiological features and the facial features using the fatigue status decoder in combination with the normalization processing sequence, and outputting a fatigue status.

[0086] In practical applications, the following hardware devices are used to implement the technical solution of this application.

[0087] The hardware device consists of sensors, core processors and display modules. Its structural block diagram is as follows: Figure 6 In addition, to facilitate implementation, only commercial products are used. The module information of the hardware equipment is shown in Table 2.

[0088] Table 2 Hardware device module information table

[0089] 1. The sensor array utilizes a high-precision sensor array to collect physiological and facial information from miners. These sensors include an electrocardiogram (ECG) sensor, a galvanic skin sensor, a blood pressure and oxygen saturation (O2) sensor, a temperature sensor, and a camera. In addition to the camera, the physiological data sensors are placed on a single plane. The arrangement of the sensors on this plane corresponds to the tops of six fingers, modeled after the shape of a human palm. The ECG and galvanic skin sensors each occupy two finger positions, while the blood pressure and oxygen saturation (O2) and temperature sensors each occupy one finger position. A camera is embedded above the display to automatically collect facial data. All multimodal data, including ECG, galvanic skin, blood pressure, O2 saturation, temperature, and facial images, is collected and fed into the core processor.

[0090] 2. The core processor builds a multimodal feature extraction and fusion framework based on the Python environment. ECG, blood pressure, and blood oxygen sensors, as well as cameras, transmit raw multimodal data via the UART communication interface between the sensors and the core processor. However, for galvano-dermal and body temperature sensors, the raw data requires conversion via an analog-to-digital converter (ADC) component before data transmission via the UART. The core processor is connected to a display via an HDMI cable to visualize both recognition results and raw data waveforms. Furthermore, a storage unit records and stores each mine worker's multimodal data, as well as emotion and fatigue status recognition results.

[0091] 3. The display provides a visual interface showing emotion classification, fatigue level, and waveform graphs. Emotion classification and fatigue level provide a clear understanding of the mine worker's current state, while the waveform graph intuitively displays the collected multimodal data.

[0092] Based on the technical solution of this application, the offline test results are as follows: During offline testing, parameter settings were adjusted multiple times to determine the optimal parameters. Finally, the network parameters were optimized using the gradient descent-based Adam optimizer, with a learning rate set to 0.001 to ensure rapid convergence and avoid losing the optimal solution. Furthermore, this offline test performed 454 training iterations, each of which involved batching four samples from the training set and updating the network weights.

[0093] Using cross entropy as the loss function, the curve of the model loss function under different input features is as follows Figure 7-Figure 9 As shown in the figure, compared to single physiological or facial features, the fusion of multiple features in the Multi-view Transformer+ fusion model results in a large number of parameters and relatively slow convergence. However, after 100 iterations, the average loss shows a clear downward trend and remains stable after 300 iterations, further confirming the effectiveness of the Multi-view Transformer+ fusion model.

[0094] The cross entropy is used as the loss function for: ;in, is the true label, Predict labels for the model, is the sample size, i is the sample serial number, is the true label of the i-th sample, is the predicted label of the i-th sample.

[0095] Offline testing was conducted over five training runs, ensuring that randomly selected samples were used in each training run. Tables 3 and 4 show the accuracy and variance of emotion and fatigue state recognition. When the multi-view Transformer+fusion model inputs both physiological and facial features, the accuracy and variance of emotion and fatigue state recognition were 93.79% ± 0.95 and 94.97% ± 0.96, respectively. Using a single physiological feature to identify miners' emotion and fatigue states achieved an accuracy of 80.83% ± 1.02 and a variance of 81.49% ± 0.49, respectively. Using a single facial feature achieved an accuracy of 88.95% ± 0.83 and a variance of 89.21% ± 1.28.

[0096] Table 3 Emotion recognition accuracy

[0097] Table 4 Fatigue recognition accuracy

[0098] Compared to using a single physiological feature or a single facial feature, fusing physiological and facial features significantly improves recognition accuracy. Specifically, the fusion method achieves higher accuracy rates for emotional state recognition of 12.96% and 4.84%, respectively, surpassing the accuracy achieved using physiological and facial features alone. Similarly, for fatigue state recognition, the fusion method outperforms using either physiological or facial features alone, achieving accuracy rates of 13.48% and 5.76%, respectively. This clearly demonstrates that multimodal feature fusion can leverage the strengths of all six modalities, resulting in more comprehensive and accurate state recognition information and deeper correlations and patterns. Consequently, it improves accuracy.

[0099] In order to help mine workers understand their current mood and fatigue status more intuitively, an online operation interface can be developed through software design. Figure 10 As shown, it can display multimodal data waveforms, facial images, emotional states, and fatigue states.

[0100] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above-mentioned method embodiments. The computer device can be a server or a terminal, and its internal structure can be as shown in FIG. Figure 11 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store data to be processed. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for identifying the status of a mine operator is implemented.

[0101] Those skilled in the art will understand that Figure 11The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application and does not constitute a limitation on the computer device to which the solution of the present application is applied. A specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps of the above-mentioned method embodiments when executing the computer program.

[0102] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0103] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0104] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0105] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0106] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.

[0107] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0108] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A method for identifying the status of mine workers, characterized in that: include: Preprocessing multimodal data of mine workers to determine preprocessed multimodal data; the multimodal data includes physiological signals and facial images; Performing feature extraction on the preprocessed multimodal data to determine a physiological feature set and a facial feature set; The physiological feature set and the facial feature set are input into a multi-view Transformer+ fusion model to determine the emotional state and fatigue state of the mine worker; the emotional state includes happiness, sadness, fear, and neutral; the fatigue state includes normal, mild, moderate, and severe.

2. The method for identifying the status of mine workers according to claim 1, characterized in that: Preprocessing the multimodal data of the mine workers to determine preprocessed multimodal data, specifically including: the preprocessed multimodal data including preprocessed physiological signals and preprocessed facial images; Intercepting valid data of the physiological signals; the physiological signals include electrocardiogram signals, skin electrical signals, blood oxygen saturation signals, body temperature and blood pressure signals; reducing the sampling frequency of the valid data, and filtering and reducing noise on the valid data collected at the sampling frequency using a wavelet packet decomposition method to determine a preprocessed physiological signal; The size of the facial image is adjusted using a Lanczos interpolation method, and the pixels of the facial image are normalized using a maximum-minimum normalization method to determine a preprocessed facial image.

3. The method for identifying the status of mine workers according to claim 2, characterized in that: Performing feature extraction on the pre-processed multimodal data to determine a physiological feature set and a facial feature set, specifically including: Performing time domain analysis and frequency domain analysis on the preprocessed physiological signal to determine the time domain characteristics and frequency domain characteristics of the preprocessed physiological signal; The time domain features and frequency domain features are selected by using the maximum correlation minimum redundancy and recursive feature elimination algorithm to construct a physiological feature set; Using a convolutional neural network to extract spatial information of the preprocessed facial image; Based on the spatial information, a temporal feature extraction network is used to extract temporal features of the preprocessed facial image; A facial feature set is constructed according to the temporal features.

4. The method for identifying the status of mine workers according to claim 1, characterized in that: The physiological feature set and the facial feature set are input into a multi-view Transformer+ fusion model to determine the emotional state and fatigue state of the mine worker, specifically including: The multi-view Transformer+ fusion model includes a multi-view embedding module and a Transformer+ model; Based on the multi-view embedding module, the physiological features in the physiological feature set and the facial features in the facial feature set are fused, and the fused features are converted into multiple different views to generate a fused feature embedding sequence; Based on the Transformer+ model, the self-attention mechanism is used to process the fused feature embedding sequence, and according to the processed embedding sequence, the emotional state decoder and the fatigue state decoder decode the physiological features and the facial features to output the emotional state and the fatigue state.

5. The method for identifying the status of mine workers according to claim 4, characterized in that: Based on the multi-view embedding module, the physiological features in the physiological feature set and the facial features in the facial feature set are fused, and the fused features are converted into multiple different views to generate a fused feature embedding sequence, specifically including: The multi-view embedding module includes a linear layer and a batch normalization layer; The linear layer is used to convert the physiological features and the facial features into an embedding sequence; the embedding sequence includes a physiological feature embedding sequence and a facial feature embedding sequence; Processing the facial feature embedding sequence using an activation function to determine a processed facial feature embedding sequence; Reconstructing and stacking the physiological feature embedding sequence and the processed facial feature embedding sequence to generate a fused feature; Based on the batch normalization layer, the fused features are batch normalized and converted into a plurality of different views to generate a fused feature embedding sequence.

6. The method for identifying the status of mine workers according to claim 4, characterized in that: Based on the Transformer+ model, the fused feature embedding sequence is processed using a self-attention mechanism. Based on the processed embedding sequence, the emotional state decoder and the fatigue state decoder decode the physiological features and the facial features to output the emotional state and the fatigue state. Previously, the following steps were also included: Assign different emotion labels and fatigue labels to historical physiological features and historical facial features; training the emotional state decoder based on the historical physiological features, the historical facial features, and the corresponding emotional labels; The fatigue state decoder is trained according to the historical physiological features, the historical facial features, and the corresponding fatigue labels.

7. The method for identifying the status of mine workers according to claim 6, characterized in that: Based on the Transformer+ model, the self-attention mechanism is used to process the fused feature embedding sequence. Based on the processed embedding sequence, the emotional state decoder and the fatigue state decoder decode the physiological features and the facial features, and output the emotional state and fatigue state, which specifically include: Convert the fused feature embedding sequence into a query vector, a key vector, and a value vector in a self-attention mechanism through a linear layer; Determining a normalization processing sequence according to the query vector, the key vector, and the value vector based on a connection layer, a normalization layer, and a feed-forward fully connected layer; Decoding the physiological features and the facial features using the emotional state decoder in combination with the normalization processing sequence to output an emotional state; In combination with the normalization processing sequence, the fatigue state decoder is used to decode the physiological features and the facial features to output the fatigue state.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for identifying the status of a mine operator according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for identifying the status of a mine operator according to any one of claims 1 to 7 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for identifying the status of a mine operator according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Flower image classification algorithm based on feature enhancement and decision fusion

    CN115099294A

  • Operating personnel on-line multi-mode identification system based on multi-mode feature fusion

    CN116226715A

  • Driver fatigue detection method based on time-space double characteristics

    CN116665191A

  • Multi-modal emotion recognition model training method and system and electronic equipment

    CN117171626A

  • Miner sign evaluation device and method based on multi-modal information fusion

    CN119007076A

Cited By

  • Miner abnormal psychological state grading early warning method

    CN120748742A