A method for classifying and early warning of abnormal psychological states in miners

By using a multimodal signal fusion method, the problem of rapid and accurate early warning of miners' psychological state in complex environments has been solved, achieving scientific and accurate early warning of miners' psychological state and improving safety management and production safety.

CN120748742BActive Publication Date: 2025-11-14XIAN UNIV OF SCI & TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511220625.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-11-14
Estimated Expiration
2045-08-29

AI Technical Summary

Technical Problem

Existing technologies cannot quickly and accurately provide early warnings and classifications of miners' psychological states in complex environments, and single-dimensional feature extraction methods cannot fully characterize miners' psychological stress responses in complex environments.

Method used

A multimodal signal fusion method is adopted to extract features of EEG signals, peripheral physiological signals, and speech signals in the time domain, frequency domain, time-frequency domain, and spatial domain, respectively. Cross-modal alignment is performed through optimal transmission theory, maximum mean difference algorithm, and kernel function. The features are then mapped to a unified-dimensional embedding space for comparative learning. The features are fused by combining bidirectional long short-term memory network and GhostNet network, and finally input into the psychological hierarchical early warning model for early warning.

Benefits of technology

It enables scientific, accurate, and early warning of miners' psychological state, improves the overall and humanistic level of safety management, promotes proactive intervention and resource optimization, and safeguards miners' physical and mental health and production safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748742B_ABST
    Figure CN120748742B_ABST
Patent Text Reader

Abstract

This invention discloses a method for graded early warning of abnormal psychological states in miners. It extracts and filters features from three modal signals processed in the time domain, frequency domain, time-frequency domain, and spatial domain to obtain corresponding first target signal feature sets. Based on optimal transmission theory, the maximum mean difference algorithm, and a selected kernel function, the first target signal feature sets are aligned locally and globally to obtain second target signal feature sets. These second target signal feature sets are then mapped to a unified, preset-dimensional embedding space for comparative learning to obtain a third target signal feature set. The third target signal feature set is then weighted, summed, and input into a bidirectional long short-term memory network to obtain initial fused features. These features are then input into a GhostNet network to obtain fused features, which are then input into the psychological state recognition module of a psychological graded early warning model to determine the miner's psychological state. Finally, these features, combined with mine environment parameters, are input into the early warning grading module of the psychological graded early warning model to obtain early warning information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of miners' psychological state analysis technology, and relates to, but is not limited to, a method for classifying and warning of abnormal psychological states in miners. Background Technology

[0002] The mechanical environment and reservoir structure of deep coal mining differ significantly from those of shallow mining, and the spatiotemporal relationships within the mining area are more complex, undoubtedly increasing the difficulty of safety management in mining operations. Statistical data shows that human-related accidents account for a high proportion of coal mine safety accidents, and abnormal psychological states are one of the main contributing factors to unsafe behaviors. Miners work long hours in enclosed, high-risk underground environments, facing significant psychological stress and frequent mood swings. These abnormal psychological states significantly reduce their risk perception and safe operational skills, increasing the accident rate.

[0003] In related technologies, the power spectral density and differential entropy characteristics of EEG signals are analyzed to identify brain activity patterns under different emotional states. Further EEG signal features are extracted for subsequent psychological grading and early warning. However, this method mainly targets single-dimensional features (such as time domain, frequency domain, or spatial domain features). Considering the challenges brought by the complex environment of mining areas, single-dimensional feature extraction methods cannot fully characterize the psychological stress response of miners in complex environments.

[0004] Therefore, how to quickly and accurately classify and issue early warnings about miners' psychological state has become an urgent problem to be solved. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a method for classifying and warning of abnormal psychological states of miners, which at least solves the problem that related technologies cannot quickly and accurately classify and warn of the psychological states of miners.

[0006] According to a first aspect of the present invention, a method for graded early warning of abnormal psychological states in miners is provided, comprising:

[0007] Feature extraction and filtering are performed on the three modal signals after processing in the time domain, frequency domain, time-frequency domain, and spatial domain respectively to obtain the corresponding first target signal feature set. The first target signal feature set includes the EEG signal feature set, the peripheral physiological signal feature set, and the speech signal feature set; the three modal signals include EEG signals, peripheral physiological signals, and speech signals.

[0008] The first target signal feature set is locally aligned across modes and globally based on optimal transmission theory, maximum mean difference algorithm and selected kernel function, respectively, to obtain the second target signal feature set;

[0009] The second target signal feature set is mapped to a unified embedding space of a preset dimension for comparative learning processing to obtain the third target signal feature set; the third target signal feature set is then weighted and summed and input into a bidirectional long short-term memory network to obtain initial fusion features with context awareness.

[0010] The initial fused features are input into the GhostNet network to obtain the fused features; and the fused features are input into the psychological state recognition module in the psychological grading early warning model to obtain the miner's psychological state.

[0011] The psychological state and mine environment parameters are input into the early warning classification module of the psychological classification early warning model to obtain early warning information.

[0012] According to a second aspect of the present invention, an electronic device is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; the memory is used to store at least one executable instruction, wherein the executable instruction causes the processor to perform an operation corresponding to the method described in the first aspect.

[0013] According to a third aspect of the present invention, a computer storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.

[0014] The solution provided in this invention performs feature extraction and filtering on the three modal signals after processing in the time domain, frequency domain, time-frequency domain, and spatial domain, respectively, to obtain a first target signal feature set corresponding to each. The first target signal feature set includes an EEG signal feature set, a peripheral physiological signal feature set, and a speech signal feature set; the three modal signals include EEG signals, peripheral physiological signals, and speech signals.

[0015] The first target signal feature set is locally aligned across modes and globally based on optimal transmission theory, maximum mean difference algorithm and selected kernel function, respectively, to obtain the second target signal feature set;

[0016] The second target signal feature set is mapped to a unified, preset-dimensional embedding space for comparative learning processing to obtain a third target signal feature set. The third target signal feature set is then weighted and summed before being input into a bidirectional long short-term memory network to obtain initial fused features with context awareness. These initial fused features are then input into a GhostNet network to obtain fused features. The fused features are then input into the psychological state recognition module of a psychological grading early warning model to obtain the miner's psychological state. The psychological state and mine environment parameters are then input into the early warning grading module of the psychological grading early warning model to obtain early warning information. In this process, the differences in multimodal physiological signals from the time domain, frequency domain, time-frequency domain, and spatial domain are analyzed. Temporal features, frequency distribution characteristics, time-varying frequency characteristics, and spatial distribution characteristics are extracted to comprehensively characterize the psychological stress response of workers in complex mining environments. A multi-angle cross-modal feature alignment strategy and a feature set contribution weight allocation method are adopted to establish a dynamic fusion method for multimodal features containing spatiotemporal dependencies, overcoming the limitations of single-modal signals and balancing the contributions of each modality. Based on the fusion results of miners' multimodal features, a basic network architecture including temporal analysis, convolutional operations, and dense connections was designed to construct a psychological state recognition module. Environmental parameters from typical mining scenarios were introduced as prior knowledge to construct an early warning grading standard. A psychological grading early warning model was built by integrating the psychological state recognition and early warning grading modules. By inputting psychological states and mining environment parameters into the psychological grading early warning model, early warning information can be obtained. This enables scientific, accurate, and early warning of psychological risks, promoting proactive intervention and resource optimization, improving the overall and humanistic level of safety management, and effectively protecting the physical and mental health of miners and production safety. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, wherein:

[0018] Figure 1 A flowchart illustrating a method for classifying and issuing early warnings of abnormal psychological states in miners, provided in an embodiment of the present invention;

[0019] Figure 2 This is a schematic diagram illustrating the effect of a long short-term memory network provided in an embodiment of the present invention;

[0020] Figure 3 This is a schematic diagram illustrating the effect of a GhostNet network provided in an embodiment of the present invention;

[0021] Figure 4This is a schematic diagram illustrating the effect of fusing multimodal features according to an embodiment of the present invention;

[0022] Figure 5 This is a schematic diagram illustrating the effect of a psychological state recognition module provided in an embodiment of the present invention;

[0023] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. The following embodiments are used to illustrate the present invention, but are not intended to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0026] It should be noted that the terms "first, second, and third" used in the embodiments of the present invention are only used to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, and third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of the present invention described herein can be implemented in an order other than that illustrated or described herein.

[0027] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which these embodiments of the invention pertain. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0028] Figure 1 This is a schematic diagram of a method for classifying and warning of abnormal psychological states of miners according to an embodiment of the present invention. The method for classifying and warning of abnormal psychological states of miners according to an embodiment of the present invention can be executed by an electronic device, such as a computer or a server.

[0029] like Figure 1As shown, the graded early warning method for abnormal psychological states of miners includes:

[0030] S101. Feature extraction and filtering are performed on the three modal signals after processing in the time domain, frequency domain, time-frequency domain, and spatial domain respectively to obtain the corresponding first target signal feature set. The first target signal feature set includes the EEG signal feature set, the peripheral physiological signal feature set, and the speech signal feature set; the three modal signals include EEG signals, peripheral physiological signals, and speech signals.

[0031] In embodiments of this invention, electroencephalogram (EEG) signals are weak electrical signals generated during the activity of clusters of neurons in the brain. Recorded via scalp electrodes, they reflect the electrical activity patterns of the cerebral cortex. Peripheral physiological signals refer to physiological activity signals generated by peripheral organs or tissues, reflecting changes in the autonomic nervous system (sympathetic / parasympathetic) and metabolic state. Speech signals are sound wave signals generated by vocal cord vibration and vocal tract modulation, containing multi-dimensional information such as language content, emotion, and identity. Peripheral physiological signals include electrocardiogram (ECG), skin conductance signals, blood oxygen saturation signals, body temperature signals, and blood pressure signals. The acquired EEG and ECG signals are stored as voltage-time series; skin conductance signals are stored as conductance-time series; blood oxygen saturation, body temperature, and blood pressure signals are stored as scalar values; and speech signals are stored as amplitude-time series. After acquiring EEG signals, peripheral physiological signals, and speech signals using different sensors, the three modalities were spatiotemporally aligned and denoised to obtain the processed three modalities. Then, feature extraction was performed on the processed three modalities in the time domain, frequency domain, time-frequency domain, and spatial domain to obtain initial EEG signal feature sets, initial peripheral physiological signal feature sets, and initial speech signal feature sets. Each initial modal signal feature set consists of feature vectors extracted from multiple domains. For example, the features in the initial EEG signal feature set consist of features extracted from the time domain, frequency domain, time-frequency domain, and spatial domain. The initial peripheral physiological signal feature sets and the initial speech signal feature sets are similar. Then, feature filtering is performed on the initial feature sets of each modality signal to obtain the final set of EEG signal features, peripheral physiological signal features, and speech signal features. The final set of EEG signal features, peripheral physiological signal features, and speech signal features are all regarded as the first target signal feature set. That is, the first target signal feature set corresponding to the EEG signal is the EEG signal feature set, the first target signal feature set corresponding to the peripheral physiological signal is the peripheral physiological signal feature set, and the first target signal feature set corresponding to the speech signal is the speech signal feature set.

[0032] Spatiotemporal alignment is a key step in ensuring the consistency of signal data collected by different sensors in time and space. During signal data acquisition, all modal signal data are marked with a unified time reference. Synchronization is achieved through timestamp sorting and interpolation. Discrete signals (blood oxygen saturation signal, body temperature signal, and blood pressure signal) are aligned to the timestamps of other continuous signals through nearest neighbor interpolation. Then, the spatiotemporally aligned three modal signals are denoised. The denoising process includes filtering, independent component analysis (ICA), and sparsity processing. Filtering removes noise from the signal that is irrelevant to the target frequency band. Notch filtering is used to remove power frequency interference, and baseline drift is filtered out through frequency band selection. The filtered signal retains the effective frequency band and has a smoother time-domain waveform. Independent component analysis (ICA) is used to process the multimodal signal separately, separating the statistically independent components in the multimodal signal. Each component represents a source signal (such as EEG artifacts, ECG artifacts, etc.), and noise reduction is achieved by component removal. Sparse processing is performed on the speech signal by sparse transformation. Taking advantage of the sparsity of the signal (most coefficients are close to zero), non-stationary noise such as plosive sounds can be effectively separated from the effective signal in the transform domain. After spatiotemporal alignment and signal denoising, the multimodal signal has a clearer physical meaning. These signals retain a unified timestamp to ensure that the time-domain statistics calculated during feature extraction are synchronized with physiological events. Finally, the three processed modal signals are obtained.

[0033] In feature extraction, the time domain, frequency domain, time-frequency domain, and spatial domain reveal the characteristics of multimodal signals from different dimensions. Time domain analysis describes the variation of multimodal signals over time; frequency domain analysis decomposes the frequency components of multimodal signals; time-frequency domain analysis jointly analyzes time-frequency characteristics; and spatial domain analysis analyzes the spatial distribution and patterns of multimodal signals. These four domains are interconnected and complementary. The final first target signal feature set integrates the key characteristics of the signal in different analytical dimensions, obtaining characteristics such as dynamic trends, periodic patterns, spectral envelope, resonant frequency, frequency change rate, and frequency band resolution.

[0034] S102. Based on the optimal transmission theory, the maximum mean difference algorithm, and the selected kernel function, the first target signal feature set is locally aligned across modes and globally aligned to obtain the second target signal feature set.

[0035] In the embodiments of the present invention, optimal transmission starts from the perspective of local detail alignment, while maximum mean difference and selection of kernel function start from the perspective of global relation alignment. The two are carried out in parallel to obtain the target signal feature set after distribution alignment, that is, the second target signal feature set. Cross-modal feature alignment eliminates the distribution differences between modes and establishes semantic consistency.

[0036] S103. Map the second target signal feature set to a unified preset dimension embedding space for comparative learning processing to obtain the third target signal feature set; and input the weighted sum of the third target signal feature set into a bidirectional long short-term memory network to obtain the initial fusion features with context awareness.

[0037] In embodiments of the present invention, the bidirectional long short-term memory network enhances the model's understanding of sequential data by combining forward and backward temporal information. Its core structure comprises two independent long short-term memory network layers: a forward processing sequence (from front to back) and a backward processing sequence (from back to front), ultimately fusing their outputs. A second target signal feature set can be projected into a shared, fixed-dimensional vector space using a learnable mapping function. Then, comparative learning is performed in this shared space to obtain a third target signal feature set. The third target signal feature set is weighted and summed according to its corresponding weights to obtain a new comprehensive feature. This new comprehensive feature is then input into the bidirectional long short-term memory network to obtain an initial fused feature with context-aware capabilities.

[0038] The process of determining the weights of the third target signal feature set is as follows: During the training phase of the Long Short-Term Memory (LSTM) network, initial weights are set for the EEG signal feature set, peripheral physiological signal feature set, and speech signal feature set, respectively. After weighted summation, the weights are input into the LTM network for training. After each N epochs of training, training is paused. During the validation phase, by resetting the weights of other signal feature sets (such as peripheral physiological signal feature set and speech signal feature set) to zero and resetting the weights of the target signal feature set (such as EEG signal feature set) to one, a single-modality-dominated fusion input is constructed and fed into the LTM network. Its classification accuracy is evaluated as the contribution. The contribution is weighted averaged and normalized to generate new fusion weights. The new fusion weights are then re-weighted and summed to perform feature fusion, and the input is fed into the LTM network to continue training for the next N epochs. After training, the final dynamically adjusted weights are obtained, which are the weights corresponding to the third target signal feature set.

[0039] like Figure 2 As shown, Figure 2 This is a schematic diagram illustrating the effect of a Long Short-Term Memory (LSTM) network provided in an embodiment of the present invention. Figure 2 middle, , and For each time step in the new integrated features input, [the feature vector is] . , and The output for each time step is shown. The new integrated features are input into the bidirectional long short-term memory network. First, the new input integrated features are preprocessed. Then, the preprocessed integrated features are simultaneously fed into the forward LSTM layer (forward long short-term memory layer) and the backward LSTM layer (backward long short-term memory layer). The forward LSTM layer captures forward context information, and the backward LSTM layer captures backward context information. At each time step, the hidden states generated by the forward and backward LSTM layers are concatenated. Finally, the concatenated feature vector is processed to obtain the initial fused features.

[0040] S104. Input the initial fused features into the GhostNet network to obtain the fused features; and input the fused features into the psychological state recognition module in the psychological grading early warning model to obtain the miner's psychological state.

[0041] In embodiments of this invention, GhostNet is a lightweight convolutional neural network that replaces redundant features generated by conventional convolution with a simple linear transformation, significantly reducing the number of parameters and computational cost while maintaining accuracy. The initially fused features are input into a regular convolutional layer in GhostNet to generate a set of basic feature maps. A linear transformation is then performed on these feature maps to generate redundant feature maps. These redundant feature maps are then concatenated with the basic feature maps to obtain a concatenated feature map. Global aggregation is performed on the concatenated feature map based on the L2 norm to calculate the relative importance of each position in the entire feature space. Using the global aggregation result, the relative importance score of the current information in the concatenated feature map is calculated. The feature normalization score is then multiplied element-wise by the concatenated feature map to obtain the fused features, which can be obtained using the following formula:

[0042] ;

[0043] In the above formula, For the i-th sub-feature map after splicing, For feature normalization scores, For the fused i-th sub-feature map, This represents the i-th sub-feature map in the basic feature map.

[0044] Furthermore, the psychological grading and early warning model includes a psychological state recognition module and an early warning grading module. The fused features are input into the psychological state recognition module in the psychological grading and early warning model to obtain the miner's psychological state, which includes four psychological states: sick, fatigued, angry, and excited.

[0045] like Figure 3 As shown, Figure 3This is a schematic diagram illustrating the effect of a GhostNet network provided in an embodiment of the present invention. Initial fused features are input into the GhostNet network, and preliminary feature extraction is performed through convolutional layers (Conv). Multiple residual blocks are then utilized. , ,..., The input features are transformed. Each residual block contains an identity mapping path that directly transmits the original information, and a path that learns new feature representations through a series of transformations (such as convolution and activation functions). The results of these two paths are added together to form the output of the residual block. After processing through multiple levels of residual blocks, the fused features are finally obtained.

[0046] like Figure 4 As shown, Figure 4 This is a schematic diagram illustrating the effect of fusing multimodal features according to an embodiment of the present invention. Figure 4 In this study, three modalities of signals—EEG, peripheral physiological signals, and speech signals—were acquired. Features were extracted and filtered from these signals in the time, frequency, time-frequency, and spatial domains, respectively, resulting in EEG, peripheral physiological, and speech signal feature sets. These feature sets were then mapped to a unified low-dimensional shared embedding space using optimal transfer theory and the maximum mean difference algorithm, respectively, for local cross-modal and global feature alignment. This enabled semantic alignment and cross-modal retrieval, ultimately yielding aligned EEG, peripheral physiological, and speech signal feature sets. These aligned feature sets were then weighted and summed according to their respective weights obtained using a weighted scoring method, resulting in a new comprehensive feature set. This new comprehensive feature set was then input into a bidirectional long short-term memory network to obtain the initial fused feature set. Finally, the initial fused feature set was input into a GhostNet network to obtain the fused feature set.

[0047] like Figure 5 As shown, Figure 5 This is a schematic diagram illustrating the effect of a psychological state recognition module provided in an embodiment of the present invention. The fused features are input into the psychological state recognition module and multiplied by a set of weights. The weighted features are then standardized through a batch normalization layer. The standardized features are input into a layer containing loop operations and a ReLU (Rectified Linear Unit) activation function. After loop operations and ReLU activation, the features are again passed through a batch normalization layer to obtain the first-processed features. The first-processed features are then multiplied by another set of weights, and then processed through the same steps as the first time to finally obtain the miner's psychological state.

[0048] S105. Input the psychological state and mine environment parameters into the early warning classification module of the psychological classification early warning model to obtain early warning information.

[0049] In an embodiment of the present invention, the mine environment parameters are the environmental parameters of the mine where the miner is currently located. The mine environment parameters include humidity, dust, noise, and illuminance. The early warning classification module sets early warning classification standards. The psychological state and mine environment parameters are input into the early warning classification module of the psychological classification early warning model. The early warning classification standards in the rule warehouse are read by the rule engine, and finally the early warning information is obtained. The early warning information includes the early warning level, the early warning color, and the early warning suggestion, such as a level two early warning, an orange warning color, and an early warning suggestion of "obvious psychological distress or abnormal behavior, it is recommended to suspend high-risk operations and arrange psychological counseling".

[0050] The first method for constructing early warning classification standards:

[0051] Specifically, the construction process of the early warning grading standard is as follows: Collect psychological health assessment data of miners under different working environments (such as noise, dust concentration, humidity, temperature, etc.). Based on preset values ​​and environmental parameters, select samples under normal environmental conditions. On these samples, calculate the median (P50), 75th percentile (P75), and 90th percentile (P90) of the psychological health assessment values. Use this to establish a baseline threshold for psychological state as the initial standard for judging normal, mildly abnormal, and severely abnormal states under ideal conditions. Obtain the weight of a single environmental parameter on the psychological state of personnel using fuzzy hierarchical analysis. Simultaneously, divide each environmental parameter into multiple intervals according to its actual range of change. Based on this, and combined with the actual parameters of the current environment, dynamically adjust the baseline threshold using factors such as weight, environmental deviation intensity (representing the degree to which the current environmental parameter value deviates from the baseline value, obtained through the current value, baseline value, and maximum allowable value), and direction of influence. Determine a reasonable final threshold. Based on the final threshold, formulate a set of early warning grading standards. The direction of influence can be represented by parameters. The symbol 'i' represents environmental factors, which can be noise, dust, and illuminance, etc., and refers to negative environmental factors (such as noise and dust). It is set to -1 for positive environmental factors (such as illuminance). It is set to 1.

[0052] The weights of individual environmental factors on psychological states were obtained using fuzzy hierarchical analysis (AHP). A fuzzy consistency matrix was constructed, and pairwise comparisons were made based on the observed influence strength of each environmental factor on the psychological state assessment value in the experimental data. Larger values ​​indicated a stronger influence of that environmental factor relative to another. Next, the weights of the fuzzy consistency matrix were calculated: the arithmetic mean of each row's elements was calculated, and the resulting rows and mean were normalized to obtain the final weights, i.e., the weights of individual environmental factors on psychological states. The fuzzy consistency matrix is ​​shown in Table 1 below, with values ​​generated randomly.

[0053] Table 1 Fuzzy Consistency Matrix

[0054] ;

[0055] It is understood that, in the embodiments of the present invention, this method proposes a multi-domain joint analysis method for miners' multimodal physiological signals based on the differences in multimodal physiological signals from the perspectives of time domain, frequency domain, time-frequency domain, and spatial domain. This method extracts the temporal characteristics, frequency distribution characteristics, time-varying frequency characteristics, and spatial distribution characteristics to comprehensively characterize the psychological stress response of workers in complex mining environments. A multi-angle cross-modal feature alignment strategy and a feature set contribution weight allocation method are adopted to establish a dynamic fusion method for multimodal features containing spatiotemporal dependencies, overcoming the limitations of single-modal signals and balancing the contributions of each modality. Based on the miners' multimodal feature fusion results, a basic network architecture including temporal analysis, convolution operations, and dense connections is designed to construct a psychological state recognition module. Environmental parameters from typical mining scenarios are introduced as prior knowledge to construct an early warning grading standard. A grading early warning model is constructed by integrating the psychological state recognition and early warning grading modules.

[0056] In some embodiments of the present invention, S101 can be implemented by S1011 to S1015, as described in the following steps.

[0057] S1011. Perform feature extraction on the three processed modal signals in the time domain, frequency domain, time-frequency domain, and spatial domain respectively to obtain the first signal feature set corresponding to each.

[0058] S1012. Calculate the first Pearson correlation coefficient between each pair of signal features within the first signal feature set, and filter out redundant signal feature pairs based on the first Pearson correlation coefficient and the first preset condition.

[0059] In some embodiments of the present invention, feature extraction is first performed on the three processed modal signals in the time domain, frequency domain, time-frequency domain and spatial domain respectively to obtain an initial first signal feature set corresponding to each. The first signal feature set contains multiple features. The first Pearson correlation coefficient between each pair of signal features in the first signal feature set is calculated. The first preset condition can be a preset value. When the first Pearson correlation coefficient satisfies the first preset condition, redundant signal feature pairs are obtained.

[0060] S1013. Based on the feature variance, obtain the signal features to be removed from the redundant signal feature pairs; and calculate the second Pearson correlation coefficient between the signal features to be removed and all features in the first signal feature set of other modal signals.

[0061] In some embodiments of the present invention, the variance of each feature in the redundant signal feature pair is calculated, the feature with larger variance is retained, and the feature with smaller variance is taken as the signal feature to be removed. Furthermore, the second Pearson correlation coefficient between the signal feature to be removed and all features in the first signal feature set of other modal signals is calculated.

[0062] For example, when the first signal feature set is the first EEG signal feature set, the signal features to be eliminated in the first EEG signal feature set are calculated, and then the second Pearson correlation coefficient between the signal features to be eliminated and all features in the first peripheral physiological signal feature set and the first speech signal feature set is calculated.

[0063] S1014. When the second Pearson correlation coefficient meets the second preset condition, the signal features to be removed are retained; otherwise, they are deleted.

[0064] S1015. Based on the retained features of the signals to be eliminated and the first set of signals features for each modal signal, obtain the second set of signals features corresponding to each of the three modal signals; and obtain the first set of signals features based on the second set of signals features.

[0065] In some embodiments of the present invention, a second preset condition is set, which can be a threshold. When the second Pearson correlation coefficient meets the second preset condition, the signal feature to be removed is retained and no further deletion is performed. When the second Pearson correlation coefficient does not meet the second preset condition, the signal feature to be removed is finally deleted. When the signal feature to be removed needs to be retained, the retained signal feature is added back to the corresponding first signal feature set, and finally, a second signal feature set corresponding to each of the three modal signals is constructed. Further, the first target signal feature set is obtained based on the second signal feature set.

[0066] In some embodiments of the present invention, obtaining the first target signal feature set based on the second signal feature set in S1015 can be implemented through S201 to S204, as described in the following steps.

[0067] S201. Calculate the third Pearson correlation coefficient of each pair of features in the second signal feature set under the current sliding window, and obtain the basic score of each signal feature according to the discrimination index of the current sliding window.

[0068] In some embodiments of the present invention, feature values ​​from a second set of signal features within the current sliding window are obtained, and then a third Pearson correlation coefficient is calculated based on the values ​​of each pair of features. A base score for each signal feature is obtained based on the discriminative index of the current sliding window. Within the current sliding window, the ability of each signal feature in the second set of signal features to distinguish different states or categories is evaluated, and this ability is quantified by a discriminative index (such as information gain and area under the curve). Then, this index value is converted into a standardized score, which serves as the base score for each signal feature within the current sliding window.

[0069] S202. Identify the first candidate signal feature selected in the previous sliding window in each signal feature, and weight the basic score corresponding to the first candidate signal feature in the basic score of the current sliding window to obtain the weighted result.

[0070] In some embodiments of the present invention, the previous candidate signal feature selected in the previous sliding window is identified in each signal feature, and then the basic score corresponding to the previous candidate signal feature is weighted in the basic score of the current sliding window to obtain a weighted result.

[0071] For example, the signal features of the current sliding window include signal feature 1, signal feature 2, ... signal feature n, each signal feature corresponds to a basic score, the previous candidate signal features include signal feature 1 and signal feature 2, the basic scores of signal feature 1 and signal feature 2 of the current sliding window are weighted to obtain the weighted result of signal feature 1 and signal feature 2, and the basic scores of other signal features remain unchanged.

[0072] S203. Based on the weighted result, the basic scores of the remaining signal features, and the third Pearson correlation coefficient, obtain the second candidate signal feature corresponding to the current sliding window; and based on the second candidate signal feature, obtain all candidate signal features corresponding to all sliding windows.

[0073] In some embodiments of the present invention, the scoring results of each signal feature under the current sliding window are obtained, the first n are selected as initial candidate signal features, and redundancy is removed from the initial candidate signal features by combining the third Pearson correlation coefficient to obtain the second candidate signal feature corresponding to the current sliding window. The candidate signal features of all sliding windows are subsequently calculated based on the second candidate signal features in the same way.

[0074] S204. Obtain the candidate frequency of each signal feature from all candidate signal features, and obtain the first target signal feature set through the signal features with high candidate frequencies.

[0075] In some embodiments of the present invention, the candidate frequency of each signal feature is calculated among the candidate signal features of all sliding windows, and the signal features with high candidate frequencies are selected to construct a first target signal feature set.

[0076] In some embodiments of the present invention, S102 can be implemented by S1021 to S1023, as described in the following steps.

[0077] S1021. Based on the EEG signal feature set, the speech signal feature set, and the optimal transmission matrix formula, the first optimal transmission matrix is ​​obtained.

[0078] S1022. Based on the set of EEG signal features, the set of peripheral physiological signal features, and the formula for the optimal transfer matrix, the second optimal transfer matrix is ​​obtained.

[0079] S1023. Obtain the second target signal feature set based on the first optimal transmission matrix, the second optimal transmission matrix, the maximum mean difference algorithm, and the selected kernel function.

[0080] In some embodiments of the present invention, the EEG signal feature set and the speech signal feature set are substituted into the optimal transfer matrix formula to obtain a first optimal transfer matrix. Then, the EEG signal feature set and the peripheral physiological signal feature set are substituted into the optimal transfer matrix formula to obtain a second optimal transfer matrix. Finally, based on the first optimal transfer matrix, the second optimal transfer matrix, the maximum mean difference algorithm, and the selected kernel function, a second target signal feature set is obtained.

[0081] In some embodiments of the present invention, S1023 can be implemented by S301 to S305, as described in the following steps.

[0082] S301. Based on the set of features of EEG signals, the set of features of speech signals, and the selected kernel function, obtain the first maximum mean difference squared distance.

[0083] In some embodiments of the present invention, the set of EEG signal features and the set of speech signal features are implicitly mapped to a high-dimensional (or even infinite-dimensional) regeneration kernel Hilbert space using a selected kernel function. In this regeneration kernel Hilbert space, they are regarded as probability distributions, and the first maximum mean difference squared distance is calculated based on the probability distribution.

[0084] S302. Based on the set of EEG signal features, the set of peripheral physiological signal features, and the selected kernel function, the second maximum mean difference squared distance is obtained.

[0085] In some embodiments of the present invention, the set of EEG signal features and the set of peripheral physiological signal features are implicitly mapped to a high-dimensional (or even infinite-dimensional) regeneration kernel Hilbert space using a selected kernel function. In this regeneration kernel Hilbert space, they are regarded as probability distributions, and the second maximum mean difference squared distance is calculated based on the probability distributions.

[0086] S303. Based on the kernel function selected from the speech signal feature set and the peripheral physiological signal feature set, the third maximum mean difference squared distance is obtained.

[0087] S304. Construct a joint alignment objective function based on the first optimal transfer matrix, the second optimal transfer matrix, the first maximum mean difference squared distance, the second maximum mean difference squared distance, and the third maximum mean difference squared distance.

[0088] S305. Optimize the EEG signal feature set, speech signal feature set, and peripheral physiological signal feature set according to the joint alignment objective function to obtain the second target signal feature set.

[0089] In some embodiments of the present invention, using a selected kernel function, the speech signal feature set and the peripheral physiological signal feature set are implicitly mapped to a high-dimensional (or even infinite-dimensional) regeneration kernel Hilbert space, where they are treated as probability distributions, and a third maximum mean difference squared distance is calculated based on the probability distributions. Then, a joint alignment function is constructed based on the first optimal transfer matrix, the second optimal transfer matrix, the first maximum mean difference squared distance, the second maximum mean difference squared distance, and the third maximum mean difference squared distance. Finally, the joint alignment objective function is used to optimize the EEG signal feature set, the speech signal feature set, and the peripheral physiological signal feature set to obtain a second target signal feature set.

[0090] In some embodiments of the present invention, the process of mapping the second target signal feature set to a unified preset dimension embedding space for comparative learning in S103 to obtain the third target signal feature set can be implemented through S1031 to S1032, as described in the following steps.

[0091] S1031. By learning the mapping function, the second target signal feature set is mapped to the shared embedding space to obtain the corresponding embedded target signal feature set.

[0092] In some embodiments of the present invention, the second target signal feature set includes a distribution-aligned set of EEG signal features, a distribution-aligned set of speech signal features, and a distribution-aligned set of peripheral physiological signal features. Different learning mapping functions are used to map the corresponding distribution-aligned sets of EEG signal features, distribution-aligned sets of speech signal features, and distribution-aligned sets of peripheral physiological signal features to obtain a shared embedding space containing the sets of EEG signal features, sets of speech signal features, and sets of peripheral physiological signal features.

[0093] S1032. Construct a contrastive loss function using the corresponding embedded target signal feature set, and optimize the learning mapping function through the contrastive loss function until a third target signal feature set is obtained.

[0094] In some embodiments of the present invention, a contrastive loss function is constructed using a set of EEG signal features, a set of speech signal features, and a set of peripheral physiological signal features in a shared embedding space, and the learning mapping function is optimized through the contrastive loss function until a third target signal feature set is obtained.

[0095] Reference Figure 6 The diagram shows a structural schematic of an electronic device according to an embodiment of the present invention. The specific embodiments of the present invention do not limit the specific implementation of the electronic device.

[0096] like Figure 6 As shown, the electronic device may include: a processor 502, a communications interface 504, a memory 506, and a communications bus 508.

[0097] in:

[0098] The processor 502, communication interface 504, and memory 506 communicate with each other via communication bus 508.

[0099] Communication interface 504 is used to communicate with other electronic devices or servers.

[0100] The processor 502 is used to execute program 510, specifically the relevant steps in the above method embodiments.

[0101] Specifically, program 510 may include program code that includes computer operation instructions.

[0102] Processor 502 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. The smart device may include one or more processors of the same type, such as one or more CPUs; or it may include processors of different types, such as one or more CPUs and one or more ASICs.

[0103] Memory 506 is used to store program 510. Memory 506 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0104] Specifically, program 510 can be used to cause processor 502 to perform the operations corresponding to the methods described in the above method embodiments.

[0105] The specific implementation of each step in program 510 can be found in the corresponding descriptions of the steps and units in the above method embodiments, and will not be repeated here. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the devices and modules described above can be referred to the corresponding process descriptions in the foregoing method embodiments, and will not be repeated here.

[0106] It should be noted that, depending on the implementation needs, the various components / steps described in the embodiments of the present invention can be broken down into more components / steps, or two or more components / steps or parts of the operation of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present invention.

[0107] The methods described above according to embodiments of the present invention can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code originally stored on a remote recording medium or a non-transitory machine-readable medium and subsequently stored on a local recording medium, downloaded via a network. Thus, the methods described herein can be processed by software stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., RAM, ROM, flash memory, etc.) capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses code used to implement the methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for executing the methods shown herein.

[0108] Those skilled in the art will recognize that the units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of the embodiments of the present invention.

[0109] The above embodiments are only used to illustrate the embodiments of the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of the present invention, and the patent protection scope of the embodiments of the present invention should be defined by the claims.

Claims

1. A method for graded early warning of abnormal psychological states in miners, characterized in that, include: Feature extraction and filtering are performed on the three modal signals after processing in the time domain, frequency domain, time-frequency domain, and spatial domain respectively to obtain the corresponding first target signal feature set. The first target signal feature set includes the EEG signal feature set, the peripheral physiological signal feature set, and the speech signal feature set; the three modal signals include EEG signals, peripheral physiological signals, and speech signals. The first target signal feature set is locally aligned across modes and globally based on optimal transmission theory, maximum mean difference algorithm and selected kernel function, respectively, to obtain the second target signal feature set; The second target signal feature set is mapped to a unified embedding space of a preset dimension for comparative learning processing to obtain the third target signal feature set. The weighted sum of the third target signal feature set is then input into a bidirectional long short-term memory network to obtain initial fusion features with context awareness. The initial fused features are input into the GhostNet network to obtain the fused features; and the fused features are input into the psychological state recognition module in the psychological grading early warning model to obtain the miner's psychological state. The psychological state and mine environment parameters are input into the early warning classification module of the psychological classification early warning model to obtain early warning information; The second target signal feature set is obtained by performing local cross-modal alignment and global feature alignment on the first target signal feature set based on optimal transmission theory, maximum mean difference algorithm, and selected kernel function, respectively, including: Based on the set of EEG signal features, the set of speech signal features, and the formula for the optimal transmission matrix, the first optimal transmission matrix is ​​obtained; Based on the set of EEG signal features, the set of peripheral physiological signal features, and the formula for the optimal transfer matrix, the second optimal transfer matrix is ​​obtained; The second target signal feature set is obtained based on the first optimal transmission matrix, the second optimal transmission matrix, the maximum mean difference algorithm, and the selected kernel function; The step of obtaining the second target signal feature set based on the first optimal transmission matrix, the second optimal transmission matrix, the maximum mean difference algorithm, and the selected kernel function includes: Based on the set of EEG signal features, the set of speech signal features, and the selected kernel function, the first maximum mean difference squared distance is obtained; Based on the set of EEG signal features, the set of peripheral physiological signal features, and the selected kernel function, the second maximum mean difference squared distance is obtained; Based on the selected kernel function of the speech signal feature set and the peripheral physiological signal feature set, the third maximum mean difference squared distance is obtained; A joint alignment objective function is constructed based on the first optimal transfer matrix, the second optimal transfer matrix, the first maximum mean squared difference distance, the second maximum mean squared difference distance, and the third maximum mean squared difference distance. The second target signal feature set is obtained by optimizing the EEG signal feature set, the speech signal feature set, and the peripheral physiological signal feature set according to the joint alignment objective function.

2. The method according to claim 1, characterized in that, The feature extraction and filtering of the processed three modal signals in the time domain, frequency domain, time-frequency domain, and spatial domain respectively are performed to obtain the corresponding first target signal feature sets, including: Feature extraction was performed on the three processed modal signals in the time domain, frequency domain, time-frequency domain, and spatial domain respectively to obtain the corresponding first signal feature set; Within the first set of signal features, the first Pearson correlation coefficient between each pair of signal features is calculated, and redundant signal feature pairs are selected based on the first Pearson correlation coefficient and the first preset condition. Based on the feature variance, the signal features to be eliminated are obtained from redundant signal feature pairs; and the second Pearson correlation coefficient between the signal features to be eliminated and all features in the first signal feature set of other modal signals is calculated; When the second Pearson correlation coefficient meets the second preset condition, the signal feature to be removed is retained; otherwise, it is deleted. Based on the retained features of the signals to be eliminated and the first set of signals features for each modality, the second set of signals features corresponding to each of the three modal signals is obtained; and the first set of signals features for the target signals is obtained based on the second set of signals features.

3. The method according to claim 2, characterized in that, The step of obtaining the first target signal feature set based on the second signal feature set includes: Calculate the third Pearson correlation coefficient of each pair of features in the second signal feature set under the current sliding window, and obtain the basic score of each signal feature according to the discrimination index of the current sliding window; In each signal feature, the first candidate signal feature selected in the previous sliding window is identified, and the basic score corresponding to the first candidate signal feature is weighted in the basic score of the current sliding window to obtain the weighted result; Based on the weighted result, the baseline scores of the remaining signal features, and the third Pearson correlation coefficient, the second candidate signal feature corresponding to the current sliding window is obtained; and based on the second candidate signal feature, all candidate signal features corresponding to all sliding windows are obtained. The candidate frequency of each signal feature is obtained from all candidate signal features, and the first target signal feature set is obtained through the signal features with high candidate frequencies.

4. The method according to claim 1, characterized in that, The step of mapping the second target signal feature set to a unified, preset-dimensional embedding space for comparative learning to obtain a third target signal feature set includes: By learning the mapping function, the second target signal feature set is mapped to the shared embedding space to obtain the corresponding embedded target signal feature set; A contrastive loss function is constructed using the corresponding embedded target signal feature set, and the learning mapping function is optimized using the contrastive loss function until a third target signal feature set is obtained.

Citation Information

Patent Citations

  • Miner emotion recognition method based on expression, electroencephalogram and voice multi-modal fusion

    CN117195148A

  • Teenager mental health data analysis and early warning system

    CN117912710A