Multi-modal heterogeneous data real-time fusion and intelligent decision-making method and system based on cloud computing

Through the real-time integration of multimodal data and intelligent decision-making methods on the cloud computing platform, the problems of multimodal data heterogeneity and emotional state recognition in the industrial Internet are solved, efficient state recognition and decision-making are achieved, and production safety and efficiency are improved.

CN120597193APending Publication Date: 2025-09-05LIAONING UNIVERSITY
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510659270.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The existing technology has problems of multimodal data heterogeneity, timing misalignment and missing in the industrial Internet, resulting in poor data fusion quality and lack of active identification and fusion capabilities of operators' emotional state, affecting production safety and efficiency.

Method used

The real-time fusion and intelligent decision-making method of multimodal heterogeneous data based on cloud computing are adopted. Through data acquisition, preprocessing, multimodal feature extraction and fusion representation, state modeling and graph structure representation, state classification and decision-making fusion, and real-time visual feedback, edge computing and cloud platform are used for data processing, and self-supervised training and decision-making are carried out in combination with graph neural networks.

Benefits of technology

It realizes high robustness and timeliness state recognition and intelligent decision-making, improves production safety and efficiency, describes the correlation between multimodal features through covariance matrix, enhances the generalization ability of the model, and forms an intelligent decision-making system with human-computer collaboration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120597193A_ABST
    Figure CN120597193A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal heterogeneous data real-time fusion and intelligent decision-making method and system based on cloud computing, and is suitable for industrial internet scenes. According to the method, multi-modal data such as physical sensors, audio and video, physiological signals and the like are collected, unified preprocessing and feature extraction are carried out, and a covariance matrix is constructed to represent a coupling relation between modals; further constructing the multi-period state into a graph structure, and realizing high-robustness state modeling by using a graph neural network and self-supervised learning; and identifying the operation state through a Riemannian geometric classification method, and carrying out risk scoring and grade judgment in combination with an emotional state and an environment index. The system supports AR visual prompt and control linkage, and the man-machine cooperation intelligent decision-making ability in an industrial scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a cloud computing-based multimodal heterogeneous data real-time fusion and intelligent decision-making method and system. Background Art

[0002] In scenarios such as the Industrial Internet, intelligent manufacturing, and smart operations and maintenance, production equipment and operators are in a dynamic state. Environmental factors, mechanical vibrations, and human emotions intersect to influence the stability and safety of production systems. Especially in critical process steps, abnormal equipment operation, operator fatigue, or inattention often contribute to failures and accidents. Therefore, achieving comprehensive perception and intelligent decision-making of the integrated "man-machine-environment" operational status while ensuring production efficiency has become a core challenge in the current evolution of industrial intelligence.

[0003] To address these issues, some technical solutions have been proposed that use multimodal data fusion and artificial intelligence methods to monitor and predict equipment and personnel status. Typical approaches include using deep neural networks to identify sensor data, combining video analysis to recognize operator expressions and behaviors, and leveraging edge computing nodes to process multi-source data in real time. These technologies achieve a certain degree of perceptual fusion of heterogeneous multi-source data and provide preliminary intelligent diagnostic capabilities.

[0004] However, existing technologies still have the following shortcomings: On the one hand, the heterogeneity, time sequence misalignment, and missing issues between multimodal data seriously restrict the quality of data fusion, which can easily lead to misjudgments or omissions. On the other hand, most existing methods are based on Euclidean space modeling, which is insufficient for modeling nonlinear coupling relationships between multiple channels, resulting in poor stability of classification results in complex scenarios. In addition, current systems often lack the ability to actively identify and integrate the emotional state of operators, and cannot form a truly human-machine collaborative intelligent decision-making system. Therefore, a new multimodal fusion method with structural modeling capabilities, emotional perception capabilities, and visual feedback capabilities is urgently needed to improve safety and decision-making efficiency in industrial scenarios. Summary of the Invention

[0005] In response to the requirements raised in the above background technology, embodiments of the present invention provide a cloud computing-based multimodal heterogeneous data real-time fusion and intelligent decision-making method and system, which are applicable to industrial Internet scenarios, especially in situations where the operating status of equipment and the behavior of personnel jointly affect production safety and efficiency. The method of the present invention specifically includes the following steps:

[0006] Step 1: Data collection and preprocessing

[0007] This system is deployed at industrial sites to collect multimodal heterogeneous data, including: physical sensor signals such as temperature, pressure, and vibration from industrial equipment; video images and voice audio from operators to identify their expressions and intonation; and physiological signals from wearable devices (such as wristbands), such as heart rate and skin charge.

[0008] The above data is uploaded to the cloud platform in real time through the edge gateway. Before uploading, the edge device completes the following pre-processing operations:

[0009] Signal filtering: Perform median filtering and bandpass filtering on each signal to remove high-frequency noise and pulse interference;

[0010] Data normalization: unify the dimensions and numerical scales of different modal data to facilitate subsequent unified modeling;

[0011] Time alignment: Based on the NTP time synchronization protocol and linear interpolation method, all modal data are mapped to a unified time axis to ensure cross-modal sample coordination and consistency.

[0012] Step 2: Multimodal feature extraction and fusion representation

[0013] After completing data preprocessing, the system performs high-level feature extraction on each modality, including:

[0014] Visual modality: CNN network is used to extract facial expression features from video frames;

[0015] Speech modality: Use LSTM or CNN-LSTM networks to extract emotional features such as intonation and intensity from speech segments;

[0016] Physiological / sensor modality: Extract statistical features (such as mean, standard deviation, spectral energy, etc.) of signals such as heart rate and skin conduction through a sliding window.

[0017] The above three types of features are used to obtain emotion or state vectors through the trained models, and are fused into a unified emotional state representation according to weighted rules to describe the current operator's psychological state and physical load level.

[0018] Step 3: State Modeling and Graph Structure Representation

[0019] The data in each time window is regarded as an independent state unit, and a statistical covariance matrix is ​​constructed to capture the joint change pattern between different modes. The covariance matrix has the following characteristics:

[0020] Characterizes the synchronous fluctuation trend between signals;

[0021] It has natural smoothness against noise and is suitable for stable learning under few-sample conditions.

[0022] To capture temporal and spatial correlations, the system further organizes multiple covariance units into a graph structure, which is defined as follows:

[0023] Each time window corresponds to a graph node, and its node feature is the vector after the corresponding covariance matrix is ​​expanded;

[0024] If two windows are located in a similar time period or come from the same operating equipment or personnel, an edge connection is established in the graph, and the weight is set according to the time or physical proximity.

[0025] The structure is trained using graph neural networks (such as GCN), and the implicit expression ability of each node is learned through multi-layer propagation and nonlinear mapping, further improving the ability to recognize different states.

[0026] To reduce reliance on manual labels, the system adopts a self-supervised training strategy:

[0027] Mask features of some nodes and train the model to restore the masked features based on their neighbor features.

[0028] A contrastive learning method is used to strengthen the consistency of embeddings between similar nodes in the graph and improve the generalization ability of the model.

[0029] Step 4: State Classification and Decision Fusion

[0030] After embedding the graph structure, the system maps the data of each time window into a fixed-dimensional representation vector. During the training phase, the system calculates the "average covariance structure" of each typical state sample (such as normal, abnormal, overheated, fatigue, etc.) as the central template for that category. In actual operation, the system compares the covariance structure of new data with the templates of each category:

[0031] If its distance is the smallest than a certain category, it is considered to belong to that state;

[0032] The size of the distance reflects the degree of similarity, which is then converted into classification confidence.

[0033] In addition, the system also takes into account:

[0034] The operator's current emotional state (e.g., fatigue, tension, etc.);

[0035] Whether the environmental variables exceed the limit (such as high temperature, severe vibration, etc.).

[0036] Finally, the system outputs a total risk score based on the weighted classification confidence, emotional risk score, and environmental indicators, and sets dual thresholds:

[0037] Below threshold 1 is normal;

[0038] Between threshold 1 and threshold 2 is a warning;

[0039] If the value is higher than threshold 2, an alarm will be triggered and the downstream safety module will be linked.

[0040] Step 5: Real-time visual feedback and linkage control

[0041] The system transmits decision results back to the local terminal through a cloud-edge-end architecture, including:

[0042] On wearable AR devices or large screens, the device and personnel status are indicated in the form of three-dimensional overlay icons;

[0043] If it is in the early warning or alarm state, the corresponding graphic (such as yellow / red warning sign) will be displayed, along with a text prompt (such as "the operator is nervous" or "the equipment is vibrating abnormally").

[0044] The visualization module is linked to the control module. When the risk score exceeds the threshold, it automatically activates safety protection, issues an audible and visual alarm, or reminds personnel to suspend operations.

[0045] Further: A cloud computing-based multimodal heterogeneous data real-time fusion and intelligent decision-making system, comprising:

[0046] The present invention also provides a cloud computing-based multimodal heterogeneous data real-time fusion and intelligent decision-making system. The system is suitable for industrial Internet scenarios, especially for applications such as collaborative monitoring and intelligent early warning of multi-source heterogeneous signals. The modules of the system are interconnected through a network to form a complete data processing and decision-making system. Specifically, the system includes the following modules:

[0047] The data acquisition module, used to collect data from various modalities at industrial sites, consists of a physical sensor submodule, an audio and video acquisition submodule, and a physiological signal acquisition submodule. The physical sensor submodule, deployed on the equipment itself or in its surroundings, collects operating condition information such as temperature, pressure, vibration, and current. The audio and video acquisition submodule uses a camera and microphone to capture operator behavioral characteristics such as facial expressions and intonation in real time. The physiological signal acquisition submodule uses wearable devices to capture physiological indicators such as heart rate and skin charge, reflecting the operator's workload and emotional state.

[0048] The data acquisition module sends the collected data to the data processing module after local buffering and time stamping.

[0049] The data processing and preprocessing module is used to unify the format, clean, synchronize and normalize multimodal data. It is specifically composed of a filtering and denoising submodule, a numerical normalization submodule, and a time synchronization and resampling submodule. The filtering and denoising submodule uses median filtering and bandpass filtering to remove pulse interference and high-frequency noise. The numerical normalization submodule normalizes each signal to zero mean and unit variance by channel to solve the problem of dimensional inconsistency. The time synchronization and resampling submodule is based on global clock calibration and aligns modal signals of different frequencies to a unified time base through linear interpolation.

[0050] The data processing and preprocessing module sends the processed data to the subsequent modeling module.

[0051] The feature modeling and state encoding module is used to extract high-level state representations from preprocessed data. It specifically includes a modal feature extraction submodule, a feature fusion submodule, and a covariance modeling submodule. The modal feature extraction submodule uses trained deep network models to extract facial expression features, audio emotion features, and physiological fluctuation features from video, voice, and physiological data respectively. The feature fusion submodule fuses the above-mentioned multi-source modal feature vectors into a unified emotion state vector to represent the operator's current psychological and physiological state. The covariance modeling submodule calculates the covariance matrix of the modal joint features based on a sliding time window to construct a statistical representation of the multimodal state.

[0052] The graph modeling and graph neural network embedding module is used to construct multiple state windows into a graph structure and extract feature embedding through the graph neural network. The module consists of a graph construction submodule, a graph convolution submodule, and a self-supervised training submodule; the graph construction submodule uses each time window as a graph node and sets edge weights based on time proximity, spatial position, or device function association; the graph convolution submodule uses a graph convolutional neural network (GCN) to propagate and aggregate node features to generate context-aware node embedding representations; the self-supervised training submodule uses masked reconstruction and graph comparison learning methods to achieve network pre-training in an unlabeled state, thereby improving the model's generalization ability.

[0053] The classification discrimination and risk fusion module is used to classify the current state and fuse multi-source risk information to output the risk level. The module includes a manifold classification submodule, a confidence calculation submodule, a decision fusion submodule, and a level judgment submodule; the manifold classification submodule compares the covariance characteristics of each window with the Riemannian geometric mean of each category obtained by offline calculation, and classifies based on the geodesic distance; the confidence calculation submodule outputs the confidence of each category match according to the distance result; the decision fusion submodule fuses the classification confidence, operator emotional risk, and environmental indicator score, and generates a total risk score through weighted calculation; the level judgment submodule sets a dual threshold, classifies the score and outputs "normal", "warning" or "alarm" status.

[0054] Visual presentation and linkage control module, which is used to visualize the judgment results and drive the linkage equipment or prompt terminal; this module includes AR visualization sub-module, linkage control sub-module, and log recording and tracing sub-module; the AR visualization sub-module superimposes the current risk status on the on-site equipment, operator or work area location, and prompts status changes through icons and text; the linkage control sub-module triggers linkage control according to the "alarm" status, including sound and light alarms, equipment shutdown instructions, operation authority restrictions and other actions; logging and tracing sub-module: uploads all judgment results, feature data, operation feedback, etc. to the cloud platform to support historical tracking and model retraining.

[0055] Beneficial effects of the present invention: The method of the present invention fully integrates the multimodal information of "equipment status - personnel status - environmental status" in industrial production sites, and realizes highly robust and timely state recognition and intelligent decision-making through covariance statistical modeling, graph neural network learning, Riemannian geometry classification and AR visualization feedback.

[0056] The covariance matrix is ​​introduced to describe the correlation between multimodal features, and the geodesic distance on the SPD manifold is used to achieve more accurate state classification;

[0057] Graph neural networks are used to depict complex relationships between multiple time periods and channels, and pre-training with unlabeled data significantly enhances generalization capabilities.

[0058] Taking the operator's emotional characteristics as one of the decision-making factors, an integrated state closed loop from "human" to "machine" is formed;

[0059] Risk warnings can be quickly rendered through the edge-cloud-end architecture, and the control system can be automatically linked in the alarm state, with real-time intervention capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0061] Figure 1 A flow chart of the method of the present invention is shown.

[0062] Figure 2 A schematic diagram of the composition of the system of the present invention is shown. DETAILED DESCRIPTION

[0063] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. It should be understood that the drawings in the present invention only serve the purpose of illustration and description and are not used to limit the scope of protection of the present invention. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in the present invention illustrate the operations implemented according to some embodiments of the present invention. It should be understood that the operations of the flowchart can be implemented out of sequence, and steps that have no logical context relationship can be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of the present invention, can add one or more other operations to the flowchart, or remove one or more operations from the flowchart.

[0064] In addition, the embodiments described in the present invention are only some of the embodiments of the present invention, rather than all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.

[0065] It should be noted that the term "including" will be used in the embodiments of the present invention to indicate the presence of the features declared thereafter, but does not exclude the addition of other features. It should also be noted that similar numbers and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In the description of the present invention, it should also be noted that the terms "first", "second", "third", etc. are only used to distinguish the description and are not to be understood as indicating or implying relative importance.

[0066] The specific details of this solution are described in detail below with reference to the embodiments.

[0067] First, this solution addresses several problems existing in the industrial Internet, including the problem that equipment operation status monitoring only focuses on quantitative physical indicators but ignores human factors, the lack of unified modeling of multi-source heterogeneous information and insufficient generalization ability of small samples, and the problem that multi-channel statistical correlation presents a nonlinear manifold structure that is difficult to accurately classify using Euclidean methods. It introduces multimodal emotion recognition based on video, voice, and physiological signals, and models the operator's psychological state and physical sensor data at the same level; drawing on the heterogeneous CSI graph pre-training method, multimodal nodes are organized into a graph structure and pre-trained on massive unlabeled data through self-supervised comparison and reconstruction tasks, significantly improving the feature robustness and generalization under small samples; following the Riemannian geometry classification on the SPD manifold, the fused features of each time period are input in the form of a covariance matrix, and the geometric mean and geodesic distance are used for hyper-parameter-free, high-precision real-time classification, and the confidence level is output; combined with the cloud-edge-end architecture, the discrimination results and emotional information are presented in real time through the AR visualization interface to form a closed-loop early warning and linkage control.

[0068] The following is a detailed description of the steps of the method in conjunction with the relevant drawings of the specification. Figure 1 The method in this case is a real-time fusion and intelligent decision-making method of multimodal heterogeneous data based on cloud computing, which includes the following steps:

[0069] Data preprocessing and timing synchronization

[0070] On the cloud computing platform, heterogeneous data needs to be formatted in a unified manner and time-series synchronized to lay the foundation for subsequent fusion. Therefore, in this step, data from different devices and networks (such as sensors, cameras, and voice streams) are uploaded to the cloud through the network interface for preprocessing.

[0071] The preprocessing process includes cleaning, normalizing and aligning the data of each modality, specifically including:

[0072] 1.1. Data cleaning: remove missing values ​​and erroneous data and perform interpolation to fill in the gaps;

[0073] Data cleaning mainly includes: denoising and outlier removal, digital filtering;

[0074] 1.11. Regarding denoising and outlier removal, since industrial field environments are often accompanied by noise sources such as electromagnetic interference and mechanical vibration, the original signal is often mixed with burst pulses or drift trends, which should be removed or smoothed. This is achieved by combining median filtering and sliding mean filtering on each time series signal x[n] (such as acceleration, temperature, pressure, etc.).

[0075] Median filtering is used to remove single-point impulse noise, and its process is shown in formula (1):

[0076] x med[n]=median{x[nk],...,x[n],...,x[n+k]}(1)

[0077] In formula (1), x[n] represents the value of the original discrete time series signal at the nth sampling point (e.g., the sensor measurement value); k represents the half-width of the median filter window (usually 1 to 3), indicating that k neighboring points on each side are selected for sorting; median{·} represents the median value of all values ​​in a given window, which is used to eliminate single-point impulse noise;

[0078] Sliding mean filtering is used to smooth random jitter, and its process is shown in formula (2):

[0079]

[0080] In formula (2), represents the value of the signal after median filtering at the n+ith position; M represents the half-width of the sliding window, which determines the number of sampling points involved in the mean calculation; the window length is 2M+1 (for example, M=5).

[0081] The median filter is used to remove outliers and mutations, and then the sliding mean filter is used to eliminate high-frequency jitter to ensure that the remaining signal is a stable representation of the real physical quantity.

[0082] 1.12. As for digital filtering, since the sensor signal may contain power frequency (50 / 60 Hz) or high-frequency interference, it is necessary to suppress it in a targeted manner. Digital filtering is performed by using a second-order or fourth-order Butterworth IIR bandpass (or bandstop) filter. The differential equation is shown in Equation (3):

[0083]

[0084] In formula (3), y[n] represents the value of the output signal at point n after filtering; x smooth [ni] represents the smoothed signal input to the filter, delayed by i sampling points; b i Represents the filter numerator coefficient (zero-point coefficient), which determines the contribution weight of current and past inputs to the output; a j Represents the filter denominator coefficient (pole coefficient), which determines the feedback weight of the past output on the current output.

[0085] Where P, Q are the numerator and denominator orders of the filter, respectively, and the coefficient b i , a jThe passband / stopband cutoff frequencies are pre-calculated based on the design. For a fourth-order bandpass filter, for example, the analog prototype (Butterworth prototype polynomial) can be mapped to discrete IIR coefficients using a bilinear transformation [formula not expanded]. After bandpass filtering, the desired frequency range (e.g., 0.5-50 Hz) is retained, power frequency and high-frequency interference are suppressed, and the signal-to-noise ratio (SNR) for subsequent feature extraction is improved.

[0086] 1.2. Feature normalization: Normalize each feature channel, such as X′=(X-μ) / σ;

[0087] Since the dimensions and numerical ranges of different sensor channels may vary greatly, directly inputting them into the fusion or learning model will cause some channels to "dominate" the results and suppress information from other channels. Therefore, each feature channel needs to be normalized.

[0088] The specific normalization process is to perform normalization on each signal or feature vector x, which includes calculating the mean and standard deviation and normalization transformation;

[0089] 1.21. The process of calculating the mean and standard deviation is shown in formula (4) and formula (5):

[0090]

[0091] In equations (4) and (5), N represents the total number of sampling points in a batch or a time window; x[n] represents the nth value of the original or filtered signal; μ represents the mean value of the channel in the current window, which is used for offset correction; σ represents the standard deviation of the channel in the current window, which is used for dimensional scaling.

[0092] 1.22, the normalization transformation process is shown in formula (6):

[0093]

[0094] In formula (6), x norm [n] represents the normalized signal value with zero mean and unit variance;

[0095] After normalization transformation, the signal has zero mean and unit variance, and the features of different channels are placed on the same numerical scale to avoid gradient bias in model training.

[0096] 1.3. Time synchronization: Based on the timestamp or time calibration scheme, queue buffering and time interpolation technology are used to synchronize multiple source signals and generate a unified time series X(t k ). At the same time, each channel data is sent to subsequent modules in real time through cloud-based stream computing.

[0097] Time synchronization is also the process of cross-modal timing alignment, which includes unified time baseline, linear interpolation alignment and multi-modal queue synchronization.

[0098] 1.31. Regarding the unified time baseline, since video, audio, and sensor devices sample independently and have different sampling rates and clock drifts, time reference synchronization is required to simultaneously use multiple sources of information at the same time. Since the data packets collected by each device contain local timestamps (sampling moments), the system performs two steps: alignment to the global clock and resampling in the cloud or local gateway to achieve a unified time baseline.

[0099] The alignment to the global clock is to calibrate the time of all devices with the master clock (such as NTP server) at startup, synchronize regularly, and record the time deviation Δt offset .

[0100] The resampling is to map the data of each device to a unified time point set {t k}.

[0101] 1.32. For linear interpolation alignment, if a certain signal x norm (t) Its original sampling time is {t i}, hoping to estimate the signal value at the same time {t k}, linear interpolation can be used, and the calculation process is shown in formula (7):

[0102]

[0103] In formula (7), t i ,t i+1 Represents the original sampling time, the closest to the target alignment time t k Two adjacent moments; t k represents the target alignment moment; x norm (t i ), x norm (t i+1 ) represent the time t i and t i+1 The normalized signal value at x align (t k ) represents the time t obtained after interpolation k signal value.

[0104] where t i ≤t k ≤t i+1 .

[0105] Linear interpolation is simple and efficient, and can smoothly map non-uniform time series data to a uniform moment, making it suitable for most real-time monitoring scenarios.

[0106] 1.33. For multimodal queue synchronization, a time sliding window queue is maintained in the cloud. Every time a new alignment sample x is generated in any channel, align (t k ), check other channels at t k Check whether there is a corresponding sample. If missing, interpolation is triggered or marked as "temporarily missing". After all channels are completed, the multi-channel alignment samples are combined into a multi-modal data packet. And push it to the downstream fusion module. This process ensures that at each alignment moment, there is complete multi-source data input into the subsequent algorithm to avoid interruption of the overall analysis due to temporary packet loss in a certain channel.

[0107] Through the above preprocessing process, the system can convert multimodal raw data of different industry standards into a unified, noise-free, numerically consistent and time-synchronized format, providing reliable and high-quality data input for subsequent heterogeneous graph modeling, self-supervised pre-training and Riemannian manifold classification.

[0108] Multimodal data collection and emotion recognition

[0109] The purpose of this step is to fuse the pre-processed multimodal signals such as speech and vision and use them to enrich the subtitles with emotion.

[0110] Specifically, the system treats inputs such as device sensor data, video, and voice as multimodal information sources and simultaneously identifies the underlying "emotions" or states, including emotional indicators such as the operator's facial expressions, intonation, and voice. This recognition process utilizes deep convolutional networks (CNNs) and long short-term memory networks (LSTMs) to extract emotional features from video and voice, respectively, and then labels the emotional states through multimodal fusion.

[0111] The above identification process specifically includes the following details:

[0112] 2.1. First, apply the emotion recognition model to the real-time collected audio and video data to obtain the emotion label sequence;

[0113] 2.2. Then calculate the feature vector of the device sensor data and combine it with the emotional feature to form a comprehensive feature The fusion method can use the attention mechanism (Attention) or the differentiable weighted method, for example: F multi =αF sensor +βF face +γF voice , where weights α, β, and γ are obtained by network training.

[0114] For example, in an industrial scenario, when mechanical operation is detected, the system collects vibration sensor data from the equipment and the operator's voice and facial video. The system analyzes the voice tone and expression, identifies emotional markers of "tension" or "normal" states, and then uses this emotional information together with the vibration signal as the basis for subsequent decision-making (such as alarm or continued monitoring).

[0115] As can be seen from the preceding description, multimodal signals include visual modalities, speech modalities, and physical sensor modalities. We will provide a detailed explanation of the acquisition and preprocessing process of these different modal signals.

[0116] Visual modality: This modality uses video / image signals. Cameras deployed at industrial sites capture video images of workers and the environment, then extract facial or emotion-related visual features frame by frame from the captured video stream. Face detection and alignment are performed first, followed by facial expression features extracted using a deep learning model (such as a convolutional neural network (CNN). CNN models can automatically learn to represent emotions such as anger, happiness, and sadness during image analysis. The model outputs an emotion category or corresponding probability score for the input facial image.

[0117] Speech modality: In the work scene, the speech of the operator or the environmental sound is obtained through audio collection devices such as microphones, and the emotional features are extracted after pre-processing the audio signal. Speech emotion features usually include acoustic parameters such as intonation, fundamental frequency (pitch), intensity, and duration (such as MFCC, chromatogram, etc.). The extracted audio features are input into a recurrent neural network (such as LSTM) or a CNN-LSTM hybrid network in the form of a time series. CNN is used to extract short-term spectral features, and LSTM is responsible for capturing long-term dependent emotional dynamics. The trained speech emotion model can map the input audio features to the output score of the corresponding emotion category. In addition, semantic sentiment analysis can be performed after conversion into text through the speech recognition module to enhance emotional judgment, but the core still relies on the implicit expression of emotions by acoustic features.

[0118] Physical sensor modality: The Industrial Internet of Things (IIoT) can also use wearable devices or environmental sensors to collect human physiological signals, such as electrocardiogram (ECG), heart rate, and electrodermal response (GSR). The collected raw signals are typically multi-channel, continuous time series data and must first be formatted: the sampling frequency is unified, the data is segmented, and noise and power frequency interference are removed. Time domain features (such as mean, variance, and peak interval) and frequency domain features (such as energy spectrum and waveform characteristics) are then calculated for each data segment. Nonlinear features (such as entropy) are also extracted when necessary. For example, a wearable bracelet emotion system performs segmented denoising on ECG, heart rate, and GSR signals before extracting linear, nonlinear, time domain, and frequency domain features. These formatted sensor feature vectors can be input into specialized classifiers (such as 1DCNN, MLP, or SVM) to output corresponding emotional tendencies or scores. In industrial scenarios, environmental monitoring sensor data (such as temperature, vibration, and pressure) can also be combined as auxiliary inputs, but core emotion determination typically focuses on physiological signals to characterize human status.

[0119] Regarding the structure and principle of the emotion recognition model, after the features of each modality are extracted, a dedicated neural network is used to classify the emotions:

[0120] For visual modalities, deep CNN architectures are usually used to process facial images, such as ResNet and Inception-ResNet, which continuously abstract and extract expression features from the convolutional layer and finally output the emotion prediction results through the fully connected layer.

[0121] For speech, the audio signal is typically converted into short speech frames (such as Mel-spectrograms or MFCCs) before being fed into an LSTM or CNN-LSTM hybrid network. The CNN extracts local features from the spectral frames, while the LSTM layer captures temporal correlations within the speech sequence, thereby learning patterns of emotional variation. Training is supervised by manually annotated emotional data, and network parameters are optimized using loss functions such as cross-entropy.

[0122] For physical sensor modalities, one-dimensional CNN or traditional machine learning models can be used to classify the pre-extracted features.

[0123] The output of each network is typically a score or probability distribution for multiple emotion categories. For example, by inputting the speech data to be tested into a trained speech emotion deep learning model, speech emotion scores corresponding to each emotion type can be obtained. Similarly, image emotion scores can be obtained for facial image data. Ultimately, the category with the highest score for each modality can be selected, or the final emotion label can be determined through subsequent fusion strategies.

[0124] For multimodal feature fusion and weighting mechanism, after obtaining the emotional feature vectors or scores of each modality, they need to be fused into a unified multimodal representation. The formula generally adopts the linear weighted form: F multi =αF sensor +βF face +γF voice . Among them, the vector F sensor , F face , F voice They represent the sensor, facial and voice features after their respective network encoding; the weights α, β, γ are the importance coefficients of the corresponding modalities.

[0125] The physical meaning of weights is to reflect the credibility or contribution of each modal information: higher weights indicate more reliable information for emotion judgment, while lower weights indicate that the information provided by the modality is relatively weak. For example, when the quality of the speech signal degrades in a noisy environment, the influence of the speech modality can be reduced by lowering β.

[0126] Weights can be set a priori or learned adaptively. For example, by introducing an attention mechanism, weights can be dynamically adjusted based on the loss gradients or task contributions of each modality during training, making the final fused features more expressive. Weighted fusion is simple and efficient, allowing for direct integration of multi-source information at the feature level and balancing the learning rates or coverage of different modalities through weighting.

[0127] In practical applications, a common strategy is to first perform "balanced fusion" (B fusion), adjust the weights according to the convergence of each modality network, and then further distribute the weights according to the actual contribution of the modality to the task through attention (A fusion). The final fused feature F multi It will be input into the subsequent classifier to output the final emotion label.

[0128] The following is an example of step one: Taking a smart manufacturing workshop as an example, the system monitors the status of workers using visual cameras, microphones, and wearable physiological sensors. First, in the data collection phase, the camera continuously captures the worker's face and posture, the voice sensor captures their commands and conversations, and a wristband records real-time heart rate and skin conductance signals. Next, preprocessing and feature extraction are performed: the camera video is decompressed into frame-by-frame images, and face detection and expression normalization are performed; the microphone audio is framed and spectrally analyzed; and the physiological signals are segmented and filtered to remove noise. The processed visual, voice, and sensor features are then fed into their respective emotion classification models. The visual branch uses a CNN to extract facial expression features and outputs emotion probabilities. The voice branch uses an LSTM model to extract intonation features and outputs an emotion score. The physiological branch inputs features such as heart rate and skin conductance into the classifier to output the emotional state.

[0129] Then enter the feature fusion stage: the emotional features or score vectors output by the three channels are fused according to formula F multi =αF sensor +βF face +γF voice A weighted summation is performed. For example, if preliminary experiments indicate that visual information is the most reliable, β can be increased. A higher α can also be assigned to physiological modalities to capture implicit stress states. Weights can be adaptively adjusted through model training or set through expert experience. The fused composite emotion vector represents the worker's overall emotional state and is fed into the final decision module to output an emotion label (such as "normal," "stressed," or "fatigued").

[0130] Finally, during the intelligent decision-making phase, the system feeds the identified emotions back to the production control system or management platform. For example, if an operator is identified as "overly stressed" or "fatigued," the system can adjust the operating pace of automated equipment, remind the worker to rest, or initiate a safety alert to prevent human error. Compared to single-channel visual or voice recognition, this multi-channel solution captures user emotions more comprehensively, effectively enhancing the reliability and accuracy of intelligent decision-making, enabling the shop floor automation system to make more appropriate adjustments based on human psychological state.

[0131] Heterogeneous graph modeling and self-supervised pre-training

[0132] The core of this step is to map the preprocessed multimodal temporal features onto a graph structure, and use graph neural networks (GNNs) to learn the association patterns between each node and the whole from a large amount of unlabeled data, thereby providing rich and robust features for subsequent classification. The following is a detailed explanation based on the construction of heterogeneous graphs and the forward calculation of graph neural networks.

[0133] 3.1 Construction of Heterogeneous Graph

[0134] In industrial scenarios, different data sources often have different physical or logical roles. We map them into nodes in the graph, where each node represents a specific perception source or its extracted key features. These nodes include sensor nodes, emotion nodes, and high-order feature nodes. The feature vector of each physical sensor (such as temperature, pressure, vibration) at time t is considered as a node, which is the sensor node. The "facial expression feature vector", "speech emotion score vector", and "physiological signal feature vector" extracted in step 2 correspond to the three emotion nodes in the graph respectively. Optionally, some "virtual nodes" representing device status or regional summaries are introduced into the graph as high-order feature nodes to capture a wide range of associations.

[0135] The "edges" between nodes represent the spatial relationship, functional dependency, or semantic association between data sources. For example, the temperature nodes of two adjacent machines are physically adjacent, and different sensor nodes on the same device have strong functional associations.

[0136] Specifically, if nodes i and j come from the same device, they are connected by an edge with a larger weight; if they are adjacent devices, they are connected by an edge with a medium weight; if there is no direct connection, no edge is connected or a "virtual edge" with a weight of zero is connected. The weights of all edges are summarized to form the adjacency matrix A, where A ij >0 indicates the existence of a connection, and the value reflects the strength of the association.

[0137] In actual GNN operations, the adjacency matrix A needs to be normalized to avoid excessive or weak information propagation caused by differences in node degrees. Commonly used symmetric normalization methods are shown in Equations (8), (9), and (10):

[0138]

[0139] In formula (8), formula (9) and formula (10), A represents the original adjacency matrix, A ij >0 indicates the strength of the direct physical / functional connection between nodes i and j; I represents the identity matrix, which is used to add self-loops in the graph to ensure that each node's own information is also involved in the propagation; represents the adjacency matrix after adding the self-loop; represents the degree matrix, diagonal elements is the sum of all connection weights of node i; Represents the symmetric normalized adjacency matrix, which is used to balance the information content of nodes with different degrees during GCN propagation.

[0140] The significance of the above normalization process is that for nodes with more connections (frequent sampling or dense associations), their information is appropriately scaled during propagation to maintain the overall network information balance.

[0141] 3.2. Graph Neural Network Forward Computation and Formula Derivation

[0142] Using the above normalized adjacency matrix We take the most commonly used graph convolutional layer (GCN) as an example, and its information update formula is shown in formula (11):

[0143]

[0144] Represents the feature matrix of all N nodes in the lth layer, the initial H (0) The modal features extracted from each node itself;

[0145] is the linear transformation weight to be learned in the lth layer;

[0146] σ(·) is a nonlinear activation function (such as ReLU);

[0147] Multiply This is equivalent to the "weighted average of neighbor features". From a practical point of view, a node will absorb the information of its directly related nodes, thereby expanding its receptive field layer by layer and learning the interaction features of the same layer or across modalities.

[0148] 3.3 Self-supervised pre-training tasks and loss design

[0149] Without manual labeling, we design two types of self-supervised tasks to enable GNN to learn the inherent rules of nodes in the graph structure:

[0150] 3.31. Node Feature Recovery (NodeReconstruction)

[0151] Idea: Randomly mask the input features of some nodes and train the network to use its neighbor information to predict the masked original features.

[0152] Formula: Let the set of obscured nodes be M, and the reconstruction feature output by the Lth layer of the network is The true feature is x i , define the reconstruction loss as shown in formula (12):

[0153]

[0154] In formula (12), M is the set of nodes that are randomly masked and used for self-supervised reconstruction; i Represents the original feature vector of node i, which is also the real feature; the reconstructed feature output by the Lth layer of the network is It also represents the node features of network prediction reconstruction; ||·|| represents the Euclidean norm of the vector;

[0155] This reconstruction loss forces the model to use the associations of nodes (i.e., physical or functional neighbors) to complete information, and can learn the statistical dependencies and spatial distribution patterns between modalities.

[0156] 3.32. Contrastive Learning

[0157] Node pairs that are “close” in time or space are used as positive samples, and other node pairs are used as negative samples. By distinguishing between positive and negative samples, the discriminative power of node representation is enhanced.

[0158] Formula: Use z i 、z jrepresents the final embedding vector of node i, j after the graph encoder, introduces the temperature parameter τ, and defines the InfoNCE contrast loss, as shown in Equation (13):

[0159]

[0160] In formula (13), P is the set of positive sample pairs; sim(·,·) is the vector similarity (such as dot product or cosine similarity);

[0161] This loss allows the model to cluster the embeddings of nodes that are closely related physically or functionally (such as different sensors on the same device), while pushing away unrelated or negatively correlated nodes, thereby enhancing the distinguishability of the representation.

[0162] 3.32 Joint Pre-training

[0163] The final self-supervised total loss is the weighted sum of two terms, and its mathematical calculation process is shown in formula (14):

[0164] ψ pre =λ rec ψ rec +λ c ψ c (14)

[0165] where λ rec ,λ c is a hyperparameter used to balance the importance of the two.

[0166] The entire training process is achieved by iteratively optimizing {W (l)} parameters, which enables the network to perform feature completion and distinguish between positive and negative node pairs, thereby capturing the intrinsic structure and evolution laws of multimodal data.

[0167] In this step, graph modeling uses nodes and edges to reproduce the spatial and functional associations between multimodal sources, naturally mapping the physical process flow, equipment layout, and human-computer interaction to the graph structure; normalization and graph convolution ensure the stability of information propagation, avoiding excessive smoothing or information blocking caused by excessive degree differences; self-supervised pre-training learns the extraction and interaction rules of each modal feature through neighbor recovery and comparative learning in the absence of label scenarios, so that the downstream classification module can achieve high accuracy with only a small amount of annotation.

[0168] After completing the above detailed steps, the graph neural network has already mastered the distribution patterns and structured relationships of multimodal and heterogeneous data in the cloud, providing a powerful, low-cost, and continuously updated feature foundation for subsequent manifold classification and intelligent decision-making modules.

[0169] Manifold feature extraction and Riemannian geometry classification

[0170] This step aims to convert the fused multimodal time series features into statistical covariance form and perform robust classification in the non-Euclidean space (manifold) constructed by symmetric positive definite (SPD) matrices. The following is a detailed explanation in three parts.

[0171] 4.1. Construction of covariance matrix

[0172] The meaning of statistical covariance: In the Industrial Internet, each aligned multimodal time window (e.g., 1 to 5 seconds) contains multiple sensor, physiological, and emotional features. These channels may have interdependencies: for example, machine vibration and temperature often rise and fall together, and an operator's heart rate and voice tone fluctuate together when they are nervous.

[0173] Covariance matrix C∈R d×d The i,j elements of describe the common fluctuation trend between the characteristic sequences of channel i and channel j. Its mathematical expression is shown in formula (15):

[0174]

[0175] where x i (t) is the value of the i-th feature at time t, is its mean; x j (t) is the value of the j-th feature at time t, is its mean; T is the window length, that is, the number of sampling points involved in the statistics;

[0176] The diagonal elements reflect the variance (energy or jitter) of a single channel, while the off-diagonal elements reflect the degree of synchronization or reverse change between the two channels, comprehensively characterizing the internal correlation structure of the multimodal system in this time period.

[0177] 4.2 Derivation of SPD manifold and Riemann distance

[0178] The SPD matrix set is not a flat Euclidean space, but a Riemannian manifold with curvature. Directly using the Euclidean distance metric in this space will distort the statistical structure and affect classification performance. The Riemannian distance is geometrically defined as the length of the shortest geodesic line connecting two points.

[0179] The Riemann distance formula is established. By giving two covariance matrices C1 and C2, the Riemann geodesic distance can be written as formula (16):

[0180]

[0181] In formula (16), and are the inverse of the square root of the matrix and itself, which is equivalent to "whitening C1" and locally flattening the space; C2 is the covariance matrix to be measured; log(·) is the logarithm of the matrix, which maps the flattened points to the symmetric matrix in the tangent space; ||·|| F Represents the Frobenius norm, which is equivalent to calculating the Euclidean length of the difference in the tangent space.

[0182] The whole derivation process first calculates C1 is transformed into the identity matrix so that the metric reflects only the relative structure of the two channels. The eigenvalue space of the SPD matrix is ​​then logarithmized using the matrix logarithm, converting the geodesic distance into a straight-line distance in tangent space. Finally, the differences in all eigenvectors are aggregated to obtain a scalar distance. This distance is equivalent to the cumulative change along the smoothest path of statistical structure change and better reflects the true distance of the coupling mode between channels than a simple Euclidean difference.

[0183] 4.3. Solution of the geometric mean of categories

[0184] On the SPD manifold, if N covariance samples of a certain category (such as "normal" or "abnormal") have been collected {C k}, its geometric mean Q * It is defined as the point that minimizes the sum of the squares of the total geodesic distances. Its mathematical expression is shown in formula (17):

[0185]

[0186] Iterative solution (Karcher mean), this optimization has no closed-form solution, and the following iterative algorithm is usually used, as shown in formula (18):

[0187]

[0188] In formula (19), C k Represents the kth covariance sample under the same category;

[0189] The iterative process is shown in Equation (19) and Equation (20):

[0190]

[0191] In formula (19), Q (t) represents the current geometric mean estimate obtained at iteration t; S represents the average of the logarithmic deviations of all samples relative to the current estimate, located in the tangent space;

[0192]

[0193] In formula (20), exp(·) is the matrix exponential, corresponding to the geodesic back to the manifold; iterate until ||S| is small enough.

[0194] Each iteration first quantifies the deviation of each sample from the current mean point using logarithmic mapping in the tangent space, then aggregates (averages) these deviations and maps them back to the original manifold, gradually finding the most representative central SPD matrix.

[0195] 4.4 Online Classification and Decision Making

[0196] Geodesic distance comparison: The geometric mean {Q normal ,Q fault ,...}.

[0197] The real-time covariance C of the newly arrived new , calculate its Riemann distance with each mean, as shown in formula (21):

[0198] d k =d B (C new ,Q k )(twenty one)

[0199] k=1,2,...

[0200] Select the category corresponding to the minimum distance as C new The predicted label of Q k represents the geometric mean of the kth class; d k Represents the geodesic distance of the new sample belonging to the kth class. The smaller the distance, the more likely it is.

[0201] Confidence estimation: The confidence score can be defined based on the distance distribution, as shown in Equation (22) and Equation (23):

[0202]

[0203] Where ε is a small constant to prevent division by zero, usually set to 10 -6 .

[0204] The smaller the distance, the more consistent the new sample is with the central statistical structure of the category, and the higher the confidence level.

[0205] This step uses the covariance matrix to quantify the dynamic correlation of multimodal features; the Riemann geodesic distance provides a true semantic "shortest path" measurement for SPD manifolds with non-zero curvature; the geometric mean iteration achieves high-precision solution of category centers on the manifold; and online classification achieves real-time discrimination of new samples through distance comparison and confidence estimation.

[0206] Real-time decision making and cloud visualization

[0207] The purpose of this step is to integrate the output of the previous steps in the cloud, combine preset rules and model inference to make real-time decisions, and visualize the decision results (such as device health status and emotional prompts) on the operating terminal or AR glasses.

[0208] The main process of this step is to upload the classification results and emotional labels to the decision engine and generate decision outputs based on business rules. For example, if it is determined that the operator's emotions are abnormal or the probability of equipment failure is high, the alarm information will be immediately pushed to the monitoring system through the message queue. At the same time, status icons or annotations (such as "Emotion: Worry" and "Equipment Temperature Too High") can be rendered on the supporting AR display device to enhance comprehensibility. The entire system is deployed on the cloud platform, using elastic computing to achieve low-latency processing, and edge nodes can synchronize data to ensure real-time performance.

[0209] The following is a detailed explanation in four parts, which includes the formal representation of the decision logic.

[0210] 5.1. Combination and weight allocation of decision inputs

[0211] The input sources of decision-making mainly consist of three parts: classification confidence, sentiment labels and environmental indicators.

[0212] The classification confidence comes from the confidence score of the category to which the new sample belongs in step 4. cat ;

[0213] The emotion label comes from the emotion probability distribution {p normal ,p tension ,p fatigue ,...};

[0214] The environmental indicators come from the sensor fusion results of steps 2 and 3, such as temperature and vibration over-limit probability score env ;

[0215] The system then divides the above three scores into weights ω obtained by expert experience or online learning. cat ,ω emo ,ω env Perform linear weighting to form the total decision score. The weighting process is shown in formula (24):

[0216] S total =ω cat score cat +ω emo max(p tension ,p fatigue )+ω env score env (twenty four)

[0217] where max(ptension ,p fatigue ) represents the most likely "negative" emotion probability at the moment, score env The degree to which each environmental sensor indicator exceeds the limit is comprehensively calculated.

[0218] This weighted formula quantifies equipment health, personnel status, and environmental risks into a unified score, facilitating subsequent threshold judgment and automated response.

[0219] 5.2 Decision Threshold and Execution Mechanism

[0220] First, set the thresholds. Based on historical data and safety regulations, determine the two-level thresholds:

[0221] Warning threshold θ warn :When S total ≥θ warn And S total <θ alarm When the system issues a "caution" level prompt;

[0222] Alarm threshold θ alarm :When S total ≥θ alarm When the emergency occurs, the "emergency" level linkage control or safety shutdown is triggered.

[0223] Usually choose θ warn ≈0.6,θ alarm ≈0.8 (the normalized total score range is [0,1][0,1]) and can be dynamically adjusted based on online feedback.

[0224] The execution logic of threshold setting is as follows:

[0225] If S total <θ alarm : Maintain normal operation without prompting;

[0226] If θ warn ≤S total ≤θ alarm : Send voice or interface prompts "Pay attention to personnel status / equipment operation" through the control system and record the log;

[0227] If S total ≥θ alarm : Immediately execute linkage strategies, such as initiating a safety shutdown or alarm, and notify site management.

[0228] 5.3 Augmented Reality Visual Mapping

[0229] The three-dimensional coordinates of key points on site (machines, operating tables, corridors) are pre-calibrated in the digital twin model, and three visual icons with different states are pre-defined: green check mark (normal), yellow exclamation mark (warning), and red lightning (alarm).

[0230] When the system decides "Warning" or "Alarm", the decision result and the coordinates of the relevant objects are sent to the AR terminal through WebSocket;

[0231] The AR terminal superimposes its own positioning (SLAM or installed base station) with the target coordinates, and renders the corresponding graphics in the user's field of view, aligned with the real scene.

[0232] The above rendering process includes: the AR device listens to cloud push and receives coordinates; calls the device SLAM interface to project the three-dimensional point to the current camera frame; draws the corresponding icon at the projection position, and can attach a text prompt (such as "device overheating"); the primitive continues to be displayed in the scene until the decision status is cleared or updated by a new decision.

[0233] The principle flow of the method of the present invention is as follows: first, multimodal heterogeneous data from sensors, audio and video, and physiological equipment are collected at the industrial site, and uniformly preprocessed through steps such as filtering, normalization, and time synchronization; then, the feature information of each mode is extracted separately, and the covariance matrix of the fused features is constructed in each time window to capture the joint fluctuation relationship between the modes; then, multiple covariance feature windows are constructed into a graph structure, feature embedding is performed through a graph neural network, and self-supervised learning is used to improve the generalization ability under unlabeled or small sample conditions; in the classification stage, the system compares the Riemannian geometric distance between the new sample and the center of each type of state, and outputs the state recognition result and confidence level; finally, the operator's emotional state and environmental risk score are combined to perform weighted fusion on the results, and a comprehensive risk level is output. The status prompt and safety response are realized through AR visualization or linkage control modules, thereby constructing a real-time intelligent decision-making mechanism for human-machine-environment integration for industrial scenarios.

[0234] See Figure 2 The present invention also provides a cloud computing-based multimodal heterogeneous data real-time fusion and intelligent decision-making system. The system is suitable for industrial Internet scenarios, especially for applications of multi-source heterogeneous signal collaborative monitoring and intelligent early warning. The modules of the system are interconnected through the network to form a complete data processing and decision-making system, specifically including the following modules:

[0235] The data acquisition module, used to collect data from various modalities at industrial sites, consists of a physical sensor submodule, an audio and video acquisition submodule, and a physiological signal acquisition submodule. The physical sensor submodule, deployed on the equipment itself or in its surroundings, collects operating condition information such as temperature, pressure, vibration, and current. The audio and video acquisition submodule uses a camera and microphone to capture operator behavioral characteristics such as facial expressions and intonation in real time. The physiological signal acquisition submodule uses wearable devices to capture physiological indicators such as heart rate and skin charge, reflecting the operator's workload and emotional state.

[0236] The data acquisition module sends the collected data to the data processing module after local buffering and time stamping.

[0237] The data processing and preprocessing module is used to unify the format, clean, synchronize and normalize multimodal data. It is specifically composed of a filtering and denoising submodule, a numerical normalization submodule, and a time synchronization and resampling submodule. The filtering and denoising submodule uses median filtering and bandpass filtering to remove pulse interference and high-frequency noise. The numerical normalization submodule normalizes each signal to zero mean and unit variance by channel to solve the problem of dimensional inconsistency. The time synchronization and resampling submodule is based on global clock calibration and aligns modal signals of different frequencies to a unified time base through linear interpolation.

[0238] The data processing and preprocessing module sends the processed data to the subsequent modeling module.

[0239] The feature modeling and state encoding module is used to extract high-level state representations from preprocessed data. It specifically includes a modal feature extraction submodule, a feature fusion submodule, and a covariance modeling submodule. The modal feature extraction submodule uses trained deep network models to extract facial expression features, audio emotion features, and physiological fluctuation features from video, voice, and physiological data respectively. The feature fusion submodule fuses the above-mentioned multi-source modal feature vectors into a unified emotion state vector to represent the operator's current psychological and physiological state. The covariance modeling submodule calculates the covariance matrix of the modal joint features based on a sliding time window to construct a statistical representation of the multimodal state.

[0240] The graph modeling and graph neural network embedding module is used to construct multiple state windows into a graph structure and extract feature embedding through the graph neural network. The module consists of a graph construction submodule, a graph convolution submodule, and a self-supervised training submodule; the graph construction submodule uses each time window as a graph node and sets edge weights based on time proximity, spatial position, or device function association; the graph convolution submodule uses a graph convolutional neural network (GCN) to propagate and aggregate node features to generate context-aware node embedding representations; the self-supervised training submodule uses masked reconstruction and graph comparison learning methods to achieve network pre-training in an unlabeled state, thereby improving the model's generalization ability.

[0241] The classification discrimination and risk fusion module is used to classify the current state and fuse multi-source risk information to output the risk level. The module includes a manifold classification submodule, a confidence calculation submodule, a decision fusion submodule, and a level judgment submodule; the manifold classification submodule compares the covariance characteristics of each window with the Riemannian geometric mean of each category obtained by offline calculation, and classifies based on the geodesic distance; the confidence calculation submodule outputs the confidence of each category match according to the distance result; the decision fusion submodule fuses the classification confidence, operator emotional risk, and environmental indicator score, and generates a total risk score through weighted calculation; the level judgment submodule sets a dual threshold, classifies the score and outputs "normal", "warning" or "alarm" status.

[0242] Visual presentation and linkage control module, which is used to visualize the judgment results and drive the linkage equipment or prompt terminal; this module includes AR visualization sub-module, linkage control sub-module, and log recording and tracing sub-module; the AR visualization sub-module superimposes the current risk status on the on-site equipment, operator or work area location, and prompts status changes through icons and text; the linkage control sub-module triggers linkage control according to the "alarm" status, including sound and light alarms, equipment shutdown instructions, operation authority restrictions and other actions; logging and tracing sub-module: uploads all judgment results, feature data, operation feedback, etc. to the cloud platform to support historical tracking and model retraining.

[0243] The above are only specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A cloud computing-based multimodal heterogeneous data real-time fusion and intelligent decision-making method, characterized by: The specific steps include: Collect raw signals from multiple data sources including physical sensors, audio and video modalities, and physiological modalities, filter and normalize the signals, and map the signals of each modality to a unified time axis through a time synchronization mechanism; Extract feature vectors from each modal signal after preprocessing, and fuse emotion-related features such as facial expressions, voice intonation, and physiological indicators into a unified emotional state vector; Construct the covariance matrix of the joint features of each modality within the preset time window to represent the multimodal joint change features within the time period; Represent each covariance matrix as a graph node, construct weighted edge relationships based on the temporal proximity or physical association between nodes, and use graph neural networks for embedding learning to obtain the representation vector of each node; Without manual labeling, the graph neural network is trained jointly through the node masking reconstruction task and the contrastive learning task. The Riemannian geometric mean of the covariance samples of known categories is calculated as the central representation of each category. Compare the covariance features of the current time window with the centers of each category, output the state confidence score based on the Riemann geodesic distance, and perform weighted fusion with the emotional state score and environmental risk score to form the final risk score; The risk score is compared with the preset warning threshold and alarm threshold to trigger the visualization of normal, warning or alarm status, and the control system is linked to issue control instructions when necessary.

2. The method according to claim 1, characterized in that The multimodal data collection includes: physical sensor data of industrial equipment, camera images and voice signals of operators, and physiological data such as heart rate and skin conduction of wearable devices.

3. The method according to claim 1, characterized in that The covariance matrix is ​​constructed by performing a mean removal process on the normalized features in each time window and then calculating the autocorrelation and cross-correlation to form a statistical matrix.

4. The method according to claim 1, wherein The graph neural network adopts a graph convolutional network structure, the input is a covariance vector representation, and the output is a node embedding vector.

5. The method according to claim 1, characterized in that The node reconstruction task of self-supervised training is achieved by randomly masking the input features of some graph nodes and minimizing their restoration error.

6. The method according to claim 1, wherein The class center calculates the geometric mean of multiple covariance matrices on the SPD manifold by Karcher iteration method.

7. The method according to claim 1, characterized in that The geodesic distance in state classification is calculated using Riemann logarithmic mapping and Frobenius norm.

8. The method according to claim 1, characterized in that The visual prompt uses AR devices or industrial terminals as a carrier, and presents the status by superimposing icons and texts through three-dimensional space coordinates.

9. The method according to claim 1, characterized in that Risk score fusion uses a weighted summation of the three scores and sets dual thresholds to classify the current status as normal, warning, or alarm.

10. A cloud computing-based multimodal heterogeneous data real-time fusion and intelligent decision-making system, characterized by: include: The data acquisition module is used to collect data from different modalities at the industrial site. The data acquisition module caches and timestamps the collected data locally and then sends it to the data processing and preprocessing module. Data processing and preprocessing module, which is used to unify the format, clean, synchronize and normalize multimodal data; Feature modeling and state encoding module, which is used to extract high-level state representation from preprocessed data; The graph modeling and graph neural network embedding module is used to construct multiple state windows into a graph structure and extract feature embeddings through a graph neural network. This module consists of a graph construction submodule, a graph convolution submodule, and a self-supervised training submodule. The graph construction submodule uses each time window as a graph node and sets edge weights based on temporal proximity, spatial location, or device function association. The graph convolution submodule uses a graph convolutional neural network to propagate and aggregate node features to generate context-aware node embedding representations. The self-supervised training submodule uses masked reconstruction and image comparison learning methods to achieve network pre-training in an unlabeled state; Classification and risk fusion module, which is used to classify the current status and fuse multi-source risk information to output the risk level; Visual presentation and linkage control module, which is used to visualize the judgment results and drive linkage devices or prompt terminals.

Citation Information

Cited By

  • High-temperature PZT preparation monitoring method and system based on full-process monitoring

    CN121038578A

  • Smart home remote control method based on Internet of Things

    CN121300109A

  • Vehicle cabin interior and exterior active perception interaction method, device and system and vehicle

    CN121375825A

  • Hardware-level network isolation and security protection method based on data processing unit

    CN121984793A