A dynamic risk weight allocation method based on multimodal hazard source identification
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-17
- Publication Date
- 2026-08-14
AI Technical Summary
[0004]有鉴于此,本申请的实施例提出了一种基于多模态危险源识别的风险权重动态分配方法,旨在克服现有工业安全监测方法中跨模态信息融合能力不足、风险权重分配静态僵化以及系统部署成本高昂的技术问题,进而能够显著提升危险源多模态识别的准确性与鲁棒性、实现风险权重的动态自适应分配、并降低系统部署成本
Smart Images

Figure CN122573084A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this application relate to the field of machine learning technology, and in particular to a method for dynamic allocation of risk weights based on multimodal hazard source identification. Background Technology
[0002] Industrial production safety monitoring and risk prevention are crucial foundations for ensuring personnel safety, reliable equipment operation, and sustainable enterprise development. Traditional industrial safety monitoring mainly relies on manual inspections and fixed threshold alarm systems, which suffer from inherent problems such as response delays, strong dependence on human experience, and difficulty in adapting to dynamic production changes. In recent years, artificial intelligence methods, represented by deep learning, have been widely applied in the field of industrial safety, mainly including target detection methods based on computer vision and sensor signal analysis methods based on industrial Internet of Things (IIoT) data.
[0003] However, potential safety risks in industrial sites often involve complex interactions of information from multiple sensing modalities, and signals from a single modality are insufficient to fully characterize the evolution of a hazard. On the other hand, risk weighting is a core aspect of industrial safety risk assessment. Current mainstream weighting methods primarily employ static or semi-static approaches such as the Analytic Hierarchy Process (AHP) and Grey Relational Analysis (GRA). Once the weight coefficients are determined, they remain unchanged for a long period, making it impossible to dynamically adjust the relative importance of risk factors based on changes in real-time monitoring data, thus limiting the accuracy and timeliness of risk assessment. Summary of the Invention
[0004] In view of this, embodiments of this application propose a dynamic risk weight allocation method based on multimodal hazard source identification, which aims to overcome the technical problems of insufficient cross-modal information fusion capability, static and rigid risk weight allocation, and high system deployment cost in existing industrial safety monitoring methods. This method can significantly improve the accuracy and robustness of multimodal hazard source identification, achieve dynamic adaptive allocation of risk weights, and reduce system deployment costs.
[0005] To achieve the above objectives, embodiments of this application propose a dynamic risk weight allocation method based on multimodal hazard source identification, the method comprising the following steps: Raw multimodal data from industrial production sites are collected, and the raw multimodal data is preprocessed and aligned to the time dimension to obtain processed multimodal data. Single-modal features are extracted from each modality in the multimodal data to obtain the single-modal feature vectors corresponding to each modality. Then, through gating fusion and cross-modal attention mechanisms, the extracted single-modal feature vectors are fused into a unified comprehensive hazard source characterization vector. Based on the comprehensive hazard source characterization vector, multiple types of hazard sources are identified, and real-time hit feature scores of each type of hazard source are output. Based on the real-time hit feature scores of multiple types of hazards and their historical evolution trajectories within a dynamic time window, a temporal state matrix is constructed, and causal temporal attention encoding is performed on the temporal state matrix to predict the future risk trends and risk evolution rates of each hazard. Based on the real-time hit feature scores, future risk trends, and hazard coupling and correlation information contained in the time-series state matrix of each hazard source, the dynamic risk weight of each type of hazard source at the current moment is obtained through a dynamic adaptation scoring function. The dynamic risk weight is used to sort and classify the responses of each type of hazard source in order to achieve industrial management.
[0006] To achieve the above objectives, embodiments of this application also propose a dynamic risk weight allocation system based on multimodal hazard source identification, the system comprising: The multimodal data acquisition and preprocessing module is used to acquire raw multimodal data from the industrial production site, and to preprocess and align the raw multimodal data according to the time dimension to obtain processed multimodal data. The cross-modal fusion hazard identification module is used to extract single-modal features from each modality of multimodal data to obtain the single-modal feature vectors corresponding to each modality. Through gating fusion and cross-modal attention mechanisms, the extracted single-modal feature vectors are fused into a unified comprehensive hazard representation vector. Based on the comprehensive hazard representation vector, multiple types of hazards are identified, and the real-time hit feature scores of each type of hazard are output. The time-series risk weight dynamic allocation module is used to construct a time-series state matrix based on the real-time hit feature scores of multiple types of hazards and their historical evolution trajectories within a dynamic time window. It then performs causal time-series attention encoding on the time-series state matrix to predict the future risk trends and risk evolution rates of each hazard. Based on the real-time hit feature scores, future risk trends, and hazard coupling and correlation information contained in the time-series state matrix, it obtains the dynamic risk weights of each type of hazard at the current moment through a dynamically adapted scoring function. These dynamic risk weights are used to rank and classify the responses of various hazard sources to achieve industrial management.
[0007] To achieve the above objectives, embodiments of this application also propose an electronic device, including: a processor and a memory, wherein the memory stores instructions executable by the processor, and the processor is configured to execute the instructions such that the electronic device can implement a risk weight dynamic allocation method based on multimodal hazard source identification as described above.
[0008] To achieve the above objectives, embodiments of this application also propose a computer-readable storage medium storing a computer program that, when executed by a processor, enables a dynamic risk weight allocation method based on multimodal hazard source identification as described above.
[0009] This application proposes a dynamic risk weight allocation method based on multimodal hazard source identification. The method involves collecting raw multimodal data from industrial production sites, preprocessing the raw multimodal data, and aligning it with the time dimension to obtain processed multimodal data. Single-modal features are extracted from each modality, and these features are fused into a unified hazard source comprehensive representation vector through gating fusion and cross-modal attention mechanisms. Based on this comprehensive hazard source representation vector, multiple types of hazard sources are identified, and real-time hit feature scores for each type of hazard source are output. A temporal state matrix is constructed based on the real-time hit feature scores of multiple hazard sources and their historical evolution trajectories within a dynamic time window. Causal temporal attention encoding is then applied to the temporal state matrix to predict the future risk trends and risk evolution rates of each hazard source. Finally, based on the real-time hit feature scores, future risk trends, and hazard source coupling and correlation information contained in the temporal state matrix, a dynamic risk weight for each type of hazard source at the current moment is obtained through a dynamically adapted scoring function.
[0010] First, this scheme adaptively fuses heterogeneous single-modal features into a unified comprehensive hazard source representation vector through gated fusion and cross-modal attention mechanisms. This solves the problem of decreased recognition rate of single modalities (e.g., relying solely on visual modality) in complex industrial environments such as lighting changes, occlusion, and noise. Compared to traditional single-modal methods, multimodal fusion can capture complementary information of hazard sources across different perceptual dimensions, thereby improving the accuracy and robustness of hazard source identification. Second, because this scheme is based on real-time hit feature scores, historical evolution trajectories, and future risk trend predictions, it dynamically adapts the scoring function to calculate the weights of various hazard sources in real time. This allows the risk weights to follow the dynamic changes in production conditions, thus overcoming the limitations of traditional single-modal methods. This approach overcomes the core deficiency of static weights in failing to reflect real-time risk changes, thus improving the timeliness and accuracy of risk assessment. Third, it models the historical evolution trajectory of hazards within a dynamic time window and predicts future risk trends through causal temporal attention encoding. Fourth, this scheme considers the historical coupling and mutual information between hazards when constructing the temporal state matrix, and integrates the co-occurrence relationships between hazards through an attention mechanism in the dynamic adaptation scoring function. When multiple hazards occur simultaneously or have causal relationships, this method can automatically amplify the risk weights of associated hazards, avoiding underestimation or overestimation caused by isolated assessments of individual hazards, making risk assessment more closely resemble the complex interactive scenarios of real industrial sites.
[0011] Based on this, the proposed solution can overcome the technical problems of insufficient cross-modal information fusion capability, static and rigid risk weight allocation, and high system deployment cost in existing industrial safety monitoring methods. It can significantly improve the accuracy and robustness of multimodal identification of hazardous sources, achieve dynamic adaptive allocation of risk weights, and reduce system deployment costs.
[0012] Optionally, single-modal features are extracted from each modality of the multimodal data to obtain single-modal feature vectors corresponding to each modality. These single-modal feature vectors are then fused into a unified hazard source comprehensive representation vector through gating fusion and cross-modal attention mechanisms. This process includes: extracting single-modal feature vectors for each modality using a lightweight backbone network; wherein the number of parameters in the lightweight backbone network is lower than the target value; and different lightweight backbone networks correspond to different target values; performing mean aggregation on all single-modal feature vectors to generate a global context vector; concatenating each single-modal feature vector with the global context vector and inputting it into the corresponding modality gating unit to obtain a modality gating value, and using the modality gating value to weight the corresponding single-modal feature vector; inputting the gated and weighted modality feature vectors into a multi-head cross-modal attention module, using features from one modality as queries and features from other modalities as keys and values, to perform fine-grained semantic alignment between modalities; and concatenating and reducing the dimensionality of the aligned modality features to output a unified hazard source comprehensive representation vector.
[0013] Optionally, based on the comprehensive hazard source representation vector, multiple types of hazard sources are identified, and real-time hit feature scores of each type of hazard source are output. This includes: based on the comprehensive hazard source representation vector, after multi-dimensional mapping through two fully connected layers, the identification probability distribution of each type of preset hazard source is output through a Softmax classifier; based on the identification probability distribution of each type of hazard source, the real-time hit feature scores of each type of hazard source are output, and at the same time, through a visual-language joint reasoning module, natural language description text corresponding to the identified hazard source is generated.
[0014] Optionally, based on the real-time hit feature scores of multiple types of hazards and their historical evolution trajectories within a dynamic time window, a temporal state matrix is constructed, and causal temporal attention encoding is applied to the temporal state matrix to predict the future risk trend and risk evolution rate of each hazard. This includes: for any target hazard among the multiple types of hazards, obtaining the hit feature score, rate of change, continuous hit duration, and historical coupling mutual information at the current moment to generate the temporal state vector of the target hazard at the current moment, and then constructing the temporal state matrix at the current moment using the temporal state vectors of all hazards; inputting the temporal state matrices of various types of hazards at the current moment and the rolling historical state matrix into a causal temporal attention model composed of a Transformer temporal encoding layer; wherein the length of the rolling historical state matrix is within a preset range; wherein the causal temporal attention model sets a causal mask to direct the information flow of historical time steps to the current moment, and uses a multi-head attention mechanism and a feedforward neural network to encode the evolution law of hazards within the time window, outputting the evolution risk trend and risk evolution rate of each hazard at a preset future moment.
[0015] Optionally, based on the real-time hit feature scores, future risk trends, risk evolution rates, and hazard coupling and correlation information contained in the time-series state matrix of each hazard source, the dynamic risk weights of various hazard sources at the current moment are obtained through a dynamically adapted scoring function, specifically including: No. Hazardous sources at all times Dynamic risk weights Calculated using the following formula: ; in, Indicates the first Class of hazardous sources; Indicates time The temporal state matrix of all hazards; Represents the first prediction made by the causal temporal attention model. Such hazardous sources in the future The evolving risk trends at any given moment; This indicates a dynamically adapted scoring function, and By integrating the current hazard source hit scores Risk trends Risk evolution rate Coupling mutual information of the probability of co-occurrence between hazard sources and other hazard sources get; For the first Basic bias terms for hazard sources.
[0016] Optionally, dynamically adapt the scoring function. The generation process includes: Target hazard source Temporal state vector As a query, retrieve the temporal state matrix of all hazards. Using these as keys and values, a scaled dot product attention computation is performed to obtain attention-aggregated features that fuse the co-occurrence coupling mutual information among hazard sources. ; Real-time hit feature score Evolutionary risk trends Risk evolution rate and attention aggregation features The data is then concatenated to form a multi-dimensional risk feature vector. ; Multidimensional risk feature vector The final dynamic adaptation scoring function is obtained through linear layer mapping. The calculation formula is as follows: ; in, The weight matrix is a learnable weight matrix; It is a biased scalar.
[0017] Optionally, raw multimodal data from the industrial production site is collected, and preprocessed and time-dimensional aligned to obtain processed multimodal data, including: video surveillance data, infrared thermal imaging data, IoT sensor data, and text log data from the industrial site. During preprocessing, frame sequences are extracted from the video surveillance data according to a fixed time window and image enhancement is performed; pixel values are normalized for the infrared thermal imaging data; sliding window slicing is used for the IoT sensor data, followed by denoising and time-frequency domain feature extraction; and word segmentation and word embedding vectorization are performed on the text log data. Frame-level time alignment is performed on the preprocessed multimodal data using a unified time reference, and missing time slices caused by sampling rate differences are filled in using interpolation to obtain time-synchronized multimodal data.
[0018] Optionally, the method provided in this embodiment is based on a three-layer collaborative architecture, which includes an edge layer, a cloud layer, and a terminal layer. The acquisition of multimodal data is performed by the terminal layer device, the preprocessing, the extraction of single-modal features, the identification of multiple types of hazards, and the dynamic allocation of risk weights are performed by the edge layer, and the cloud layer is used to perform model updates and long-term historical data analysis, and to send the updated model parameters to the edge layer. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies of this application will be briefly introduced below. Obviously, the following drawings are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. The drawings described herein are only used to explain this application and are not intended to limit this application.
[0020] Figure 1 This is a flowchart of a method for dynamically allocating risk weights based on multimodal hazard source identification, provided in one embodiment of this application; Figure 2 This is a structural diagram of a cross-modal gated attention fusion network provided in one embodiment of this application; Figure 3 This is a structural diagram of the Transformer coding layer in a temporal attention model provided in one embodiment of this application; Figure 4 This is a structural diagram of a three-tier collaborative architecture provided in one embodiment of this application; Figure 5 This is a schematic diagram of a risk weight dynamic allocation system based on multimodal hazard source identification provided in another embodiment of this application; Figure 6 This is a schematic diagram of the structure of an electronic device provided in another embodiment of this application. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. Those skilled in the art will understand that many technical details have been presented in the embodiments of this application to facilitate better understanding. However, the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments. The division of the following embodiments is for ease of description and should not constitute any limitation on the specific implementation of this application. The following embodiments can be combined with and referenced by each other without contradiction.
[0022] The following is a unified explanation of the English abbreviations mentioned in this application.
[0023] 1. Softmax Classifier: A linear multi-classifier based on the Softmax function, which maps input features to the probabilities of each class and outputs a sum of 1. It is often used in the output layer of neural networks.
[0024] 2. Transformer Temporal Encoding Layer: A sequence encoding module built on the self-attention mechanism of the Transformer architecture, used to capture long-range dependencies in time series.
[0025] 3. Word2Vec: A static word embedding model that learns distributed vector representations of words from a large-scale corpus, where semantically similar words are closer together in the vector space.
[0026] 4. FastText: An improved word representation model based on Word2Vec, it introduces sub-word units, which can better handle out-of-vocabulary words and low-frequency words, and has better word vector quality and training speed.
[0027] 5. DistilBERT: A lightweight pre-trained language model obtained by compressing the Bidirectional Encoder Representations from Transformers (BERT) model through knowledge distillation. It has about 60% of the number of parameters of BERT, improves inference speed by about 60%, and has little loss of accuracy.
[0028] 6. CLIP-ViT variant: refers to the use of Contrastive Language–Image Pre-training (CLP). The visual Transformer model is pre-trained using the CLIP (Contrast Learning Platform) contrastive learning framework.
[0029] 7. YOLOv11n: The nano-lightweight specification of the 11th version of the YOLO (You Only Look Once) series of object detection algorithms. "n" stands for nano, which has the smallest number of parameters and computational cost, and is suitable for real-time object detection scenarios at the edge.
[0030] 8. Xavier Initialization: One of the most classic weight initialization methods in deep learning. Its core purpose is to keep the signal variance of each layer of the neural network stable, avoiding the vanishing or exploding gradient problems.
[0031] Industrial production safety monitoring and risk prevention are crucial foundations for ensuring personnel safety, reliable equipment operation, and sustainable enterprise development. Traditional industrial safety monitoring mainly relies on manual inspections and fixed threshold alarm systems, which suffer from inherent problems such as response delays, strong dependence on human experience, and difficulty in adapting to dynamic production changes. In recent years, artificial intelligence methods, represented by deep learning, have been widely applied in the field of industrial safety, mainly including target detection methods based on computer vision and sensor signal analysis methods based on industrial Internet of Things (IIoT) data.
[0032] However, potential safety risks in industrial sites often involve complex interactions of information from multiple sensing modalities, and signals from a single modality are insufficient to fully characterize the evolution of a hazard. On the other hand, risk weighting is a core aspect of industrial safety risk assessment. Current mainstream weighting methods primarily employ static or semi-static approaches such as the Analytic Hierarchy Process (AHP) and Grey Relational Analysis (GRA). Once the weight coefficients are determined, they remain unchanged for a long period, making it impossible to dynamically adjust the relative importance of risk factors based on changes in real-time monitoring data, thus limiting the accuracy and timeliness of risk assessment.
[0033] In view of this, embodiments of this application propose a dynamic risk weight allocation method based on multimodal hazard source identification, which aims to overcome the technical problems of insufficient cross-modal information fusion capability, static and rigid risk weight allocation, and high system deployment cost in existing industrial safety monitoring methods. This method can significantly improve the accuracy and robustness of multimodal hazard source identification, achieve dynamic adaptive allocation of risk weights, and reduce system deployment costs.
[0034] One embodiment of this application proposes a dynamic risk weight allocation method based on multimodal hazard source identification, applied to electronic devices. The electronic device can be a terminal layer or a server; this embodiment and subsequent embodiments will use a server as an example. The implementation details of the dynamic risk weight allocation method based on multimodal hazard source identification proposed in this embodiment are described below. The following implementation details are provided for ease of understanding and are not essential for implementing this solution.
[0035] The specific process of the risk weight dynamic allocation method based on multimodal hazard source identification proposed in this embodiment can be described as follows: Figure 1 As shown, it includes: Step 101: Collect raw multimodal data from the industrial production site, and perform preprocessing and time dimension alignment on the raw multimodal data to obtain processed multimodal data.
[0036] In this embodiment, to comprehensively perceive the safety status of the industrial production site, it is first necessary to collect data from multiple heterogeneous data sources. These heterogeneous data sources may include visible light cameras, infrared thermal imagers, and various IoT sensors.
[0037] For example, an industrial production site can be an industrial workshop for chemical or mechanical processing.
[0038] For example, raw multimodal data may include video surveillance data, infrared thermal imaging data, IoT sensor data, and text log data.
[0039] Video surveillance data can be obtained by capturing video frame sequences from visible light cameras deployed in production areas. For example, in a chemical plant, cameras can be pointed at critical areas such as reaction vessels, pipe connections, and operating platforms.
[0040] Infrared thermal imaging data can be thermal images acquired by an infrared thermal imager, used to monitor the surface temperature distribution of equipment, abnormal hot spots, etc.
[0041] IoT sensor data can be time-series signals collected by various IoT sensors deployed on devices, including but not limited to vibration sensors, temperature sensors, pressure sensors, gas concentration sensors, etc.
[0042] Text log data can include structured or unstructured text information such as equipment operation logs, maintenance records, operating procedures, and safety specifications.
[0043] In one possible embodiment, step 101 includes: accessing video surveillance data, infrared thermal imaging data, IoT sensor data, and text log data from the industrial site as multimodal data; during preprocessing, extracting frame sequences from the video surveillance data according to a fixed time window and performing image enhancement processing, performing pixel value normalization processing on the infrared thermal imaging data, performing denoising and time-frequency domain feature extraction on the IoT sensor data using a sliding window slice, and performing word segmentation and word embedding vectorization processing on the text log data; performing frame-level time alignment on the preprocessed multimodal data with a unified time reference, and using interpolation to fill in missing time slices caused by sampling rate differences, thereby obtaining time-synchronized multimodal data, and thus obtaining preprocessed and aligned multimodal data.
[0044] Understandably, since different modal data have different sampling frequencies, data formats, and physical meanings, they need to be preprocessed and aligned in terms of time dimension.
[0045] During preprocessing, video surveillance data and infrared thermal imaging data can be denoised and their dimensions normalized; for example, video frame sequences can be extracted according to a fixed time window (e.g., 30 frames per batch). For nighttime or low-light scenes, adaptive histogram equalization is applied for image enhancement. Infrared images are normalized, mapping pixel values to... Interval.
[0046] It can filter, fill in missing values, and normalize IoT sensor data; for example, it can use a sliding window mechanism to slice continuous sensor signals. The window length can be adaptively adjusted according to the characteristics of the monitored object, such as 2 to 10 seconds. It can perform bandpass filtering to denoise high-frequency signals such as vibration, current, and voltage; and perform outlier removal and linear interpolation to fill in slowly varying signals such as temperature and pressure.
[0047] It can perform operations such as word segmentation and stop word removal on text data. For example, it can segment and vectorize log text, and use lightweight word embedding models (such as Word2Vec, FastText, or DistilBERT) to convert the text into 512-dimensional embedding vectors.
[0048] During the temporal alignment process, interpolation or resampling techniques can be used to align all preprocessed modal data to a common time reference. For example, all data can be aligned at 1-second intervals to ensure that, at the same timestamp, each modal data reflects the same moment in the field. Considering the differences in sampling rates among modalities, for time slices where a particular modality's data is missing, nearest neighbor interpolation or linear interpolation can be used to fill in the gaps, ensuring spatial synchronization of multimodal features in the temporal dimension.
[0049] Step 102: Extract single-modal features from each modal data in the multimodal data to obtain the single-modal feature vector corresponding to each modal data. Then, through gating fusion and cross-modal attention mechanism, fuse the extracted single-modal feature vectors into a unified comprehensive hazard source characterization vector.
[0050] This step is the core of achieving deep fusion of multimodal information. Its specific implementation relies on a carefully designed cross-modal feature fusion network, the structure of which can be found by referring to [reference needed]. Figure 2 The diagram shows the architecture of a cross-modal gated attention fusion network. This architecture includes a visual feature branch, an infrared thermal imaging feature branch, a sensor temporal feature branch, a text semantic feature branch, a multi-head cross-modal cross-attention module, a feature fusion fully connected layer, and a vector output module.
[0051] The system extracts features from four modalities: visual feature branch, infrared thermal imaging feature branch, sensor temporal feature branch, and text semantic feature branch. These features are extracted from the preprocessed and aligned original visual images (i.e., video surveillance data), infrared thermal imaging data, IoT sensor data, and text log data, respectively. Then, within each feature branch, the corresponding modal features are concatenated with the global context vector. Based on the modal gating units in the feature branches, the gated weighted modal features are determined and input into the multi-head cross-modal attention module. This module uses features from one modality as the query and features from the other modalities as keys and values for cross-modal attention operations. Finally, the output features of the multi-head cross-modal attention module are concatenated and subjected to dimensionality reduction and nonlinear transformation through a feature fusion fully connected layer. The output module then outputs a unified comprehensive hazard source representation vector.
[0052] In one possible embodiment, step 102 includes: Step 1021: Use a lightweight backbone network to extract the single-modal feature vectors of each modality data.
[0053] The lightweight backbone network has fewer parameters than the target value, and the target value varies for different lightweight backbone networks. For example, the target value can be the number of parameters of a traditional backbone network.
[0054] For example, if the lightweight backbone network is a CLIP-ViT distillation model for the visual branch, the target value can be 8 million (M); if the lightweight backbone network is a lightweight ResNet-18 model for the infrared branch, the target value can be 3.2M; if the lightweight backbone network is a TCN sensor encoder, the target value can be 500,000; if the lightweight backbone network is a DistilBERT text encoder, the target value can be 12M.
[0055] To reduce computational load and adapt to edge deployment, this embodiment uses a lightweight backbone network to extract features from four modalities of data: For feature extraction from video surveillance data, a lightweight vision-language model with fewer parameters than the target value (e.g., a distilled CLIP-ViT variant) can be used as a visual encoder to extract 512-dimensional global scene features and local region features from video frames. Simultaneously, lightweight detection models such as YOLOv11n are introduced to quickly locate key regions in the image, performing subsequent fine-grained feature encoding only on the selected regions, thus reducing the burden of full-image inference.
[0056] For feature extraction of infrared thermal imaging data, an optimized lightweight ResNet-18 structure is used to extract features such as device hotspot distribution and abnormal temperature areas from infrared images, and output a 256-dimensional feature vector.
[0057] For feature extraction from IoT sensor data, a Temporal Convolutional Network (TCN) is used to encode the sensor's time window data. The TCN employs causal convolution and dilated convolution design, with the number of parameters controlled below 500,000, and outputs a 256-dimensional feature vector.
[0058] For text log data, a lightweight BERT model with knowledge distillation is used as the text encoder to output a 512-dimensional text semantic feature vector.
[0059] Through the above steps, we can obtain the feature vector set of the four modes. These correspond to the single-modal features of video surveillance data, infrared thermal imaging data, IoT sensor data, and text log data, respectively.
[0060] Step 1022: Perform mean aggregation on all single-modal feature vectors to generate a global context vector.
[0061] For example, first, a global context vector is computed. Specifically, the features of the four modalities Each of the four mapped features is unified to the same dimension using a single-layer linear mapping. Then, the elements of these four mapped features are aggregated using element-wise mean to obtain an initial global aggregate vector. Finally, a sigmoid activation layer is used to normalize the initial global aggregate vector to obtain a global context vector of fixed dimension. Global context vector Used to represent overall scene information for all modalities.
[0062] Step 1023: Concatenate each single-modal feature vector with the global context vector and input it into the corresponding modal gating unit to obtain the modal gating value, and use the modal gating value to weight the corresponding single-modal feature vector.
[0063] This step aims to bridge the semantic gap between multimodal data and achieve adaptive feature fusion. Its core innovation lies in the modality gating unit and cross-modal attention mechanism.
[0064] First, design a modal gating unit for each modality. This unit uses the current modal features and global context vector The concatenated result is used as input, through a The function calculates a gate value. : ; in, and These are learnable parameters. Indicates the type of mode; gate value The range of values is This represents the reliability of the modality's contribution to hazard identification in the current time slice. Then, using... Weight the corresponding single-modal features: ; This leads to the obtained modal features after gating and weighting. , , , .
[0065] Step 1024: Input the gated weighted feature vectors of each modality into the multi-head cross-modal attention module, using the features of one modality as the query and the features of the other modalities as the keys and values, to perform fine-grained semantic alignment between modalities.
[0066] The four modal features are gated and weighted. , , , The input is fed into a multi-head cross-modal attention module. This module uses a 4-8 head structure, with fixed preferred parameters: 4 heads for both multi-head cross-modal attention and causal temporal attention, a scaling factor of 64, and dropout = 0.1. The module uses features from one modality as the query and features from the other three modalities as the keys and values, performing cross-modal attention operations. For example, the attention output for the visual modality is calculated as follows: ; In this way, each modality can obtain complementary information from other modalities, achieving fine-grained semantic alignment between modalities.
[0067] Step 1025: Concatenate and reduce the dimensionality of the aligned modal feature vectors to output a unified comprehensive hazard source characterization vector.
[0068] Finally, the output feature vectors of the multi-head cross-modal attention are concatenated and then subjected to dimensionality reduction and nonlinear transformation through a two-layer fully connected network to output a unified comprehensive hazard source representation vector. It is understandable that the comprehensive characterization vector of a hazard source... It is a comprehensive and robust digital description of the current safety status of industrial sites.
[0069] It is understandable that this embodiment, through the complete data flow of multimodal data from the original input to the unified representation vector, can clearly demonstrate the three core innovative structures of modal gating unit, global context vector, and multi-head cross-modal attention, intuitively demonstrating the adaptive weighted fusion logic of each modality, and solving the semantic gap problem of multi-source heterogeneous data.
[0070] Step 103: Based on the comprehensive hazard source characterization vector, identify multiple types of hazard sources and output the real-time hit feature scores of each type of hazard source.
[0071] In one possible embodiment, step 103 includes: based on the comprehensive characterization vector of the hazard source, after multidimensional mapping through two fully connected layer networks, outputting the identification probability distribution of various preset hazard sources through a Softmax classifier; based on the identification probability distribution of various hazard sources, outputting the real-time hit feature score of various hazard sources, and simultaneously generating natural language description text corresponding to the identified hazard sources through a visual-language joint reasoning module.
[0072] Specifically, the comprehensive characterization vector of the hazard source can be... The input is fed into a hazard identification network consisting of a two-layer fully connected network and a Softmax classifier. The hidden layer dimensions in the two fully connected layers are 256 and 128, respectively. The Softmax classifier in this network outputs the identification probability distribution of various preset hazard sources.
[0073] For example, the probability distribution of the identification of various preset hazards can indicate that the current scene belongs to a preset hazard. The probability of each category of hazard sources. The preset hazard source categories can be configured according to the application scenario. For example, the preset hazard source categories include: not wearing a safety helmet, not wearing protective clothing, unauthorized operation of machinery, intrusion into a dangerous area, localized overheating of equipment, abnormal vibration, and toxic gas leakage, etc.
[0074] The probability distribution of identification of various preset hazard sources corresponds to the first... The probability value of a hazard source is defined as its real-time hit feature score. .
[0075] Optionally, a natural language text description of the discovered hazard can also be generated simultaneously through the visual language joint reasoning module.
[0076] The visual-language joint reasoning module is used to generate natural language text descriptions for identified hazards, such as "In area A of the reactor, a worker was detected not wearing a safety helmet."
[0077] Step 104: Based on the real-time hit feature scores of multiple types of hazards and their historical evolution trajectories within the dynamic time window, construct a time-series state matrix and perform causal time-series attention encoding on the time-series state matrix to predict the future risk trends and risk evolution rates of each hazard.
[0078] It is understood that this embodiment aims to solve the problem of static allocation of risk weights.
[0079] For example, based on the cascaded input of real-time multi-hazard source identification results, the historical evolution trajectory of the hazard sources within a dynamic time window is combined. The real-time multi-hazard source identification results can include hazard source category labels and their hit feature scores. The historical evolution trajectory can include identified hazard source categories, hit feature scores, and changes in risk status. A hazard source status tracking updater is constructed using a temporal attention encoder and a gated loop unit to achieve real-time risk weight updates for each hazard source on the production site.
[0080] In one possible embodiment, step 104 includes: Step 1041: For any target hazard source among multiple hazard sources, obtain the hit feature score, rate of change, continuous hit time length and historical coupling mutual information at the current moment to generate the temporal state vector of the target hazard source at the current moment, and then construct the temporal state matrix at the current moment through the temporal state vectors of all hazard sources.
[0081] Among them, the causal temporal attention model sets a causal mask to direct information from historical time steps to the current moment, and uses a multi-head attention mechanism and a feedforward neural network to encode the evolution pattern of hazards within the time window, outputting the evolution risk trend and risk evolution rate of each hazard at a future preset time.
[0082] Specifically, when constructing the temporal state matrix at the current moment, a hazard source database can be built. : ; in, This represents the total number of preset hazard source categories.
[0083] At any moment , define the first Hit characteristic score of hazard source This is provided by the identification probability distribution. Then, the hit score values for the corresponding categories can be extracted according to the hazard source classification index.
[0084] Construction dimension Temporal state matrix .in, This represents the dimension of the risk evolution state of a hazard source.
[0085] The risk evolution state dimension consists of dynamic feature values built into the system, including the rate of evolution of hazard hit feature scores within the cumulative time window. That is, the rate of risk evolution at the current moment; the duration of continuous hits by the hazard source. and sources of danger Historical coupling mutual information with other hazard sources Three dimensions.
[0086] Among them, the rate of change within the cumulative time window This can represent the difference between the scores of two consecutive hits. The duration of consecutive hits by a hazard source. This can represent a time-step accumulation count starting from the first hit. The historical coupling and mutual information between the hazard source and other hazard sources. It can represent the hazard source within a dynamic time window. With the source of danger The co-occurrence correlation strength. The temporal state matrix is specifically expressed as: .
[0087] Current risk evolution rate Indicates the first Hazardous sources at all times The risk evolution rate characterizes the trend and speed of change of the hit feature score within the sliding historical window. It is calculated using the sliding window least squares slope, which can suppress rate distortion caused by single-point noise fluctuations. The calculation formula is as follows: ; in, Indicates the time step length of the sliding history window; Indicates the index of discrete time steps within the window; Indicates the first Hazardous sources at all times Hit feature score, range of values .
[0088] Understandable, This indicates that the risk of this type of hazard source is on the rise; the larger the value, the faster the risk is rising. This indicates that the risk is trending downwards; This indicates that the risk status remains stable.
[0089] Historical Coupling Mutual Information The mutual information obtained from the binary hit state is used to quantify the degree to which multiple hazard sources couple and amplify the risk. Specifically: For the Class and the For hazard-like entities, the mutual information between them is calculated based on the frequency statistics of their states within a window. : ; in, Indicates the first Hazard hit status is The marginal probability is obtained from the frequency statistics within the window; Indicates the first Hazard hit status is The marginal probability; Indicates the first Hazard hit status is And the first Hazard hit status is The joint probability.
[0090] For example, when the feature score is hit in real time If the condition is met, it is determined as a hit state 1; otherwise, it is 0.
[0091] It is understandable that mutual information The larger the value, the stronger the co-occurrence correlation between the two types of hazards and the higher the coupling risk.
[0092] Based on this, the first Hazardous sources at all times Historical Coupling Mutual Information Take the maximum value of the mutual information between this type of hazard and all other hazard sources: ; in, This represents the total number of preset hazard source categories.
[0093] Step 1042: Input the temporal state matrix of various hazards at the current moment and the rolling historical state matrix into the causal temporal attention model composed of the Transformer temporal coding layer.
[0094] The length of the scrolling historical state matrix is within a preset range. For example, the preset range can be 30 to 120 time steps.
[0095] The causal temporal attention model is based on a multi-head causal attention mechanism, which enables the model to capture the dynamic coupling relationship between hazard source categories over a long period of time.
[0096] For example, the temporal state matrix of the hazard source at the current moment can be used. and rolling history state matrix The input is fed into a causal temporal attention model consisting of a Transformer temporal coding layer. In this model, the hazard state vector at each time step is linearly mapped to generate a time step embedding vector, and then absolute or relative position encoding is added to preserve temporal order information.
[0097] After obtaining the time-series state matrix at the current moment in the above embodiments, the length of the dynamic time window can be retained as follows: Rolling historical state matrix To construct the evolution trajectory of hazard sources. Dynamic time window. The number of time steps can be set from 30 to 120. Then, the historical co-occurrence relationship between the hazard to be weighted and other hazard sources is encoded through a temporal self-attention mechanism, which serves as a coupling factor for fine-tuning the risk weights of similar hazard sources in the future.
[0098] First, the state vector at each time step is linearly mapped and superimposed with positional encoding to preserve temporal sequence information. Then, it is processed through a multi-head attention layer with causal masking. Causal masking ensures that the model can only see historical information up to the current moment when making predictions, preventing information leakage. This attention mechanism can capture the dynamic coupling relationships between different hazard categories over long time spans.
[0099] For example, the structure of the Transformer encoding layer in a causal temporal attention model is as follows: Figure 3 As shown, the internal structure of the Transformer temporal coding layer includes: a linear mapping layer, a positional coding layer, a multi-head causal attention layer, and two feedforward neural networks. The linear mapping layer connects to the input layer of the causal temporal attention model, and the two feedforward neural networks connect to the output layer of the causal temporal attention model.
[0100] Linear mapping layers can process temporal state matrices Perform linear transformation to generate original time step features The positional encoding layer is used to overlay relative positional codes, preserving temporal sequence information. The multi-head causal attention layer uses a four-head causal mask to allow only historical timestep information to flow to the current moment, preventing future information leakage. Two feedforward neural network layers are used to perform nonlinear transformations of the features, outputting the encoded temporal features. .
[0101] It should be noted that during the encoding process, the four-dimensional state vector of the hazard source at each time step is mapped to a unified feature space, and the positional encoding is superimposed to distinguish the temporal sequence; the causal multi-head attention traverses the state of the hazard source within the entire historical window to capture the long-term co-occurrence coupling mutual information of different hazard sources; and a two-layer feedforward neural network mines the nonlinear laws of risk evolution and outputs the encoded features required for predicting future risk trends.
[0102] The updated formula can be summarized as follows: ; ; ; ; The attention output is nonlinearly transformed using a two-layer feedforward network to output the future location of the hazard source. Evolutionary risk trends at any moment and the rate of risk evolution .
[0103] Step 105: Based on the real-time hit feature scores, future risk trends, and hazard coupling and correlation information contained in the temporal state matrix of each hazard source, the dynamic risk weight of each type of hazard source at the current moment is obtained by dynamically adapting the scoring function.
[0104] Among them, dynamic risk weights are used to sort and classify various hazards for response, so as to achieve industrial management.
[0105] For example, this step enables real-time dynamic allocation and updating of risk weights. First, a dynamically adaptable scoring function is defined. The dynamic adaptation scoring function can calculate the first digit using multi-dimensional information. At present, the type of hazard source The comprehensive risk score. The hazard coupling and correlation information contained in the time-series state matrix can be the co-occurrence coupling mutual information of hazard sources. .
[0106] For example, It can be set to 10 time steps to predict risk trends over the next 10 seconds.
[0107] Dynamically adapt scoring function The generation process includes: First, identify the target hazard source Temporal state vector As a query, retrieve the temporal state matrix of all hazards. Using these as keys and values, the scaled dot product attention computation yields the fused mutual information of co-occurrence coupling among hazard sources. Attention aggregation features Understandably, this feature quantifies the risk gain of other hazards on the current hazard, and embeds the co-occurrence coupling mutual information of the hazards. .
[0108] Then, the real-time hit feature score will be... Evolutionary risk trends Risk evolution rate and attention aggregation features The data is then concatenated to form a multi-dimensional risk feature vector. ; in, .
[0109] Finally, the multi-dimensional risk feature vector The final dynamic adaptation scoring function is obtained through linear layer mapping. The calculation formula is as follows: ; in, The weight matrix is a learnable weight matrix; It is a biased scalar.
[0110] For example, Xavier can be used for initialization. Initialize to 0; training can use the AdamW optimizer.
[0111] Based on dynamic adaptation scoring function , No. Hazardous sources at all times Dynamic risk weights Calculated using the following Softmax normalization formula: ; in, Indicates the first Class of hazardous sources; Indicates time The temporal state matrix of all hazards; Represents the first prediction made by the causal temporal attention model. Such hazardous sources in the future The evolving risk trends at any given moment.
[0112] in, Indicates the first Class of hazardous sources; Indicates time The temporal state matrix of all hazards; Represents the first prediction made by the causal temporal attention model. Such hazardous sources in the future The evolving risk trends at any given moment; This indicates a dynamically adapted scoring function, and By integrating the current hazard source hit scores Risk trends Risk evolution rate Coupling mutual information of the probability of co-occurrence between hazard sources and other hazard sources get; For the first Basic bias terms for hazard sources.
[0113] For example, basic bias terms This represents the inherent severity level of the hazard, which can prevent hazards that occur infrequently but have serious consequences from being over-punished upon their first triggering. It should be noted that the inherent severity level can be preset according to industry safety standards, such as GB 6441-1986 (i.e., the Classification Standard for Enterprise Employee Injury and Fatal Accidents).
[0114] For example, when the basic bias term A value of 1.2 indicates a severity level of critical; the basic bias item. A value of 0.6 indicates a severity level of moderate severity; the base bias term. A value of 0.2 indicates a severity level of mild to severe.
[0115] The dynamic risk weight is calculated using the above formula. .
[0116] This dynamic risk weight satisfy: ; That is, the sum of the dynamic risk weights of all hazards is 1.
[0117] Optionally, the weight update cycle can be dynamically set according to the production scenario, for example, 1-10 seconds.
[0118] Finally, based on the calculated dynamic risk weights, various hazards can be ranked, and corresponding graded response strategies can be generated. For example, the top three hazards with the highest weights can be highlighted in red on the risk situation map, triggering corresponding audible and visual alarms or linking with industrial control systems to implement protective measures.
[0119] To achieve low-cost, high-efficiency deployment of industrial safety monitoring systems and ensure rapid on-site response for risk weight allocation tasks with extremely high real-time requirements, the method provided in this embodiment is based on a three-layer collaborative architecture including an edge layer, a cloud layer, and a terminal layer.
[0120] In one possible embodiment, the method provided in this embodiment is executed based on a three-layer collaborative architecture, which includes an edge layer, a cloud layer, and a terminal layer; the acquisition of multimodal data is performed by the terminal layer device, the preprocessing, the extraction of single-modal features, the identification of multiple types of hazards, and the dynamic allocation of risk weights are performed by the edge layer, and the cloud layer is used to perform model updates and long-term historical data analysis.
[0121] like Figure 4 As shown, Figure 4 This is a structural diagram of the three-layer collaborative architecture in this embodiment. The three-layer collaborative architecture includes an edge layer, a cloud layer, and a terminal layer; The terminal layer consists of various sensing terminal devices deployed at the industrial production site, used to perform the raw acquisition of multimodal data as described in step 101. The terminal layer includes visible light cameras, infrared thermal imagers, and various IoT sensors.
[0122] Understandably, the terminal layer is only responsible for the raw data collection and pushing the raw data streams (i.e., video surveillance data, infrared thermal imaging data, IoT sensor data, and text log data mentioned above) to the edge in real time via communication networks such as industrial Ethernet, 5G, or Wi-Fi 6. Since the terminal layer itself does not perform any complex computing tasks, it can reduce the computing power requirements of the terminal hardware, making it compatible with the existing monitoring equipment and sensor nodes in the factory without the need for large-scale hardware replacement.
[0123] The edge layer typically consists of edge intelligent gateways, industrial control computers, or edge servers deployed in production sites or factory server rooms. The edge layer is responsible for all core real-time computing tasks, from data preprocessing to risk weight allocation. For example, it performs data preprocessing and time alignment in step 101 of the above embodiments, single-modal feature extraction and cross-modal fusion in step 102, hazard identification in step 103, time-series state modeling and risk trend prediction in step 104, and dynamic risk weight allocation in step 105.
[0124] Understandably, after performing the aforementioned calculations, the edge layer will send the final decision results (e.g., risk weight list, hazard identification results, and graded alarm signals) to the terminal layer's display screen or audible and visual alarm, while simultaneously uploading key time-series data and a summary of model inference results to the cloud. To ensure real-time performance, the model deployed at the edge layer can be pruned, quantized, and its parameter count controlled, thereby reducing latency control in the entire end-to-end decision-making closed loop.
[0125] The cloud layer consists of high-performance computing clusters deployed in remote data centers or cloud servers, used to execute non-real-time tasks that require a large amount of computing power and global data.
[0126] For example, the cloud layer can train and iteratively update lightweight models deployed at the edge (e.g., cross-modal fusion networks, causal temporal attention models) to improve the model's generalization ability and recognition accuracy in new scenarios. After training is complete, the cloud will distribute the updated model parameters to the edge for deployment and updates.
[0127] This application proposes a dynamic risk weight allocation method based on multimodal hazard source identification. The method involves collecting raw multimodal data from industrial production sites, preprocessing the raw multimodal data, and aligning it with the time dimension to obtain processed multimodal data. Single-modal features are extracted from each modality, and these features are fused into a unified hazard source comprehensive representation vector through gating fusion and cross-modal attention mechanisms. Based on this comprehensive hazard source representation vector, multiple types of hazard sources are identified, and real-time hit feature scores for each type of hazard source are output. A temporal state matrix is constructed based on the real-time hit feature scores of multiple hazard sources and their historical evolution trajectories within a dynamic time window. Causal temporal attention encoding is then applied to the temporal state matrix to predict the future risk trends and risk evolution rates of each hazard source. Finally, based on the real-time hit feature scores, future risk trends, and hazard source coupling and correlation information contained in the temporal state matrix, a dynamic risk weight for each type of hazard source at the current moment is obtained through a dynamically adapted scoring function.
[0128] First, this scheme adaptively fuses heterogeneous single-modal features into a unified comprehensive hazard source representation vector through gated fusion and cross-modal attention mechanisms, thereby solving the problem of decreased recognition rate of single modality (e.g., relying solely on visual modality) in complex industrial environments such as lighting changes, occlusion, and noise. Compared with traditional single-modal methods, multimodal fusion can capture complementary information of hazard sources in different perceptual dimensions, thereby improving the accuracy and robustness of hazard source identification.
[0129] Second, because this scheme is based on real-time hit feature scores, historical evolution trajectories, and future risk trend predictions, it calculates the weights of various hazards in real time through dynamic adaptation of the scoring function. This allows the risk weights to follow the dynamic changes in production conditions, overcoming the core defect that static weights cannot reflect real-time risk changes and improving the timeliness and accuracy of risk assessment.
[0130] Third, by using causal temporal attention coding, the historical evolution trajectory of hazards within a dynamic time window is modeled, and future risk trends are predicted.
[0131] Fourth, this scheme considers the historical coupling and mutual information between hazard sources when constructing the time-series state matrix, and integrates the co-occurrence correlation between hazard sources through an attention mechanism in the dynamic adaptation scoring function; when multiple hazard sources occur at the same time or have causal relationships, this method can automatically amplify the risk weight of associated hazard sources, avoid the underestimation or overestimation caused by isolated evaluation of each hazard source, and make risk assessment closer to the complex interaction scenarios of real industrial sites.
[0132] Based on this, the proposed solution can overcome the technical problems of insufficient cross-modal information fusion capability, static and rigid risk weight allocation, and high system deployment cost in existing industrial safety monitoring methods. In this way, it can significantly improve the accuracy and robustness of multimodal identification of hazardous sources, realize dynamic adaptive allocation of risk weights, and reduce system deployment costs.
[0133] The steps described above are for clarity only. In implementation, they can be combined into one step, or some steps can be broken down into multiple steps, as long as they involve the same logical relationship, they are all within the scope of protection of this application. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, without changing the core design of the algorithm and process, are also within the scope of protection of this application.
[0134] Another embodiment of this application proposes a dynamic risk weight allocation system based on multimodal hazard source identification. The details of this dynamic risk weight allocation system based on multimodal hazard source identification are described below. The following implementation details are provided for ease of understanding and are not essential for implementing this example. Figure 5 This is a schematic diagram of the structure of a risk weight dynamic allocation system based on multimodal hazard source identification proposed in this embodiment, including: The multimodal data acquisition and preprocessing module 210 is used to acquire raw multimodal data from the industrial production site, and to preprocess and align the raw multimodal data according to the time dimension to obtain processed multimodal data. The cross-modal fusion hazard identification module 220 is used to extract single-modal features from each modality of multimodal data to obtain the single-modal feature vectors corresponding to each modality. Through gating fusion and cross-modal attention mechanism, the extracted single-modal feature vectors are fused into a unified hazard comprehensive characterization vector. Based on the hazard comprehensive characterization vector, multiple types of hazards are identified, and the real-time hit feature scores of each type of hazard are output. The time-series risk weight dynamic allocation module 230 is used to construct a time-series state matrix based on the real-time hit feature scores of multiple types of hazards and their historical evolution trajectories within a dynamic time window, and to perform causal time-series attention encoding on the time-series state matrix to predict the future risk trends and risk evolution rates of each hazard. Based on the real-time hit feature scores, future risk trends, risk evolution rates of each hazard, and the hazard coupling and correlation information contained in the time-series state matrix, the dynamic risk weights of each type of hazard at the current moment are obtained through a dynamically adapted scoring function. The dynamic risk weights are used to sort and classify the responses of various types of hazards to achieve industrial management.
[0135] It is not difficult to see that this embodiment is a system embodiment corresponding to the above method embodiments, and this embodiment can be implemented in conjunction with the above method embodiments. The relevant technical details and technical effects mentioned in the above method embodiments are still valid in this embodiment, and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above method embodiments.
[0136] It is worth mentioning that all modules and units involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, to highlight the innovative aspects of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed in this application; however, this does not mean that other units do not exist in this embodiment.
[0137] Another embodiment of this application provides an electronic device, such as Figure 6 As shown, it includes: a memory 31 and a processor 32. The memory 31 stores instructions that the processor 32 can execute. When the processor 32 is configured to execute the instructions, the electronic device can implement a risk weight dynamic allocation method based on multimodal hazard source identification as described in the above method embodiment.
[0138] The memory and processor are connected via a bus, which includes any number of interconnecting buses and bridges, connecting various circuits of one or more processors and the memory. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single component or multiple components, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which further receives data and transmits it to the processor.
[0139] The processor manages the bus and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory is used to store data used by the processor during operation.
[0140] Another embodiment of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, can implement a dynamic risk weight allocation method based on multimodal hazard source identification as described in the above method embodiments.
[0141] That is, those skilled in the art will understand that all or part of the steps in the above method embodiments can be implemented by a program instructing related hardware. The program is stored in a storage medium and includes several instructions to cause a device (such as a microcontroller, chip, etc.) or processor to execute all or part of the steps of the method described in the method embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.
[0142] Those skilled in the art will understand that the above embodiments are specific implementations of this application, and in practical applications, various changes can be made in form and detail without departing from the spirit and scope of this application. For those skilled in the art, several improvements and modifications can be made without departing from the principles of this application, and these improvements and modifications are also considered to be within the scope of protection of this application.
Claims
1. A method for dynamic allocation of risk weights based on multimodal hazard source identification, characterized in that, The method includes: Raw multimodal data from industrial production sites are collected, and the raw multimodal data is preprocessed and aligned to the time dimension to obtain processed multimodal data. Single-modal features are extracted from each modality in the multimodal data to obtain the single-modal feature vectors corresponding to each modality. Then, through gating fusion and cross-modal attention mechanisms, the extracted single-modal feature vectors are fused into a unified comprehensive hazard source characterization vector. Based on the comprehensive hazard source characterization vector, multiple types of hazard sources are identified, and real-time hit feature scores of each type of hazard source are output. Based on the real-time hit feature scores of multiple types of hazards and their historical evolution trajectories within a dynamic time window, a temporal state matrix is constructed, and causal temporal attention encoding is performed on the temporal state matrix to predict the future risk trends and risk evolution rates of each hazard. Based on the real-time hit feature scores, future risk trends, risk evolution rates, and hazard coupling and correlation information contained in the time-series state matrix of each hazard source, the dynamic risk weights of each type of hazard source at the current moment are obtained through a dynamically adapted scoring function. The dynamic risk weights are used to sort and classify the responses of each type of hazard source in order to achieve industrial management.
2. The method according to claim 1, characterized in that, The process involves extracting single-modal features from each modality of the multimodal data to obtain single-modal feature vectors corresponding to each modality. Then, through gated fusion and cross-modal attention mechanisms, the extracted single-modal feature vectors are fused into a unified hazard source comprehensive characterization vector, including: A lightweight backbone network is used to extract single-modal feature vectors for each modality of data; the number of parameters of the lightweight backbone network is lower than the target value, and different lightweight backbone networks correspond to different target values; The mean of all unimodal feature vectors is aggregated to generate a global context vector. Each single-modal feature vector is concatenated with the global context vector and input into the corresponding modal gating unit to obtain the modal gating value. The modal gating value is then used to weight the corresponding single-modal feature vector. The gated and weighted feature vectors of each modality are input into the multi-head cross-modal cross-attention module. The feature of one modality is used as the query, and the features of the other modalities are used as the key and value to perform fine-grained semantic alignment between modalities. The aligned modal feature vectors are concatenated and their dimensionality reduced to output a unified comprehensive hazard source characterization vector.
3. The method according to claim 1, characterized in that, The method of identifying multiple types of hazards based on the comprehensive hazard source representation vector and outputting real-time hit feature scores for each type of hazard source includes: Based on the comprehensive hazard source representation vector, after multidimensional mapping through two fully connected layers, the identification probability distribution of various preset hazard sources is output through the Softmax classifier; Based on the probability distribution of various hazard sources, the system outputs real-time hit feature scores for each hazard source. Simultaneously, through the visual-language joint reasoning module, it generates natural language description text for the corresponding identified hazard sources.
4. The method according to claim 1, characterized in that, The method involves constructing a temporal state matrix based on the real-time hit feature scores of multiple hazard sources and their historical evolution trajectories within a dynamic time window, and then performing causal temporal attention encoding on the temporal state matrix to predict the future risk trends and risk evolution rates of each hazard source, including: For any target hazard source among multiple hazard sources, the hit feature score, rate of change, continuous hit duration and historical coupling mutual information at the current moment are obtained to generate the temporal state vector of the target hazard source at the current moment. Then, the temporal state matrix at the current moment is constructed by the temporal state vectors of all hazard sources. The temporal state matrices of various hazards at the current moment and the rolling historical state matrices are input into a causal temporal attention model composed of a Transformer temporal encoding layer; wherein the length of the rolling historical state matrix is within a preset range. Among them, the causal time-series attention model sets a causal mask to direct information from historical time steps to the current moment, and uses a multi-head attention mechanism and a feedforward neural network to encode the evolution pattern of hazards within the time window, outputting the evolution risk trend and risk evolution rate of each hazard at a future preset time.
5. The method according to claim 4, characterized in that, Based on the real-time hit feature scores, future risk trends, risk evolution rates, and hazard coupling and correlation information contained in the time-series state matrix of each hazard source, the dynamic risk weights of each type of hazard source at the current moment are obtained through a dynamically adapted scoring function, specifically including: No. Hazardous sources at all times Dynamic risk weights Calculated using the following formula: ; in, Indicates the first Class of hazardous sources; Indicates time The temporal state matrix of all hazards; Represents the first prediction made by the causal temporal attention model. Such hazardous sources in the future The evolving risk trends at any given moment; This indicates a dynamically adaptable scoring function that integrates the current hazard hit scores. Evolutionary risk trends Risk evolution rate Co-occurrence coupling mutual information with hazard sources get; For the first The basic bias term for a class of hazards is used to characterize the inherent risk severity level of the corresponding hazard.
6. The method according to claim 5, characterized in that, Dynamically adapt scoring function The generation process includes: Target hazard source Temporal state vector As a query, retrieve the temporal state matrix of all hazards. As keys and values, the co-occurrence coupling mutual information of fused hazard sources is obtained through scaling dot product attention computation. Attention aggregation features ; Real-time hit feature score Evolutionary risk trends Risk evolution rate and attention aggregation features The data is then concatenated to form a multi-dimensional risk feature vector. ; Multidimensional risk feature vector The final dynamic adaptation scoring function is obtained through linear layer mapping. The calculation formula is as follows: ; in, The weight matrix is a learnable weight matrix; It is a biased scalar.
7. The method according to any one of claims 1 to 6, characterized in that, The process involves collecting raw multimodal data from the industrial production site, preprocessing the raw multimodal data, and aligning it to the time dimension to obtain processed multimodal data, including: It integrates video surveillance data, infrared thermal imaging data, IoT sensor data, and text log data from industrial sites as multimodal data; During the preprocessing process, the video surveillance data is extracted into frame sequences according to a fixed time window and image enhancement processing is performed. The infrared thermal imaging data is normalized for pixel values. The IoT sensor data is sliced using a sliding window and denoised and extracted for time and frequency domain features. The text log data is segmented and vectorized for word embedding. The preprocessed modal data is frame-level time aligned using a unified time reference. Missing time slices caused by sampling rate differences are filled in by interpolation to obtain time-synchronized multimodal data.
8. The method according to any one of claims 1 to 6, characterized in that, The method is executed based on a three-layer collaborative architecture, which includes an edge layer, a cloud layer, and a terminal layer. The acquisition of multimodal data is performed by the terminal layer device, while the preprocessing, single-modal feature extraction, identification of multiple types of hazards and dynamic allocation of risk weights are performed by the edge layer. The cloud layer is used to perform model updates and long-term historical data analysis, and sends the updated model parameters to the edge layer.
9. A risk weight dynamic allocation system based on multimodal hazard source identification, characterized in that, The system includes: The multimodal data acquisition and preprocessing module is used to acquire raw multimodal data from the industrial production site, and to preprocess and align the raw multimodal data according to the time dimension to obtain processed multimodal data. The cross-modal fusion hazard identification module is used to extract single-modal features from each modality of multimodal data to obtain the single-modal feature vectors corresponding to each modality. Through gating fusion and cross-modal attention mechanisms, the extracted single-modal feature vectors are fused into a unified comprehensive hazard representation vector. Based on the comprehensive hazard representation vector, multiple types of hazards are identified, and the real-time hit feature scores of each type of hazard are output. The time-series risk weight dynamic allocation module is used to construct a time-series state matrix based on the real-time hit feature scores of multiple types of hazards and their historical evolution trajectories within a dynamic time window. It then performs causal time-series attention encoding on the time-series state matrix to predict the future risk trends and risk evolution rates of each hazard. Based on the real-time hit feature scores, future risk trends, and hazard coupling and correlation information contained in the time-series state matrix, it obtains the dynamic risk weights of each type of hazard at the current moment through a dynamically adapted scoring function. These dynamic risk weights are used to rank and classify the responses of various hazard sources to achieve industrial management.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it can implement a dynamic risk weight allocation method based on multimodal hazard source identification as described in any one of claims 1 to 8.