An industrial production safety risk early warning method based on multi-modal fusion

CN122839262APending Publication Date: 2026-09-29HEFEI QINGSHEN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610993132.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

现有预警方法在处理多源监测数据时,普遍存在特征冗余大、噪声干扰强、关键风险属性提取不足以及风险判别边界不清晰的问题,导致误报率和漏报率较高,难以满足工业生产场景对实时性、准确性和稳定性的要求

Benefits of technology

本发明通过对视频图像数据、温度数据、压力数据、振动数据、气体浓度数据、电气运行数据以及作业记录数据多源异构监测数据进行统一预处理、跨模态信息流耦合分析以及候选安全风险属性筛选,能够在人员行为、设备状态、环境变化和作业过程多个维度上实现对工业生产安全风险的综合感知。相较于现有主要依赖单一数据源、固定规则或简单经验判断的预警方式,本发明利用跨模态信息流耦合、风险轨迹点阵迭代标注以及邻域粗糙集属性约简方法,对复杂工况下的高维风险信息进行有效提炼和逐层压缩,不仅提高了关键安全风险属性提取的准确性,也增强了对复合型、隐蔽型和渐进型风险的识别能力,有效提升了工业生产安全风险感知的全面性、稳定性和适应复杂工况变化的能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122839262A_ABST
    Figure CN122839262A_ABST
Patent Text Reader

Abstract

The application discloses a kind of industrial production safety risk early warning method based on multi-modal fusion, it is related to industrial safety early warning technical field, comprising: acquisition multi-source heterogeneous monitoring data, pre-processing obtains multimodal security sample library;Cross-modal information flow coupling graph is constructed, and candidate security risk attribute library is obtained;According to iteration labeling to flow-diffusion rule, labeled training sample set and sample set to be judged are obtained;Based on neighborhood rough set attribute reduction, key security risk attribute set is obtained;Improved learning vector quantization network is constructed, and optimized security risk prototype library is obtained;Local extremum-curvature joint analysis is executed, and real-time security risk determination result table is obtained;Security risk grading early warning instruction is generated, and sent to monitoring alarm platform.The application is processed by neighborhood rough set attribute reduction and improved learning vector quantization network cooperation, and the accurate identification and grading early warning of complex security risk of industrial production are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial safety early warning technology, and in particular to an industrial production safety risk early warning method based on multimodal fusion. Background Technology

[0002] As industrial production continues to evolve towards automation, intelligence, and continuity, the types of equipment, operational processes, and environmental conditions in the production process are becoming increasingly complex. Safety risks in industrial settings are also characterized by diversified sources, intertwined impact chains, and dynamic evolution. In actual production, safety risks are often not directly triggered by a single abnormal factor, but rather are the result of multiple factors working together, including abnormal personnel behavior, equipment operation fluctuations, changes in environmental parameters, and deficiencies in work management. Existing industrial production safety monitoring methods typically rely on a single data source, such as temperature, pressure, vibration, gas concentration, electrical operating status, or video images, or perform simple combination analysis of a small number of data sources, using preset rules, fixed thresholds, or a single judgment logic to output alarm results. While current methods can detect some overt anomalies, the strong coupling relationships between different risk factors in the temporal, spatial, and operational dimensions make it difficult for single data sources or simple combinations to comprehensively reflect the true safety status of industrial production processes. They struggle to identify complex risks, hidden risks, and progressive risks that accumulate gradually from multiple factors, resulting in insufficient perception capabilities and limited coverage in complex industrial scenarios.

[0003] As the scale of data acquired through industrial monitoring continues to increase, the dimensions, types, and update frequency of monitoring data have also improved. Extracting effective information closely related to safety risks from multi-source heterogeneous data has become a crucial technical challenge in the field of industrial safety early warning. Existing early warning methods generally suffer from problems such as high feature redundancy, strong noise interference, insufficient extraction of key risk attributes, and unclear risk discrimination boundaries when processing multi-source monitoring data. This leads to high false alarm and false negative rates, making it difficult to meet the real-time, accuracy, and stability requirements of industrial production scenarios. While some technical solutions attempt to introduce multiple monitoring data for joint analysis, they lack unified and effective processing paths in areas such as multimodal data fusion, candidate risk attribute construction, key attribute reduction, risk prototype optimization, and risk level determination. This makes it difficult to fully utilize the inherent correlations between different modalities of data and to generate stable and reliable early warning results under complex operating conditions and dynamic scenarios.

[0004] Therefore, how to provide an industrial production safety risk early warning method based on multimodal fusion is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] One objective of this invention is to propose a multimodal fusion-based method for early warning of safety risks in industrial production. This invention comprehensively utilizes multi-source heterogeneous monitoring data preprocessing, cross-modal information flow coupling analysis, neighborhood rough set attribute reduction, and improved learning vector quantization network data analysis and intelligent discrimination techniques. It details the entire process from multi-source heterogeneous monitoring data acquisition, preprocessing, candidate safety risk attribute extraction, key safety risk attribute screening, risk prototype optimization, to safety risk classification and early warning output. In the attribute reduction stage, it innovatively introduces high-order neighborhood spectrum segmentation iteration and information potential field contraction reduction. In the risk judgment stage, it innovatively introduces entropy positive mass flow balance rearrangement, curvature equilibrium arc surface correction, and hyperspectral entropy cointegration optimization, achieving efficient identification and accurate early warning of complex safety risks in industrial production. Compared with existing technologies, this invention has advantages such as strong multimodal risk perception capability, accurate extraction of key risk attributes, low false alarm and false negative rates, strong adaptability to complex working conditions, and ease of engineering application deployment.

[0006] An industrial production safety risk early warning method based on multimodal fusion according to an embodiment of the present invention includes: Collect multi-source heterogeneous monitoring data from industrial production sites, preprocess the multi-source heterogeneous monitoring data, and obtain a multimodal safety sample library; A cross-modal information flow coupling graph is constructed in a multimodal security sample library, and a candidate security risk attribute library is obtained based on the change of mutual information gain; Based on the candidate safety risk attribute library, risk trajectory points are mapped and iteratively labeled according to the convection-diffusion rule to obtain the labeled training sample set and the sample set to be judged; Based on neighborhood rough set attribute reduction, neighborhood relations are reconstructed and attributes are filtered in the labeled training sample set. High-order neighborhood spectrum segmentation iteration is introduced to perform spectrum decomposition, segmentation and update. Information potential energy field contraction reduction is used to perform boundary contraction, attribute compression and redundancy removal to obtain the key security risk attribute set. An initial risk prototype set is generated by the set of key safety risk attributes. An improved learning vector quantization network is constructed. The initial risk prototype set is then subjected to prototype rearrangement, boundary trimming and correlation optimization. Entropy positive mass flow balance rearrangement is introduced for global adjustment. Curvature balance processing is performed based on curvature equilibrium arc surface correction. Hyperspectral entropy cointegration optimization is used to uniformly optimize the risk prototypes, resulting in an optimized safety risk prototype library. The sample set to be judged is matched with the optimized safety risk prototype library, the continuous distance surface of the risk prototype is calculated, and the local extremum-curvature joint analysis is performed to obtain the real-time safety risk judgment result table. Based on the real-time security risk assessment results table, generate corresponding security risk classification early warning instructions and send the security risk classification early warning instructions to the monitoring and alarm platform.

[0007] Optionally, the multi-source heterogeneous monitoring data includes video image data, temperature data, pressure data, vibration data, gas concentration data, electrical operation data, and work record data.

[0008] Optionally, the preprocessing of multi-source heterogeneous monitoring data includes time alignment, outlier removal, missing value completion, noise reduction filtering, standardization, and structured coding.

[0009] Optionally, obtaining the multimodal security sample library includes: Video image data, temperature data, pressure data, vibration data, gas concentration data, electrical operation data, and work record data are collected separately. Timestamps are added to each type of data according to a unified sampling clock to obtain the original monitoring dataset. The original monitoring dataset is subjected to time alignment, outlier removal, missing value completion, noise reduction filtering, standardization and structured coding. The standardization process is to subtract the sample mean of the monitoring variable from the actual value of each monitoring variable, and then divide by the sample standard deviation of the monitoring variable to obtain the preprocessed dataset. The preprocessed dataset is segmented and combined according to a unified time window. Video image data, temperature data, pressure data, vibration data, gas concentration data, electrical operation data, and work record data within the same time window are correlated and uniformly encapsulated to generate a multimodal safety sample library.

[0010] Optionally, obtaining the candidate security risk attribute library includes: The attribute data and corresponding risk labels of all samples are extracted from the multimodal safety sample library. A cross-modal information flow coupling graph is constructed with each candidate attribute as a node. The conditional mutual information between each candidate attribute and the joint mutual information between each candidate attribute and the risk label are calculated. The mutual information gain value is calculated based on the conditional mutual information and joint mutual information corresponding to each candidate attribute. Candidate attributes are introduced one by one in descending order of mutual information gain value. After each candidate attribute is introduced, the mutual information gain value of the remaining candidate attributes is recalculated. During the process of introducing candidate attributes one by one, the change in mutual information gain before and after two consecutive introductions is continuously recorded. When the change in mutual information gain after two consecutive introductions is less than the change in mutual information gain after the previous introduction, the introduction is stopped, and all introduced candidate attributes are combined into a candidate security risk attribute library.

[0011] Optionally, obtaining the labeled training sample set and the sample set to be judged includes: Read the candidate security risk attribute values, timestamp information and risk label information corresponding to all samples from the candidate security risk attribute library, project each sample onto a unified temporal coordinate system in chronological order, and use the value of each sample on each candidate security risk attribute as a spatial coordinate component to generate a risk trajectory matrix corresponding to the time sequence. In the risk trajectory matrix, the sample with the highest risk label confidence is selected as the anchor sample. Starting from each anchor sample, convection-diffusion iterative labeling is performed along the time continuity direction and the attribute adjacency direction. Unlabeled samples that maintain a continuous connection with the anchor sample are assigned the risk label of the corresponding anchor sample layer by layer. After each round of expansion, the iteration layer number corresponding to the labeled sample is recorded until all samples are labeled. All labeled samples are divided according to the iteration layer number. Samples with odd layer numbers are assigned to the labeled training sample set, and samples with even layer numbers are assigned to the undecided sample set.

[0012] Optionally, the obtained set of key security risk attributes includes: Read the candidate safety risk attribute values ​​and corresponding risk labels of each sample in the labeled training sample set, construct an initial neighborhood set according to the attribute distance relationship between samples, and build a neighborhood association graph with samples as nodes and neighborhood association relationships as edges to form the initial neighborhood structure; The high-order neighborhood spectrum segmentation iteration is performed on the initial neighborhood structure, the neighborhood association graph is spectral decomposed, the neighborhood association edges are segmented and updated according to the spectral decomposition results, unstable neighborhood connections are gradually stripped away, and the sample neighborhood relations are reconstructed after each segmentation to obtain the purified neighborhood structure after spectral segmentation. Based on the clean neighborhood structure, the information potential field is reduced by contraction. An information potential field is formed according to the local distribution state of each sample in the clean neighborhood structure. The neighborhood boundary of each sample is adjusted along the contraction direction of the information potential field to obtain the contracted neighborhood structure. The ability of each candidate attribute to distinguish risk labels is calculated based on the shrinking neighborhood structure. Redundant attributes are deleted and valid attributes are retained. Neighborhood boundary adjustment and attribute screening are repeated until the set of remaining attributes is stable, thus obtaining the set of key safety risk attributes.

[0013] Optionally, the obtained optimized security risk prototype library includes: An improved learning vector quantization network is constructed based on a set of key safety risk attributes. An initial risk prototype set is generated according to the risk category. Entropy positive mass flow balance rearrangement is set in the prototype distribution layer of the learning vector quantization network. Curvature equilibrium arc surface correction is set in the category boundary adjustment layer. Hypergraph entropy cointegration optimization is performed in the prototype association and collaboration layer to form an improved learning vector quantization network. The initial risk prototype set is processed by entropy positive mass flow balancing rearrangement. The position of each risk prototype in the attribute space is adjusted according to the overall distribution relationship between each training sample and each risk prototype, resulting in a rearranged risk prototype set. The rearranged risk prototype set is processed by curvature equalization arc surface correction. Based on the category boundary structure formed between risk prototypes of the same category and risk prototypes of different categories, continuous trimming is performed on the boundary curvature region and the positional relationship of adjacent risk prototypes is adjusted simultaneously to obtain the boundary trimming risk prototype set. By using hypergraph entropy cointegration optimization to process the risk prototype set of boundary trimming, the correlation between training samples, key safety risk attributes and risk prototypes is uniformly coordinated and organized, and correlation optimization is performed on all risk prototypes to obtain the optimized risk prototype set. The labeled training sample set is input into the improved learning vector quantization network. The optimized risk prototype set is used as the initial training basis. The position and category boundary of each risk prototype are iteratively updated according to the category affiliation relationship between the sample and the risk prototype until the position change of each risk prototype meets the stopping condition. The optimized safety risk prototype library is then output.

[0014] Optionally, the real-time security risk assessment result table includes: Read the key safety risk attribute values ​​corresponding to each sample in the sample set to be judged, and read the prototype coordinates and prototype categories corresponding to each risk prototype in the optimized safety risk prototype library. Match each sample to be judged with all risk prototypes to obtain the prototype distance sequence corresponding to each sample to be judged. Based on the prototype distance sequence corresponding to each sample to be judged, the distance values ​​of each risk prototype are continuously interpolated and expanded in the local neighborhood of the sample to be judged according to the spatial continuous distribution relationship, generating a continuous distance surface corresponding to each sample to be judged. Local extremum search and curvature joint analysis are performed on the continuous distance surface to extract local minimum points, local maximum points and corresponding curvature change regions in the continuous distance surface. The risk prototype category corresponding to the local minimum point with the greatest attraction intensity is determined, and the risk prototype category is used as the risk category of the corresponding sample to be judged. Based on the risk category corresponding to each sample to be judged, the curvature change state of the continuous distance surface, and the distance change relationship between the risk prototype of the same category and the nearest risk prototype of a different category, the risk level corresponding to each sample to be judged is determined. The sample identifier, risk category, risk level and judgment time are correlated to generate a real-time safety risk judgment result table.

[0015] Optionally, generating the corresponding security risk classification early warning instruction includes: Each sample to be judged is classified and categorized according to its risk level. Samples to be judged that belong to the same risk category and the same risk level at the same judgment time are aggregated to generate a graded risk event set. Based on the number of samples to be judged, the corresponding risk category, the corresponding risk level, and the judgment time in each graded risk event set, a safety risk graded early warning instruction corresponding to each graded risk event set is generated. The event number, risk category identifier, risk level identifier, occurrence time, and sample source identifier are written into each safety risk graded early warning instruction. The warning instructions for each level of safety risk are sent to the monitoring and alarm platform in the order of the judgment time. The monitoring and alarm platform displays, stores and outputs alarms for each level of safety risk warning instruction.

[0016] The beneficial effects of this invention are: This invention achieves comprehensive perception of industrial production safety risks across multiple dimensions, including personnel behavior, equipment status, environmental changes, and operational processes, by uniformly preprocessing multi-source heterogeneous monitoring data such as video image data, temperature data, pressure data, vibration data, gas concentration data, electrical operation data, and work record data, coupled with cross-modal information flow analysis and candidate safety risk attribute screening. Compared to existing early warning methods that mainly rely on a single data source, fixed rules, or simple experience-based judgments, this invention utilizes cross-modal information flow coupling, iterative labeling of risk trajectory lattices, and neighborhood rough set attribute reduction methods to effectively extract and compress high-dimensional risk information under complex operating conditions. This not only improves the accuracy of extracting key safety risk attributes but also enhances the ability to identify complex, hidden, and progressive risks, effectively improving the comprehensiveness, stability, and adaptability of industrial production safety risk perception to changes in complex operating conditions.

[0017] This invention constructs an improved learning vector quantization network and combines it with entropy-positive mass flow balancing rearrangement, curvature equilibrium arc surface correction, and hyperspectral entropy cointegration optimization to uniformly optimize the risk prototype distribution, category boundary morphology, and higher-order correlation relationships, enabling more accurate risk category determination and risk level classification. Based on the real-time safety risk determination result table obtained from continuous distance surfaces and local extrema-curvature joint analysis, corresponding safety risk classification early warning instructions can be further generated and sent to the monitoring and alarm platform, achieving efficient connection between risk identification, risk classification, and early warning output. Compared with existing technologies, this invention has advantages such as low false alarm and false negative rates, strong adaptability to complex scenarios, high real-time and reliable early warning results, outstanding risk evolution identification capabilities, and ease of engineering deployment and field application promotion. Attached Figure Description

[0018] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of an industrial production safety risk early warning method based on multimodal fusion proposed in this invention; Figure 2 This is a structural block diagram of the neighborhood rough set attribute reduction of an industrial production safety risk early warning method based on multimodal fusion proposed in this invention; Figure 3 This is a functional diagram of the improved learning vector quantization network for an industrial production safety risk early warning method based on multimodal fusion proposed in this invention. Detailed Implementation

[0019] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0020] refer to Figure 1 , Figure 2 and Figure 3 A method for early warning of industrial production safety risks based on multimodal fusion, comprising: Collect multi-source heterogeneous monitoring data from industrial production sites, preprocess the multi-source heterogeneous monitoring data, and obtain a multimodal safety sample library; A cross-modal information flow coupling graph is constructed in a multimodal security sample library, and a candidate security risk attribute library is obtained based on the change of mutual information gain; Based on the candidate safety risk attribute library, risk trajectory points are mapped and iteratively labeled according to the convection-diffusion rule to obtain the labeled training sample set and the sample set to be judged; Based on neighborhood rough set attribute reduction, neighborhood relations are reconstructed and attributes are filtered in the labeled training sample set. High-order neighborhood spectrum segmentation iteration is introduced to perform spectrum decomposition, segmentation and update. Information potential energy field contraction reduction is used to perform boundary contraction, attribute compression and redundancy removal to obtain the key security risk attribute set. An initial risk prototype set is generated by the set of key safety risk attributes. An improved learning vector quantization network is constructed. The initial risk prototype set is then subjected to prototype rearrangement, boundary trimming and correlation optimization. Entropy positive mass flow balance rearrangement is introduced for global adjustment. Curvature balance processing is performed based on curvature equilibrium arc surface correction. Hyperspectral entropy cointegration optimization is used to uniformly optimize the risk prototypes, resulting in an optimized safety risk prototype library. The sample set to be judged is matched with the optimized safety risk prototype library, the continuous distance surface of the risk prototype is calculated, and the local extremum-curvature joint analysis is performed to obtain the real-time safety risk judgment result table. Based on the real-time security risk assessment results table, generate corresponding security risk classification early warning instructions and send the security risk classification early warning instructions to the monitoring and alarm platform.

[0021] In this embodiment, the multi-source heterogeneous monitoring data includes video image data, temperature data, pressure data, vibration data, gas concentration data, electrical operation data, and work record data.

[0022] In this embodiment, the preprocessing of multi-source heterogeneous monitoring data includes time alignment, outlier removal, missing value completion, noise reduction filtering, standardization, and structured coding.

[0023] In this embodiment, obtaining the multimodal security sample library includes: Video image data, temperature data, pressure data, vibration data, gas concentration data, electrical operation data, and work record data are collected separately. Timestamps are added to each type of data according to a unified sampling clock to obtain the original monitoring dataset. The original monitoring dataset undergoes time alignment, outlier removal, missing value completion, denoising filtering, standardization, and structured coding. Standardization involves subtracting the sample mean of each monitoring variable from its actual value, then dividing by the sample standard deviation to obtain the preprocessed dataset. The structured coding specifically involves: Data of various types is aggregated in a unified time window. Continuous numerical data is directly used as the corresponding field value. Video image data is extracted into target category, target quantity, target location coordinates, and target area percentage. The target area percentage is calculated by dividing the number of pixels in the target area by the total number of pixels in the image. The average value of the corresponding fields of multiple frames in the same time window is taken. The operation record and alarm log are converted into integer codes according to the preset event coding table. The occurrence frequency of each event in the current time window is counted. Status data is binary encoded with 1 for occurrence and 0 for non-occurrence. Multi-category discrete data is encoded in a one-hot manner. The data is concatenated into a fixed-length feature vector according to the field order. When a field is missing, continuous numerical fields are filled with 0, status fields are filled with 0, event count fields are filled with 0, and category coding fields are all set to 0 to obtain the structured coding result. The preprocessed dataset is segmented and combined according to a unified time window. Video image data, temperature data, pressure data, vibration data, gas concentration data, electrical operation data, and work record data within the same time window are correlated and uniformly encapsulated to generate a multimodal safety sample library.

[0024] In this embodiment, obtaining the candidate security risk attribute library includes: Attribute data and corresponding risk labels of all samples are extracted from the multimodal safety sample library. A cross-modal information flow coupling graph is constructed using each candidate attribute as a node. The conditional mutual information between each candidate attribute and the joint mutual information between each candidate attribute and the risk label are calculated, where: A cross-modal information flow coupling graph is constructed using each candidate attribute as a node. Specifically, candidate attribute values ​​of all samples are read from the multimodal security sample library. An attribute node set is established with one candidate attribute corresponding to one node. Any two candidate attributes are traversed pairwise. If the candidate attributes come from different modalities, a connection edge is directly established between the corresponding nodes. If the candidate attributes come from the same modality, it is further determined whether there is a correlation between them and other modal candidate attributes. If there is, a connection edge is established between the nodes. After the connection edge is established, the connection edge is assigned a value according to the synchronous change of the candidate attributes in all samples and the degree of correlation under the risk label condition. This results in a cross-modal information flow coupling graph where nodes represent candidate attributes, connection edges represent cross-modal coupling relationships between candidate attributes, and edge values ​​represent the strength of coupling. The conditional mutual information between candidate attributes and the joint mutual information between candidate attributes and risk labels are calculated as follows: The number of times any two candidate attributes co-occur under different risk labels is counted; the number of times each candidate attribute appears under each risk label is counted; the information correlation between candidate attributes is calculated under each risk label; and the information correlation under each label is summed according to the proportion of each risk label in the entire sample to obtain the conditional mutual information between candidate attributes. The larger the conditional mutual information value, the stronger the correlation between candidate attributes even when the risk label is known. When calculating the joint mutual information between each candidate attribute and risk label, the number of times each candidate attribute and each risk label appear in combination is counted first; then the number of times each candidate attribute appears alone and each risk label appears alone are counted. The information contribution is accumulated item by item according to the degree of difference between the joint occurrence and the individual occurrence to obtain the joint mutual information between candidate attributes and risk labels. The mutual information gain value is calculated based on the conditional mutual information and joint mutual information corresponding to each candidate attribute. Candidate attributes are then introduced one by one in descending order of mutual information gain value. After each candidate attribute is introduced, the mutual information gain value is recalculated for the remaining candidate attributes. Specifically, the calculation of the mutual information gain value, involving the introduction of candidate attributes one by one in descending order of mutual information gain value, is as follows: For each candidate attribute, read the joint mutual information between it and the risk label, then read the conditional mutual information between the candidate attribute and all candidate attributes. Use the joint mutual information between the candidate attribute and the risk label as the effective information, and the average of all conditional mutual information between the candidate attributes as the redundant information. Subtract the redundant information from the effective information to obtain the mutual information gain value of the candidate attribute. After calculating the mutual information gain value of all candidate attributes, sort them from largest to smallest according to the mutual information gain value. Take the first ranked candidate attribute and introduce it into the attribute set first. After introducing each candidate attribute, delete the introduced attribute from the remaining candidate attributes and recalculate the average conditional mutual information between each remaining candidate attribute and the currently introduced attribute set. Subtract the updated average conditional mutual information from the original joint mutual information of the candidate attribute to obtain the new mutual information gain value of the candidate attribute. Sort them from largest to smallest according to the updated mutual information gain value and introduce the first ranked candidate attribute. Repeat the process until the introduction process is completed. During the process of introducing candidate attributes one by one, the change in mutual information gain before and after two consecutive introductions is continuously recorded. When the change in mutual information gain after two consecutive introductions is less than the change in mutual information gain after the previous introduction, the introduction is stopped, and all introduced candidate attributes are combined into a candidate security risk attribute library.

[0025] In this embodiment, obtaining the labeled training sample set and the sample set to be judged includes: The candidate security risk attribute values, timestamp information, and risk label information corresponding to all samples are read from the candidate security risk attribute library. Each sample is projected onto a unified temporal coordinate system in chronological order. Using the value of each sample on each candidate security risk attribute as a spatial coordinate component, a risk trajectory point matrix corresponding to the time sequence is generated. Specifically, the generation of the risk trajectory point matrix corresponding to the time sequence is as follows: Read the timestamp information of all samples and sort them in ascending order of timestamp. Use the sorted positions as discrete numbers on the time axis. Read all candidate safety risk attribute values ​​corresponding to each sample and use each attribute value as a spatial coordinate component in the order of attribute arrangement. Make each sample correspond to a trajectory point composed of time coordinates and attribute coordinates. Connect the trajectory points of adjacent samples in the sorted time order to form the risk change trajectory of the sample over time. Arrange the trajectory points of all samples according to the time sequence number and retain the risk label information corresponding to each trajectory point to obtain a risk trajectory point matrix consistent with the time order. Each row in the point matrix corresponds to the sample trajectory point under the time sequence number, and each column corresponds to the candidate safety risk attribute dimension. In the risk trajectory matrix, the sample with the highest risk label confidence is selected as the anchor sample. Starting from each anchor sample, convection-diffusion iterative labeling is performed along the time continuity direction and the attribute adjacency direction. Unlabeled samples that maintain a continuous connection with the anchor sample are assigned the corresponding risk label layer by layer. After each round of expansion, the iteration layer number corresponding to the labeled sample is recorded until all samples are labeled. The convection-diffusion iterative labeling process is as follows: The sample with the highest risk label confidence is selected from the risk trajectory matrix as the anchor sample. The location of the anchor sample is recorded as the first layer. With the anchor sample as the center, the sample is expanded outward in two directions. The samples at time points adjacent to the current sample are selected along the time continuity direction. The unlabeled samples with the smallest attribute distance to the current sample are selected along the attribute adjacency direction. For each candidate unlabeled sample, the time distance between it and the current labeled sample is calculated first, and then the attribute Euclidean distance is calculated. The time distance and attribute distance are combined into a comprehensive diffusion distance. The unlabeled sample with the smallest comprehensive diffusion distance and a continuous connection with the current labeled sample is assigned a risk label to the current anchor sample. After each outward expansion, all samples with newly assigned labels in this round are recorded as the next layer. The expansion continues with the current layer as the new starting point until all samples in the risk trajectory matrix have obtained risk labels and corresponding iteration layer numbers. All labeled samples are divided according to the iteration layer number. Samples with odd layer numbers are assigned to the labeled training sample set, and samples with even layer numbers are assigned to the undecided sample set.

[0026] In this embodiment, obtaining the set of key security risk attributes includes: Read the candidate safety risk attribute values ​​and corresponding risk labels of each sample in the labeled training sample set, construct an initial neighborhood set according to the attribute distance relationship between samples, and build a neighborhood association graph with samples as nodes and neighborhood association relationships as edges to form the initial neighborhood structure. The formation of the initial neighborhood structure is specifically as follows: Read the candidate safety risk attribute values ​​of all samples in the labeled training sample set, represent each sample as an attribute vector, calculate the attribute distance between any two samples, and obtain the attribute distance by taking the square root of the sum of the squares of the differences between the candidate safety risk attributes. For each sample, sort it in ascending order of attribute distance with all samples, select the top 5 samples with the smallest distance as the initial neighborhood samples of the sample, take each sample as a node, establish neighborhood association edges between the sample and the initial neighborhood samples, and use the reciprocal of the attribute distance between samples as the corresponding edge value, so that the smaller the distance between samples, the stronger the neighborhood association. After traversing all samples, the initial neighborhood structure composed of sample nodes, neighborhood association edges and edge values ​​is obtained. A high-order neighborhood spectral segmentation iteration is performed on the initial neighborhood structure. Spectral decomposition is performed on the neighborhood association graph. Based on the spectral decomposition results, neighborhood association edges are segmented and updated, gradually stripping away unstable neighborhood connections. After each segmentation, the sample neighborhood relations are reconstructed to obtain the purified neighborhood structure after spectral segmentation, where: The initial neighborhood structure is subjected to a high-order neighborhood spectral segmentation iteration. The neighborhood association graph is decomposed spectrally, specifically as follows: a neighborhood association matrix is ​​constructed based on the neighborhood association edges and corresponding edge values ​​between all sample nodes. The sum of edge values ​​for each sample node is calculated to form the node degree value. The node degree value is subtracted from the neighborhood association matrix to obtain the graph Laplacian matrix. The graph Laplacian matrix is ​​decomposed into eigenvalues ​​to obtain multiple eigenvalues ​​and their corresponding eigenvectors sorted by size. The eigenvectors corresponding to the three smallest non-zero eigenvalues ​​are selected as high-order spectral features. The values ​​of each sample node on the three eigenvectors are combined to form new spectral coordinates. The spectral distance between any two sample nodes is recalculated in the spectral coordinate space. The relationships between nodes are rearranged according to the spectral distance from smallest to largest. After completing one round of spectral decomposition, the updated spectral distance result is used as the input for the next round of segmentation. This process is repeated until the spectral distance sorting results of two consecutive rounds are consistent, thus completing the high-order neighborhood spectral segmentation iteration. Based on the spectral decomposition results, the neighborhood association edges are segmented and updated. Specifically, the spectral coordinates of the sample nodes at both ends of each neighborhood association edge are read, and the spectral distance between the two nodes is calculated. If the risk labels of the two nodes at both ends of the neighborhood association edge are different, and the spectral distance of the edge is greater than 1.2 times the average spectral distance of all neighborhood association edges, the neighborhood association edge is deleted. If the risk labels of the two nodes at both ends of the neighborhood association edge are the same, and the spectral distance of the edge is less than 0.8 times the average spectral distance of all neighborhood association edges, the neighborhood association edge is retained, and the edge value is updated to the product of the original edge value and the inverse of the spectral distance. For the remaining neighborhood association edges, the connection relationship is retained, but the edge value is multiplied by 0.5 and updated by decay. After traversing all neighborhood association edges, the top 5 associated nodes are reselected as new neighborhood samples for each sample node according to the updated edge value from largest to smallest. The sample neighborhood relationship is reconstructed to obtain the purified neighborhood structure after spectral segmentation. Based on the cleaned neighborhood structure, information potential field contraction reduction is performed. An information potential field is formed according to the local distribution state of each sample within the cleaned neighborhood structure. The neighborhood boundaries of each sample are adjusted along the contraction direction of the information potential field to obtain the contracted neighborhood structure. Specifically, the information potential field contraction reduction based on the cleaned neighborhood structure is performed as follows: For each sample, the number of neighboring samples, the sum of neighboring edge values, and the average attribute distance to the neighborhood center are counted in the purified neighborhood structure. The information aggregation intensity of the sample is obtained by multiplying the number of neighboring samples by the sum of neighboring edge values, dividing by the average attribute distance, and adding 0.001. The negative number of the information aggregation intensity is used as the information potential energy value of the sample. The information potential energy difference between each sample and each neighboring sample is compared. The direction in which the information potential energy decreases the fastest is selected as the contraction direction. Neighboring connection edges that are opposite to the contraction direction and whose edge values ​​are lower than the average value of all neighboring edges of the sample are deleted. After completing one round of boundary contraction, the neighborhood center, average attribute distance, and information aggregation intensity are recalculated for the remaining neighboring samples. The next round of contraction is performed according to the updated results. If the set of neighboring samples remaining after two consecutive rounds of contraction does not change, the boundary contraction of the sample is stopped. When all samples meet the stopping condition, the contracted neighborhood structure is obtained. Based on the shrinking neighborhood structure, the distinguishing ability of each candidate attribute for risk labels is calculated. Redundant attributes are removed and valid attributes are retained. The neighborhood boundary adjustment and attribute screening are repeated until the set of remaining attributes is stable, resulting in the set of key safety risk attributes. The distinguishing ability of each candidate attribute for risk labels is calculated as follows: For each candidate attribute, the intra-class dispersion among samples with the same risk label and the inter-class separation among samples with different risk labels are statistically analyzed. The intra-class dispersion is obtained by calculating the average absolute deviation between the attribute value and the label mean under the same risk label. The inter-class separation is obtained by calculating the pairwise absolute differences between the corresponding means of different risk labels and taking the average. The initial discrimination value of the candidate attribute is obtained by dividing the inter-class separation value by the intra-class dispersion value and adding 0.001. Combined with the shrinking neighborhood structure, the proportion of samples with the same label that are clustered and samples with different labels that are separated in each sample neighborhood is statistically analyzed. The proportion is multiplied by the initial discrimination value to obtain the final discrimination power value of the candidate attribute. The larger the discrimination power value, the stronger the discrimination effect of the candidate attribute on the risk label.

[0027] In this embodiment, obtaining the optimized security risk prototype library includes: An improved learning vector quantization network is constructed based on a set of key safety risk attributes. An initial risk prototype set is generated according to risk categories. Entropy-positive mass flow balancing rearrangement is implemented in the prototype distribution layer of the learning vector quantization network, curvature equilibrium arc surface correction is implemented in the category boundary adjustment layer, and hypergraph entropy cointegration optimization is performed in the prototype association and coordination layer, thus forming the improved learning vector quantization network. Specifically, the initial risk prototype set is generated according to risk categories as follows: The labeled training sample set is grouped according to risk category. For each risk category, the attribute values ​​of all samples on the key safety risk attribute set are read. The sample mean of each key safety risk attribute under the risk category is calculated to form the central attribute vector of the risk category. The attribute distance between each sample and the central attribute vector is calculated and sorted in ascending order of distance. The top 3 samples with the smallest distance are selected and the arithmetic mean of the attribute vectors corresponding to the 3 samples is taken as the initial risk prototype vector of the risk category. If the number of samples in a risk category is less than 3, the average of the attribute vectors of all samples in the category is directly used as the initial risk prototype vector. After repeating the process for all risk categories, the initial risk prototype set composed of the initial risk prototype vectors of each risk category is obtained. The initial risk prototype set is processed by entropy positive mass flow balancing rearrangement. Based on the overall distribution relationship between each training sample and each risk prototype, the position of each risk prototype in the attribute space is adjusted to obtain the rearranged risk prototype set. Specifically, the rearranged risk prototype set is as follows: Calculate the attribute distance between each training sample and all risk prototypes, inverse the distance and exponentialize it, normalize all exponentialized results for the same sample to obtain the quality allocation ratio of the sample to each risk prototype, summarize the quality of all training samples allocated to the risk prototype for each risk prototype, multiply the attribute vector of each sample by the corresponding quality allocation ratio and sum them, then divide by the total quality received by the risk prototype to obtain the new position of the risk prototype. After completing one round of updating the positions of all risk prototypes, recalculate the position distance of the training sample to the new risk prototype and repeat the process until the change in the position of each risk prototype is less than 0.001 in two consecutive rounds, and obtain the rearranged risk prototype set. The rearranged risk prototype set is processed by curvature equalization arc surface correction. Based on the category boundary structure formed between risk prototypes of the same category and risk prototypes of different categories, continuous trimming is performed on the boundary curvature region, and the positional relationship of adjacent risk prototypes is adjusted synchronously to obtain the boundary trimmed risk prototype set. Specifically, the boundary trimmed risk prototype set is as follows: For each risk prototype, find the two closest risk prototypes of the same category and the two closest risk prototypes of different categories. The risk prototype and its adjacent risk prototypes together form a local boundary segment. Calculate the rate of change of the angle between adjacent lines in the local boundary segment. Divide the rate of change of the angle by the average length of the corresponding line and add 0.001 to obtain the local boundary curvature value of the risk prototype's location. When the local boundary curvature value is greater than 1.1 times the average local boundary curvature of all risk prototypes, move the risk prototype along the local center direction formed by the risk prototypes of the same category, and simultaneously adjust it along the direction away from the nearest risk prototype of different category. The movement amount is 0.5 times the difference between the current local boundary curvature value and the average local boundary curvature of all risk prototypes. After completing one round of position adjustment for all risk prototypes, recalculate the adjacency relationship and local boundary curvature value between each risk prototype, and repeat the adjustment process until the average change of local boundary curvature of all risk prototypes in two consecutive rounds is less than 0.001, thus obtaining the boundary-adjusted risk prototype set. The risk prototype set for boundary trimming is processed by hypergraph entropy cointegration optimization. The correlation between training samples, key safety risk attributes, and risk prototypes is uniformly coordinated and organized. Correlation optimization is then performed on all risk prototypes to obtain an optimized risk prototype set, which is specifically as follows: Using training samples, key security risk attributes, and boundary trimming risk prototypes as three types of nodes, a hyperedge is established for each training sample, along with the top five key security risk attributes with the largest corresponding attribute values ​​and the nearest risk prototype. The association frequency, associated attribute count, and associated sample count of each risk prototype across all hyperedges are counted. The hypergraph association entropy of each risk prototype is calculated. First, the association ratio between the risk prototype and each hyperedge is calculated. Then, the information dispersion of all association ratios is calculated and accumulated to obtain the hypergraph association entropy value of the risk prototype. If the hypergraph association entropy value of a risk prototype is greater than the hypergraph association entropy value of all risk prototypes, the risk prototype is considered a higher risk prototype. If the entropy value is 1.1 times the average value, the risk prototype is moved towards the center of the sample with the most associations and towards the attribute center formed by the high-frequency association attributes. The amount of movement in each direction is 0.5 times the average distance in the corresponding direction. If the hypergraph association entropy value of the risk prototype is less than 0.9 times the average hypergraph association entropy value of all risk prototypes, the position of the risk prototype remains unchanged. After completing one round of updating all risk prototypes, the hyperedges, association counts, and hypergraph association entropy values ​​are recalculated and optimization continues until the average change in the position of all risk prototypes in two consecutive rounds is less than 0.001, thus obtaining the optimized risk prototype set. The labeled training sample set is input into the improved learning vector quantization network, using the optimized risk prototype set as the initial training basis. The position and category boundary of each risk prototype are iteratively updated according to the category affiliation relationship between the samples and the risk prototypes until the position changes of each risk prototype satisfy the stopping condition. The optimized safety risk prototype library is then output, specifically as follows: The labeled training sample set is input one by one into the improved learning vector quantization network. For each training sample, the attribute distance between it and all risk prototypes is calculated. The risk prototype with the smallest distance is selected as the current matching prototype. If the risk category of the current matching prototype is consistent with the risk label of the training sample, the risk prototype is moved along the direction of the training sample by 0.1 times the attribute difference between the current risk prototype and the training sample. If the risk categories are inconsistent, the risk prototype is moved away from the training sample by 0.1 times the attribute difference between the current risk prototype and the training sample. After one round of updating all training samples, the boundary positions between adjacent risk prototypes of the same category and risk prototypes of different categories are recalculated based on the updated risk prototypes. The updated risk prototype coordinates, risk category labels, and boundary adjacency relationships are saved as the training result of one round. The training process is repeated until the average change in the position of all risk prototypes in two consecutive rounds is less than 0.001. The coordinates, risk categories, key safety risk attribute dimensions, and category boundary relationships of all risk prototypes are stored uniformly, and the optimized safety risk prototype library is output.

[0028] In this embodiment, obtaining the real-time security risk assessment result table includes: Read the key safety risk attribute values ​​corresponding to each sample in the sample set to be judged, and read the prototype coordinates and prototype category corresponding to each risk prototype in the optimized safety risk prototype library. Match each sample to be judged with all risk prototypes to obtain the prototype distance sequence corresponding to each sample to be judged. Specifically, the prototype distance sequence corresponding to each sample to be judged is as follows: For each sample to be judged, all key safety risk attribute values ​​are read, and sample attribute vectors are formed according to the attribute order consistent with the prototype coordinates in the optimized safety risk prototype library. The distance between the sample attribute vector and the prototype coordinates of all risk prototypes is calculated one by one. The distance value is obtained by summing the squared differences of each attribute and then taking the square root. After the distance calculation between the sample to be judged and all risk prototypes is completed, all the obtained distance values ​​are stored in the order of the risk prototypes in the optimized safety risk prototype library to form the prototype distance sequence corresponding to the sample to be judged. After repeating the process for all samples in the sample set to be judged, the prototype distance sequence corresponding to each sample to be judged is obtained. Based on the prototype distance sequence corresponding to each sample to be judged, the distance values ​​of each risk prototype are continuously interpolated and expanded according to the spatial continuous distribution relationship in the local neighborhood of the sample to be judged, generating a continuous distance surface corresponding to each sample to be judged. Specifically, the generation of the continuous distance surface corresponding to each sample to be judged is as follows: Centered on the sample to be judged, the five samples with the smallest attribute distance in the sample set to be judged are selected as local neighborhood samples. The prototype distance sequence corresponding to the sample to be judged and the local neighborhood samples is read. The first two dominant attribute values ​​of each sample in the critical safety risk attribute space in the local neighborhood are used as plane coordinates. The distance value corresponding to each risk prototype is used as the height value. A discrete distance point set is constructed for each risk prototype. Continuous interpolation is performed on the discrete distance point set according to the distance reciprocal weighting method. For any point in the plane, the Euclidean distance between the current point and all discrete distance points is calculated. The distance value of each discrete distance point is multiplied by the corresponding distance reciprocal and summed. Then, it is divided by the sum of all distance reciprocals to obtain the interpolation height. After repeating the calculation for all interpolation points in the local neighborhood plane, a continuous distance surface corresponding to the risk prototype is formed. The continuous distance surfaces corresponding to all risk prototypes are uniformly unfolded to obtain the continuous distance surface corresponding to the sample to be judged. Local extremum search and curvature joint analysis are performed on the continuous distance surface to extract local minima, local maxima, and corresponding curvature variation regions. The risk prototype category corresponding to the local minima with the highest attraction intensity is determined, and this risk prototype category is used as the risk category for the corresponding sample to be judged. Specifically, the local extremum search and curvature joint analysis are performed as follows: The height value of each grid point is read point by point on the discrete grid of the continuous distance surface. The grid point is compared with the height values ​​of the eight surrounding grid points. When the height value of the grid point is less than the height values ​​of the eight surrounding grid points, the grid point is recorded as a local minimum point. When the height value of the grid point is greater than the height values ​​of the eight surrounding grid points, the grid point is recorded as a local maximum point. The second-order difference values ​​of the location in the two planar directions are calculated for each local minimum point and local maximum point. The second-order difference values ​​in the two directions are added together to obtain the curvature value of the current point. The corresponding curvature change region is determined according to the continuous change range of the curvature value in the adjacent grid. The attraction intensity is calculated for all local minimum points. The attraction intensity is the result of subtracting the height value of the local minimum point from the average height value of the eight surrounding grid points. The risk prototype category corresponding to the local minimum point with the largest attraction intensity is selected as the risk category of the sample to be judged. Based on the risk category corresponding to each sample to be judged, the curvature change state of the continuous distance surface, and the distance change relationship between the risk prototype of the same category and the nearest risk prototype of a different category, the risk level corresponding to each sample to be judged is determined. The sample identifier, risk category, risk level, and judgment time are correlated to generate a real-time safety risk judgment result table. The specific determination of the risk level corresponding to each sample to be judged is as follows: Read the identified risk categories of the samples to be judged, extract the curvature values ​​of the continuous distance surface near the corresponding local minimum points, calculate the distance difference between the risk prototype of the same category to which the sample belongs and the nearest risk prototype of a different category, and use the curvature value and the distance difference together as the basis for risk level determination. The larger the curvature value, the more drastic the boundary change of the region where the sample is located; the smaller the distance difference, the closer the sample is to the boundary of the risk of a different category. Calculate the comprehensive risk value for all samples to be judged. The comprehensive risk value is the result of dividing the curvature value by the distance difference and adding 0.001. Sort the samples according to the comprehensive risk value from largest to smallest. The samples with the comprehensive risk value in the top 20% are judged as high-risk, those in the top 20% to 50% are judged as medium-risk, and those in the bottom 50% are judged as low-risk. This gives the risk level corresponding to each sample to be judged.

[0029] In this embodiment, generating the corresponding security risk classification early warning instruction includes: The samples to be judged are classified and categorized according to their risk levels. Samples belonging to the same risk category and risk level at the same judgment time are aggregated to generate a graded risk event set. The generation of the graded risk event set is as follows: Read all pending sample records from the real-time security risk assessment result table, sort them in ascending order of assessment time, and group them one by one using assessment time, risk category, and risk level as three aggregation keys. Pending samples with the same assessment time, risk category, and risk level are grouped into the same event group. After each event group is formed, count the number of pending samples in the event group and use the number as the event size value of the event group. Generate a unique event number for each event group and bind the event number, corresponding assessment time, risk category, risk level, and event size value. After completing the grouping of all pending samples, a hierarchical risk event set consisting of multiple event groups is obtained. Based on the number of samples to be judged, the corresponding risk category, the corresponding risk level, and the judgment time in each graded risk event set, a safety risk graded early warning instruction corresponding to each graded risk event set is generated. Each safety risk graded early warning instruction includes the event number, risk category identifier, risk level identifier, occurrence time, and sample source identifier. Specifically, the generation of the safety risk graded early warning instruction corresponding to each graded risk event set is as follows: First, read the number of samples to be judged, risk category, risk level, and judgment time for each graded risk event set. Determine the warning priority based on the risk level: high risk level corresponds to Level 1 warning, medium risk level corresponds to Level 2 warning, and low risk level corresponds to Level 3 warning. Multiply the number of samples to be judged by the event scale coefficient to obtain the warning intensity value of the graded risk event. The preset event scale coefficient is 0.75. Generate the corresponding warning instruction record according to the fixed field order of event number, risk category identifier, risk level identifier, occurrence time, sample source identifier, warning priority, and warning intensity value. The sample source identifier is obtained by combining the source identifiers of all samples to be judged in the graded risk event set in chronological order. After processing all graded risk event sets, obtain the safety risk graded warning instruction corresponding to each graded risk event set. The warning instructions for each level of safety risk are sent to the monitoring and alarm platform in the order of the judgment time. The monitoring and alarm platform displays, stores and outputs alarms for each level of safety risk warning instruction.

[0030] Example 1: To verify the feasibility of this invention in practice, it was applied to a continuous industrial production cycle. A batch of multi-source heterogeneous monitoring data from the production line monitoring terminal was received, including 4320 frames of video image data with a resolution of 1280×720 pixels. The current batch of data covers one complete continuous production cycle, forming 2640 basic sampling moments. In the raw data, the temperature range was 61.3-86.7, the pressure range was 1.08-1.46, the effective value of vibration ranged from 1.9-6.8, the gas concentration ranged from 3-38, and the current ranged from 36.4-53.2. Approximately 12.6% of the frames in the video data had steam obstruction or local reflection. The operation log included valve position switching, inspection stops, and manual operation confirmations. In the current scenario, there were early signs of slow leakage at the connection point. Although the single temperature or pressure signal did not exceed the traditional alarm line, multiple signals simultaneously exhibited weak abnormal coupling, which is consistent with the characteristics of the problems that this invention aims to solve: insufficient perception of a single data source, difficulty in identifying complex risks, and high false alarm and missed alarm rates.

[0031] After the data enters the processing flow, time alignment is first performed to unify all monitoring data into one sampling period, compressing the time drift from a maximum of 0.83 sampling periods to 0. After outlier removal, 27 temperature spikes, 19 pressure spikes, 31 vibration spikes, and 16 gas concentration spikes are deleted; the missing data completion rates are 0.42%, 0.37%, 0.55%, 0.28%, and 0.31%, respectively. After denoising filtering, the mean square fluctuation of the temperature sequence is reduced from 2.84 to 1.91, and the mean square fluctuation of the vibration sequence is reduced from 0.73 to 0.48. After standardization, the mean of each continuous variable is close to 0, and the standard deviation is close to 1. During structured coding, 10 consecutive sampling times are used as one time window, forming a total of 864 time window samples. Each sample contains 29-dimensional structured fields, including 7-dimensional video extraction fields, 15-dimensional environmental and equipment continuous fields, 4-dimensional status fields, and 3-dimensional event counting fields, ultimately generating 864 sets of multimodal safety sample libraries. Compared with traditional methods that only retain four types of data—temperature, pressure, vibration, and gas concentration—the present invention increases the sample information dimension to 29 dimensions.

[0032] In the candidate safety risk attribute extraction stage, a cross-modal information flow coupling graph was constructed using 29 candidate attributes as nodes, establishing a total of 246 connection edges, including 173 cross-modal connection edges. After conditional mutual information and joint mutual information calculations, the top 12 attributes with the highest mutual information gain were: local temperature rise slope (0.418), pressure fluctuation amplitude (0.387), vibration growth rate (0.365), gas concentration rise (0.352), current offset (0.336), image vapor area ratio (0.319), number of abnormal targets in the image (0.287), valve position switching count (0.266), inspection dwell time (0.251), temperature and pressure coupling deviation (0.244), vibration and current coupling deviation (0.231), and gas concentration change rate (0.219). Attributes were introduced one by one, from largest to smallest, based on mutual information gain. After the 13th attribute was introduced, the change in mutual information gain decreased from 0.021 to 0.008, and then further to 0.005, at which point the introduction was stopped, resulting in a 12-dimensional candidate security risk attribute library. Compared to the traditional manually selected six fixed features, this invention retains more coupled attributes sensitive to risk evolution.

[0033] In the sample annotation phase, 864 sets of samples were projected into a risk trajectory matrix, with each set of samples corresponding to one trajectory point. The 14 samples with the highest confidence levels were selected as anchor points from the trajectory matrix: 6 normal anchor points, 4 low-risk anchor points, 2 medium-risk anchor points, and 2 high-risk anchor points. After iterative annotation according to the convection-diffusion rule, the first round of annotation expanded 79 sets of samples, the second round expanded 146 sets, the third round expanded 208 sets, the fourth round expanded 243 sets, and the fifth round expanded 188 sets, completing the annotation of all samples. Subsequently, the samples were divided into odd and even layers, resulting in a training sample set of 432 sets and a judgment sample set of 432 sets. Among the training samples, there were 278 normal samples, 86 low-risk samples, 43 medium-risk samples, and 25 high-risk samples. This ensures the continuity of the risk evolution chain between training and judgment samples, which is superior to traditional random partitioning.

[0034] In the key security risk attribute screening stage, an initial neighborhood structure was first constructed, selecting the top 5 nearest neighbors for each sample to form 2160 initial neighborhood edges. After three rounds of iterative high-order neighborhood spectrum segmentation, 428 unstable neighborhood edges were removed, leaving 1732 cleaned neighborhood edges. The average spectral distance of samples with the same label decreased from 0.413 to 0.267, while the average spectral distance of samples with different labels increased from 0.521 to 0.708. Subsequently, information potential field contraction and reduction were performed. After four consecutive rounds of contraction, the average neighborhood size of the sample decreased from 5 to 3.2, three redundant attributes were removed, and finally, nine key security risk attributes were retained. Compared with before reduction, the inter-class separation improved by 31.4%, and the intra-class dispersion decreased by 22.7%.

[0035] In the risk prototype construction and optimization phase, four initial risk prototypes were generated according to four risk categories. After six rounds of entropy-positive mass flow balancing rearrangement, the average position change of the prototypes decreased from 0.184 to 0.007; after four rounds of curvature equilibrium arc surface correction, the average local boundary curvature decreased from 0.362 to 0.194; after five rounds of hyperspectral entropy cointegration optimization, the average prototype correlation entropy decreased from 1.283 to 0.914, ultimately outputting an optimized safety risk prototype library. Subsequently, 432 sets of samples to be judged were matched one by one with the four risk prototypes to obtain prototype distance sequences, and continuous distance surfaces were generated in the local neighborhood. Local extremum-curvature joint analysis showed that 61 sets of samples were judged as high-risk, 89 sets as medium-risk, 102 sets as low-risk, and the remaining 180 sets as normal. The system further aggregated and generated a set of 47 graded risk events, including 9 Level 1 warning events, 14 Level 2 warning events, and 24 Level 3 warning events. Of the Level 1 alerts output by the monitoring and alarm platform, 8 correspond to actual abnormal operating conditions, and the other is a false alarm event.

[0036] In the comparative experiment, the number of training samples and the number of samples to be judged were both 432. The traditional method used single-source threshold warning, fixed features, and ordinary learning vector quantization classification. The following results were obtained on the same batch of samples to be judged: high-risk sample detection rate: 71.2% for the traditional method, 94.8% for this invention; overall recognition accuracy: 81.6% for the traditional method, 95.6% for this invention; false alarm rate: 12.3% for the traditional method, 4.1% for this invention; false negative rate: 14.8% for the traditional method, 3.9% for this invention; average warning lead time: 3.1 time windows for the traditional method, 7.4 time windows for this invention; training time after key attribute dimension compression: 12.8 seconds for the traditional method, 9.6 seconds for this invention; level 1 warning effectiveness rate: 66.7% for the traditional method, 88.9% for this invention. In a typical risk evolution process, the traditional method only triggered an alarm when the temperature reached 82.4°C, the pressure reached 1.41 ppm, and the gas concentration reached 29 ppm. However, the present invention issued a Level 1 warning under the combined conditions of a temperature of 78.6°C, a pressure of 1.33 ppm, a gas concentration of 18 ppm, a current of 49.1 ppm, and a steam area ratio of 0.12 ppm, six time windows ahead of schedule. This demonstrates that the present invention exhibits a clear data change process in each processing step and improves the identification and graded warning capabilities for complex safety risks in complex industrial production environments, verifying the engineering feasibility and effectiveness of the method.

[0037] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for early warning of industrial production safety risks based on multimodal fusion, characterized in that, include: Collect multi-source heterogeneous monitoring data from industrial production sites, preprocess the multi-source heterogeneous monitoring data, and obtain a multimodal safety sample library; A cross-modal information flow coupling graph is constructed in a multimodal security sample library, and a candidate security risk attribute library is obtained based on the change of mutual information gain; Based on the candidate safety risk attribute library, risk trajectory points are mapped and iteratively labeled according to the convection-diffusion rule to obtain the labeled training sample set and the sample set to be judged; Based on neighborhood rough set attribute reduction, neighborhood relations are reconstructed and attributes are filtered in the labeled training sample set. High-order neighborhood spectrum segmentation iteration is introduced to perform spectrum decomposition, segmentation and update. Information potential energy field contraction reduction is used to perform boundary contraction, attribute compression and redundancy removal to obtain the key security risk attribute set. An initial risk prototype set is generated by the set of key safety risk attributes. An improved learning vector quantization network is constructed. The initial risk prototype set is then subjected to prototype rearrangement, boundary trimming and correlation optimization. Entropy positive mass flow balance rearrangement is introduced for global adjustment. Curvature balance processing is performed based on curvature equilibrium arc surface correction. Hyperspectral entropy cointegration optimization is used to uniformly optimize the risk prototypes, resulting in an optimized safety risk prototype library. The sample set to be judged is matched with the optimized safety risk prototype library, the continuous distance surface of the risk prototype is calculated, and the local extremum-curvature joint analysis is performed to obtain the real-time safety risk judgment result table. Based on the real-time security risk assessment results table, generate corresponding security risk classification early warning instructions and send the security risk classification early warning instructions to the monitoring and alarm platform.

2. The industrial production safety risk early warning method based on multimodal fusion according to claim 1, characterized in that, The multi-source heterogeneous monitoring data includes video image data, temperature data, pressure data, vibration data, gas concentration data, electrical operation data, and work record data.

3. The industrial production safety risk early warning method based on multimodal fusion according to claim 1, characterized in that, The preprocessing of multi-source heterogeneous monitoring data includes time alignment, outlier removal, missing value completion, noise reduction filtering, standardization, and structured coding.

4. The industrial production safety risk early warning method based on multimodal fusion according to claim 1, characterized in that, The obtained multimodal security sample library includes: Video image data, temperature data, pressure data, vibration data, gas concentration data, electrical operation data, and work record data are collected separately. Timestamps are added to each type of data according to a unified sampling clock to obtain the original monitoring dataset. The original monitoring dataset is subjected to time alignment, outlier removal, missing value completion, noise reduction filtering, standardization and structured coding. The standardization process is to subtract the sample mean of the monitoring variable from the actual value of each monitoring variable, and then divide by the sample standard deviation of the monitoring variable to obtain the preprocessed dataset. The preprocessed dataset is segmented and combined according to a unified time window. Video image data, temperature data, pressure data, vibration data, gas concentration data, electrical operation data, and work record data within the same time window are correlated and uniformly encapsulated to generate a multimodal safety sample library.

5. The industrial production safety risk early warning method based on multimodal fusion according to claim 1, characterized in that, The obtained candidate security risk attribute library includes: The attribute data and corresponding risk labels of all samples are extracted from the multimodal safety sample library. A cross-modal information flow coupling graph is constructed with each candidate attribute as a node. The conditional mutual information between each candidate attribute and the joint mutual information between each candidate attribute and the risk label are calculated. The mutual information gain value is calculated based on the conditional mutual information and joint mutual information corresponding to each candidate attribute. Candidate attributes are introduced one by one in descending order of mutual information gain value. After each candidate attribute is introduced, the mutual information gain value of the remaining candidate attributes is recalculated. During the process of introducing candidate attributes one by one, the change in mutual information gain before and after two consecutive introductions is continuously recorded. When the change in mutual information gain after two consecutive introductions is less than the change in mutual information gain after the previous introduction, the introduction is stopped, and all introduced candidate attributes are combined into a candidate security risk attribute library.

6. The industrial production safety risk early warning method based on multimodal fusion according to claim 1, characterized in that, The process of obtaining the labeled training sample set and the sample set to be judged includes: Read the candidate security risk attribute values, timestamp information and risk label information corresponding to all samples from the candidate security risk attribute library, project each sample onto a unified temporal coordinate system in chronological order, and use the value of each sample on each candidate security risk attribute as a spatial coordinate component to generate a risk trajectory matrix corresponding to the time sequence. In the risk trajectory matrix, the sample with the highest risk label confidence is selected as the anchor sample. Starting from each anchor sample, convection-diffusion iterative labeling is performed along the time continuity direction and the attribute adjacency direction. Unlabeled samples that maintain a continuous connection with the anchor sample are assigned the risk label of the corresponding anchor sample layer by layer. After each round of expansion, the iteration layer number corresponding to the labeled sample is recorded until all samples are labeled. All labeled samples are divided according to the iteration layer number. Samples with odd layer numbers are assigned to the labeled training sample set, and samples with even layer numbers are assigned to the undecided sample set.

7. The industrial production safety risk early warning method based on multimodal fusion according to claim 1, characterized in that, The obtained set of key security risk attributes includes: Read the candidate safety risk attribute values ​​and corresponding risk labels of each sample in the labeled training sample set, construct an initial neighborhood set according to the attribute distance relationship between samples, and build a neighborhood association graph with samples as nodes and neighborhood association relationships as edges to form the initial neighborhood structure; The high-order neighborhood spectrum segmentation iteration is performed on the initial neighborhood structure, the neighborhood association graph is spectral decomposed, the neighborhood association edges are segmented and updated according to the spectral decomposition results, unstable neighborhood connections are gradually stripped away, and the sample neighborhood relations are reconstructed after each segmentation to obtain the purified neighborhood structure after spectral segmentation. Based on the clean neighborhood structure, the information potential field is reduced by contraction. An information potential field is formed according to the local distribution state of each sample in the clean neighborhood structure. The neighborhood boundary of each sample is adjusted along the contraction direction of the information potential field to obtain the contracted neighborhood structure. The ability of each candidate attribute to distinguish risk labels is calculated based on the shrinking neighborhood structure. Redundant attributes are deleted and valid attributes are retained. Neighborhood boundary adjustment and attribute screening are repeated until the set of remaining attributes is stable, thus obtaining the set of key safety risk attributes.

8. The industrial production safety risk early warning method based on multimodal fusion according to claim 1, characterized in that, The optimized security risk prototype library includes: An improved learning vector quantization network is constructed based on a set of key safety risk attributes. An initial risk prototype set is generated according to the risk category. Entropy positive mass flow balance rearrangement is set in the prototype distribution layer of the learning vector quantization network. Curvature equilibrium arc surface correction is set in the category boundary adjustment layer. Hypergraph entropy cointegration optimization is performed in the prototype association and collaboration layer to form an improved learning vector quantization network. The initial risk prototype set is processed by entropy positive mass flow balancing rearrangement. The position of each risk prototype in the attribute space is adjusted according to the overall distribution relationship between each training sample and each risk prototype, resulting in a rearranged risk prototype set. The rearranged risk prototype set is processed by curvature equalization arc surface correction. Based on the category boundary structure formed between risk prototypes of the same category and risk prototypes of different categories, continuous trimming is performed on the boundary curvature region and the positional relationship of adjacent risk prototypes is adjusted simultaneously to obtain the boundary trimming risk prototype set. By using hypergraph entropy cointegration optimization to process the risk prototype set of boundary trimming, the correlation between training samples, key safety risk attributes and risk prototypes is uniformly coordinated and organized, and correlation optimization is performed on all risk prototypes to obtain the optimized risk prototype set. The labeled training sample set is input into the improved learning vector quantization network. The optimized risk prototype set is used as the initial training basis. The position and category boundary of each risk prototype are iteratively updated according to the category affiliation relationship between the sample and the risk prototype until the position change of each risk prototype meets the stopping condition. The optimized safety risk prototype library is then output.

9. The industrial production safety risk early warning method based on multimodal fusion according to claim 1, characterized in that, The obtained real-time security risk assessment result table includes: Read the key safety risk attribute values ​​corresponding to each sample in the sample set to be judged, and read the prototype coordinates and prototype categories corresponding to each risk prototype in the optimized safety risk prototype library. Match each sample to be judged with all risk prototypes to obtain the prototype distance sequence corresponding to each sample to be judged. Based on the prototype distance sequence corresponding to each sample to be judged, the distance values ​​of each risk prototype are continuously interpolated and expanded in the local neighborhood of the sample to be judged according to the spatial continuous distribution relationship, generating a continuous distance surface corresponding to each sample to be judged. Local extremum search and curvature joint analysis are performed on the continuous distance surface to extract local minimum points, local maximum points and corresponding curvature change regions in the continuous distance surface. The risk prototype category corresponding to the local minimum point with the greatest attraction intensity is determined, and the risk prototype category is used as the risk category of the corresponding sample to be judged. Based on the risk category corresponding to each sample to be judged, the curvature change state of the continuous distance surface, and the distance change relationship between the risk prototype of the same category and the nearest risk prototype of a different category, the risk level corresponding to each sample to be judged is determined. The sample identifier, risk category, risk level and judgment time are correlated to generate a real-time safety risk judgment result table.

10. The industrial production safety risk early warning method based on multimodal fusion according to claim 1, characterized in that, The generation of corresponding security risk classification early warning instructions includes: Each sample to be judged is classified and categorized according to its risk level. Samples to be judged that belong to the same risk category and the same risk level at the same judgment time are aggregated to generate a graded risk event set. Based on the number of samples to be judged, the corresponding risk category, the corresponding risk level, and the judgment time in each graded risk event set, a safety risk graded early warning instruction corresponding to each graded risk event set is generated. The event number, risk category identifier, risk level identifier, occurrence time, and sample source identifier are written into each safety risk graded early warning instruction. The warning instructions for each level of safety risk are sent to the monitoring and alarm platform in the order of the judgment time. The monitoring and alarm platform displays, stores and outputs alarms for each level of safety risk warning instruction.