Security risk supervision method and device, electronic equipment and storage medium
By constructing a collaborative architecture of cloud layer, edge layer, and device layer, cross-modal feature extraction, spatiotemporal alignment, and dynamic weighted fusion of multimodal data are achieved. This solves the problems of low recognition reliability and insufficient decision-making in existing coal mine video monitoring systems in complex underground environments, improves recognition accuracy and decision-making efficiency, and reduces costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIFANG WEIJIAMAO COAL POWER CO LTD
- Filing Date
- 2025-12-02
- Publication Date
- 2026-04-21
AI Technical Summary
Existing coal mine video monitoring systems are ill-suited to the complex underground environment, which is susceptible to changes in lighting and dust interference. They suffer from limited perception dimensions, lack of multimodal data fusion, and insufficient intelligent decision-making capabilities, resulting in low reliability, high cost, and difficulty in proactive early warning and efficient response.
Construct a collaborative architecture of cloud layer, edge layer, and device layer, deploy large visual models and multimodal models, perform pre-training and industry fine-tuning, realize cross-modal feature extraction, spatiotemporal alignment and dynamic weighted fusion of multimodal data, generate risk event identification results, and trigger alarms and decision-making schemes.
It improves the accuracy and generalization ability of coal mine operation safety risk identification, reduces the threshold for AI application and the total life cycle cost, and realizes the leap from passive alarm to proactive early warning and accurate decision-making, thus ensuring the safety of coal mine operations.
Smart Images

Figure CN121903791A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of safety management technology, and in particular to a method, apparatus, electronic device and storage medium for monitoring safety risks. Background Technology
[0002] In the field of coal mine operation safety monitoring, intelligent monitoring technology for industrial safety production is a key support for ensuring underground operation safety and reducing accident risks. Currently, mainstream coal mine video monitoring systems mainly rely on traditional computer vision algorithms or dedicated small deep learning models (such as YOLO, CNN, etc.) to carry out safety risk monitoring. The core application scenarios are concentrated on the detection and identification of single targets such as safety helmet wearing detection, flame recognition, and smoke detection, which provides certain technical support for basic safety supervision of coal mines.
[0003] However, existing technologies struggle to meet the demands for precise, efficient, and intelligent safety supervision in the complex environment of underground coal mines. Firstly, existing models are typically trained in specific scenarios, making them ill-suited to the complex and ever-changing lighting conditions, dust interference, and equipment obstructions encountered underground. When new violations or undefined anomalies occur, the system must re-collect data, label samples, and retrain the model, resulting in long development cycles and high deployment costs. Secondly, the perception dimension is limited, lacking multimodal collaboration. Existing systems primarily rely on video image information, lacking effective fusion of multi-dimensional information such as sound, equipment operating parameters, and environmental sensor data. When video quality deteriorates due to the underground environment, the system's recognition reliability significantly decreases, making comprehensive judgments of complex working conditions impossible. Finally, intelligent decision-making capabilities are insufficient. Most existing systems remain at the "perception-alarm" level, lacking the ability to understand the semantic context of events, unable to perform root cause analysis, trend prediction, or provide handling suggestions. They still heavily rely on human experience for judgment and response, leading to low emergency response efficiency and the potential for overlooking potential risks. Summary of the Invention
[0004] This application provides a security risk monitoring method, device, electronic device, and storage medium. It can solve the problems in related technologies, which are poor in generalization ability, high cost of adapting to new risk types, low recognition reliability, and difficulty in proactive early warning and efficient handling, due to reliance on traditional computer vision algorithms or small deep learning models, single perception dimension, lack of multimodal data spatiotemporal alignment and dynamic fusion mechanism and intelligent decision support.
[0005] According to a first aspect of this application, a method for monitoring security risks is provided, comprising: A collaborative architecture is constructed, consisting of a cloud layer, an edge layer, and a device layer. The cloud layer deploys large visual models and multimodal models for pre-training and industry-specific fine-tuning. The edge layer deploys application models and multimodal information fusion algorithms for data analysis, data fusion, and event reasoning. The device layer deploys multimodal data acquisition devices to acquire multimodal data. Cross-modal feature extraction is performed on multimodal data to generate multi-dimensional features that include visual features, auditory features, device parameter features, and text semantic features; Based on a multimodal information fusion algorithm, timestamps and spatial coordinates of multi-dimensional features are accurately matched to achieve spatiotemporal alignment. Based on the application model, the multi-dimensional features after spatiotemporal alignment are dynamically weighted and fused, and the fusion weights are adjusted according to the confidence of each modality to generate risk event identification results. When the identification result is abnormal, an alarm is triggered, a decision plan is generated, and it is pushed to the scheduling terminal.
[0006] According to a second aspect of this application, a security risk monitoring device is provided, comprising: The building module is configured to construct a collaborative architecture of cloud layer, edge layer, and device layer. The cloud layer deploys large visual models and multimodal large models for pre-training and industry fine-tuning. The edge layer deploys application models and multimodal information fusion algorithms for data analysis, data fusion, and event reasoning. The device layer deploys multimodal data acquisition devices to acquire multimodal data. The extraction module is configured to perform cross-modal feature extraction on multimodal data, generating multi-dimensional features that include visual features, auditory features, device parameter features, and text semantic features. The alignment module is configured to accurately match timestamps and spatial coordinates of multi-dimensional features based on a multimodal information fusion algorithm to achieve spatiotemporal alignment. The first generation module is configured to dynamically weight and fuse the spatiotemporally aligned multi-dimensional features based on the application model, and adjust the fusion weights according to the confidence of each modality to generate risk event identification results. The second generation module is configured to trigger an alarm and generate a decision plan and push it to the scheduling terminal when the identification result is abnormal.
[0007] According to a third aspect of this application, an electronic device is provided, comprising: At least one processor; and memory that is communicatively connected to at least one processor; The memory stores instructions that can be executed by at least one processor, which enables the at least one processor to perform the security risk monitoring method described in the first aspect.
[0008] According to a fourth aspect of this application, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the security risk monitoring method of the first aspect described above.
[0009] According to a fifth aspect of this application, a computer program product is provided, including a computer program that, when executed by a processor, implements the security risk monitoring method as described in the first aspect above.
[0010] This application provides a security risk monitoring method, device, electronic device, and storage medium, comprising: constructing a collaborative architecture of cloud layer, edge layer, and device layer, wherein the cloud layer deploys a large visual model and a large multimodal model for pre-training and industry fine-tuning; the edge layer deploys an application model and a multimodal information fusion algorithm for data analysis, data fusion, and event reasoning; and the device layer deploys multimodal data acquisition devices to acquire multimodal data; performing cross-modal feature extraction on the multimodal data to generate multi-dimensional features including visual features, auditory features, device parameter features, and text semantic features; accurately matching the timestamps and spatial coordinates of the multi-dimensional features based on the multimodal information fusion algorithm to achieve spatiotemporal alignment; dynamically weighting and fusing the spatiotemporally aligned multi-dimensional features based on the application model, adjusting the fusion weights according to the confidence level of each modality to generate risk event identification results; and triggering an alarm and generating a decision plan and pushing it to the scheduling terminal when the identification result is abnormal. This application constructs a collaborative architecture encompassing the cloud layer, edge layer, and device layer. It extracts cross-modal features from multimodal data to generate multi-dimensional features including visual, auditory, device parameter, and textual semantic features. Then, based on a multimodal information fusion algorithm, it accurately matches the timestamps and spatial coordinates of these multi-dimensional features to achieve spatiotemporal alignment. Subsequently, an application model dynamically weights and fuses the spatiotemporally aligned multi-dimensional features to generate risk event identification results. When the identification result is abnormal, an alarm is triggered, and a decision plan is generated and pushed to the scheduling end. Therefore, this application solves the problems in related technologies that rely solely on traditional computer vision algorithms or small deep learning models, have limited perception dimensions, lack spatiotemporal alignment and dynamic fusion mechanisms for multimodal data, and lack intelligent decision support. These problems result in poor scenario generalization ability, high adaptation costs for new risk types, low identification reliability, and difficulty in proactive early warning and efficient handling. This achieves the technical effect of improving the accuracy and generalization ability of coal mine operation safety risk identification, reducing the threshold and lifecycle cost of AI applications, and realizing a leap from passive alarms to proactive early warning and precise decision support, thus ensuring coal mine operation safety.
[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0012] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 A flowchart illustrating a security risk monitoring method provided in an embodiment of this application; Figure 2 A flowchart illustrating another security risk monitoring method provided in this application embodiment; Figure 3 This is a schematic diagram of a security risk monitoring device provided in an embodiment of this application. Detailed Implementation
[0014] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of this application, including various details to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0015] The following description, with reference to the accompanying drawings, describes a security risk monitoring method, apparatus, electronic device, and storage medium according to embodiments of this application.
[0016] Figure 1 This is a flowchart illustrating a security risk monitoring method provided in an embodiment of this application.
[0017] like Figure 1 As shown, the method includes the following steps: Step 101: Construct a collaborative architecture of cloud layer, edge layer, and device layer. In the cloud layer, deploy large visual models and multimodal models for pre-training and industry fine-tuning. In the edge layer, deploy application models and multimodal information fusion algorithms for data analysis, data fusion, and event reasoning. In the device layer, deploy multimodal data acquisition devices to acquire multimodal data.
[0018] In some embodiments, constructing a collaborative architecture of cloud layer, edge layer, and device layer is the fundamental support for realizing intelligent supervision of coal mine operation safety. Each layer is equipped with dedicated components and core capabilities according to its functional positioning. The device layer, as the "sensory nerve endings" of data acquisition, is equipped with various multimodal data acquisition devices. Specifically, these include high-definition network cameras deployed at key locations such as coal conveyor belts, mine pits, and pump rooms in the coal mine to collect visual data of the operation scene; fiber optic acoustic and temperature sensing systems for belt conveyor monitoring, which can capture acoustic and temperature information during belt operation; Kuanghong smart wristbands for personnel positioning and vital sign monitoring, which can acquire personnel location and physical status data in real time; Kuanghong smart drainage control boxes and their connected intrinsically safe level / flow sensors, which can collect level and flow parameters of the drainage system; in addition, various PLCs (programmable logic controllers) are included, whose output equipment operating parameters (OT data) can reflect the real-time operating status of the equipment. These devices work together to achieve comprehensive acquisition of multi-dimensional data from the coal mine. The edge layer is deployed in the local data center of the coal mine, equipped with AI inference servers (such as servers equipped with Kunpeng 920 CPU and Atlas 3001Pro inference cards, a single unit can support 100 video streams) and the coal mine AI safety monitoring platform. The core deployment includes application models and multimodal information fusion algorithms. On the one hand, it receives various types of data transmitted from the device layer, conducts real-time video analysis through application models, and completes data fusion and event inference with the help of multimodal information fusion algorithms. On the other hand, it can perform rapid control operations to the device layer (such as linking sound and light alarms), and can operate independently when the network is interrupted, meeting the low latency and high reliability requirements of the supervision business. The cloud layer relies on industrial internet platforms (such as the Ordos Industrial Internet Platform) to build powerful computing support, deploying large-scale visual models and multimodal models. The visual model, based on the Transformer architecture, learns general visual representations through self-supervised pre-training on massive amounts of internet images and videos, possessing capabilities such as image segmentation and object detection. The multimodal model, also based on the Transformer architecture, integrates text, image, and audio data for joint training to master cross-modal correlation characteristics. Both models, through pre-training and industry fine-tuning on massive amounts of coal mine data, form a model capability adapted to coal mine scenarios. The trained and optimized model is then pushed to the edge layer through a cloud-edge collaboration mechanism. The beneficial effects of this collaborative architecture are that it achieves layered collaboration in data acquisition, real-time processing, and model training. This ensures comprehensive acquisition of multimodal data, reduces response latency through real-time processing at the edge layer, and enhances the system's adaptability to coal mine scenarios through large-scale model training in the cloud layer, providing a stable and efficient architectural foundation for subsequent safety supervision.
[0019] Step 102: Perform cross-modal feature extraction on the multimodal data to generate multi-dimensional features that include visual features, auditory features, device parameter features, and text semantic features.
[0020] In some embodiments, cross-modal feature extraction of the multimodal data is a key step in transforming various types of raw data collected at the device layer into effective information that can be used for subsequent analysis. Its core is to adopt appropriate extraction methods for multimodal data from different sources and of different types to generate visual features, auditory features, device parameter features, and text semantic features respectively, forming a feature set covering multiple dimensions. The extraction of visual features is based on video streams captured by high-definition network cameras deployed at the equipment layer. The video frames are processed using an edge-layer application model. For example, it identifies image details such as "stagnant coal flow at the feed inlet and overflowing coal lumps" in coal conveyor belt scenarios, or visual information such as "the head outline of personnel not wearing safety helmets" and "the movement trajectory of personnel crossing the conveyor belt" in personnel operation scenarios. These concrete image contents are transformed into machine-recognizable feature vectors. Auditory features are derived from audio data captured by a fiber optic acoustic signature temperature sensing system. This system can monitor the sound signals of equipment operating in the coal mine in real time. When extracting features, it focuses on analyzing the frequency, amplitude, and variation patterns of the acoustic signature. For example, it analyzes specific frequency bands when "friction noise increases abnormally" during conveyor belt operation, or the amplitude variation characteristics of "abnormal water flow impact sound" in the pump room, thus forming an auditory feature representation. The parameter feature extraction targets structured data such as equipment operating parameters (OT data) output by PLC and values collected by intrinsically safe level / flow sensors. During the extraction process, trend analysis and key value capture are performed on these time-series data. For example, "the range of abnormally high current" and "the duration of current exceeding the threshold" are extracted from belt motor current data, and "the rate of rapid rise in liquid level" is extracted from liquid level sensor data, transforming the equipment operating status into quantifiable feature indicators. Text semantic features are generated based on unstructured text data such as coal mine safety regulations, accident reports, and equipment maintenance records. With the help of the semantic understanding capabilities of cloud-based multimodal large models, word embedding and semantic encoding are performed on the text content to extract key semantic information such as "typical handling procedures for coal blockage in belt conveyors" and "common causes of water pump failures," forming feature data that reflects the core meaning of the text.
[0021] This cross-modal feature extraction step comprehensively covers the visual, auditory, equipment operation, and textual information dimensions in coal mine operation scenarios, avoiding the limitations of single-modal features. It provides a complete and effective data foundation for subsequent multi-dimensional feature fusion analysis, which helps to improve the comprehensiveness and accuracy of subsequent risk event identification.
[0022] Step 103: Based on the multimodal information fusion algorithm, timestamps and spatial coordinates of multi-dimensional features are accurately matched to achieve spatiotemporal alignment.
[0023] In some embodiments, timestamp matching needs to be based on the unified industrial clock system within the coal mine, and the acquisition timestamps of each modal feature need to be calibrated. For example, the video frames output by the high-definition network camera at the equipment layer will carry a shooting timestamp, the friction sound audio data recorded by the fiber optic acoustic sensor system will have an acquisition timestamp, and the belt motor current parameters transmitted by the PLC will also have a generation timestamp. The multimodal information fusion algorithm will read these original timestamps, and through the deviation calculation with the unified clock, fine-tune the timestamp of each feature data to ensure that the visual features, auditory features, and equipment parameter features corresponding to the on-site events that occur at the same time (such as an abnormality at a certain position of the belt) are completely consistent in their timestamps, avoiding misalignment problems such as "a video frame at one time corresponds to current data at another time" caused by equipment clock deviation. Spatial coordinate matching requires establishing a spatial coordinate mapping system underground in the coal mine, associating the acquisition location of each modal feature with the actual underground space. For example, high-definition network cameras installed in the coal conveyor belt unloading area have a clearly defined monitoring coverage coordinate range; each sensor node of the fiber optic acoustic signature sensing system is laid along the belt, corresponding to specific belt section coordinates; equipment parameter features, such as belt motors and water pumps, also have fixed underground installation coordinates. The algorithm uses a preset coordinate mapping table to associate and bind the camera monitoring coordinates corresponding to visual features, the sensor node coordinates corresponding to auditory features, and the equipment location coordinates corresponding to equipment parameter features. This ensures that each modal feature in a certain spatial area (within 5 meters of the unloading area) can accurately correspond to the same underground spatial range, avoiding situations where "video features in area A are mistakenly associated with acoustic signature features in area B" due to spatial location mismatch. The multimodal information fusion algorithm automatically completes the above timestamp calibration and coordinate binding process without manual intervention, further ensuring the efficiency and accuracy of alignment. The beneficial effect of this step is that it solves the problem of spatiotemporal misalignment of multidimensional features caused by differences in clock and installation location of different acquisition devices, ensuring that each modal feature can collaboratively reflect the on-site state at the same time and in the same space, providing a reliable data association basis for subsequent dynamic weighted fusion, and reducing the risk of misjudgment caused by spatiotemporal deviation.
[0024] Step 104: Dynamically weighted and fused the spatiotemporally aligned multi-dimensional features based on the application model, and adjusted the fusion weights according to the confidence of each modality to generate risk event identification results.
[0025] In some embodiments, the dynamic weighted fusion of spatiotemporally aligned multi-dimensional features based on the application model is the core process of a lightweight application model deployed at the edge layer. This model integrates visual, auditory, equipment parameter, and textual semantic features that have achieved precise temporal and spatial matching, and outputs risk event identification results through non-fixed weighted fusion calculations. The application model first assesses the reliability of each modal feature and generates a modal confidence score. This confidence score is automatically determined by the model based on the clarity, completeness, and stability of the features. For example, in a coal mine "roller jamming" scenario, if the spatiotemporally aligned visual features are unclear due to belt obstruction causing "abnormal roller rotation," the application model will assign a low confidence score (e.g., 0.3). Conversely, the synchronous auditory features ("roller jamming noise" captured by the fiber optic acoustic signature sensor) are interference-free and feature-clear, and the equipment parameter features ("roller drive motor speed fluctuation" data transmitted by the PLC) are continuous and without abnormal fluctuations. The application model will assign high confidence scores to these two types of features (e.g., 0.92 and 0.88, respectively). Subsequently, the application model dynamically adjusts the fusion weights of each modality feature based on the confidence level, following the logic of "the higher the confidence level, the greater the weight." For example, it assigns 10% weight to low-confidence visual features, 50% weight to high-confidence auditory features, and 40% weight to equipment parameter features, rather than using equal weights. After weight allocation, the application model performs a weighted calculation on the analysis results of each modality feature. If the overall result exceeds a preset risk judgment threshold (e.g., 0.75), a risk event identification result of "roller jamming" is generated; if it does not exceed the threshold, it is judged as a normal operating state. This avoids misjudgments caused by unreliable single modalities or fixed-weight fusion, allowing high-reliability modalities to play a full role and effectively improving the accuracy and anti-interference capability of coal mine risk event identification.
[0026] Step 105: When the identification result is abnormal, trigger an alarm, generate a decision plan, and push it to the scheduling terminal.
[0027] In some embodiments, when the identification result is determined to be abnormal, the abnormal handling process of the edge-layer coal mine AI safety monitoring platform will be immediately initiated. First, the alarm mechanism is triggered—the platform automatically links the audible and visual alarms in the corresponding area of the coal mine site, activating audible and visual warnings at the location of the abnormal event (such as the feed inlet of a coal conveyor belt blockage, or the water accumulation area in the pump room). At the same time, alarm information containing the abnormality type, precise location of occurrence, identification timestamp, and core judgment criteria is generated, allowing on-site operators to quickly detect the abnormality and grasp the basic situation. Subsequently, the platform calls the decision-making scheme generation module. This module relies on a built-in knowledge base that integrates coal mine safety regulations, historical accident handling cases, and expert experience, and analyzes the associated data of the abnormal event to generate a structured decision-making scheme. For example, for the "coal conveyor belt blockage" abnormality, a scheme will be generated that "prioritizes stopping the operation of the conveyor belt and notifies maintenance personnel to check the coal flow channel at the feed inlet; if the coal blockage is severe, the backup coal cleaning equipment should be activated to assist in clearing the blockage." The scheme clearly defines the operation priority and specific steps to avoid vague guidance. Finally, the platform, through its AI assistant module, simultaneously pushes alarm information and decision-making solutions to the coal mine dispatch terminal—which includes the operation terminal, display screen, and mobile work equipment of the coal mine dispatch center—in the form of voice broadcasts and text / image messages. This ensures that dispatchers can obtain complete handling information without additional queries. This upgrades anomaly response from "alarm only" to "alarm + precise decision-making," reducing manual analysis time, improving the timeliness and standardization of handling coal mine anomalies, and reducing the risk of accident escalation.
[0028] Compared with related technologies, this embodiment constructs a collaborative architecture of cloud layer, edge layer, and device layer. The cloud layer deploys large-scale visual models and multimodal models for pre-training and industry-specific fine-tuning. The edge layer deploys application models and multimodal information fusion algorithms for data analysis, data fusion, and event reasoning. The device layer deploys multimodal data acquisition devices to acquire multimodal data. Cross-modal feature extraction is performed on the multimodal data to generate multi-dimensional features including visual features, auditory features, device parameter features, and textual semantic features. Based on the multimodal information fusion algorithm, the timestamps and spatial coordinates of the multi-dimensional features are accurately matched to achieve spatiotemporal alignment. The application model dynamically weights and fuses the spatiotemporally aligned multi-dimensional features, adjusting the fusion weights according to the confidence level of each modality to generate risk event identification results. When the identification result is abnormal, an alarm is triggered, a decision plan is generated, and pushed to the scheduling terminal. It can solve the problems in related technologies that rely solely on traditional computer vision algorithms or small deep learning models, have a single perception dimension, lack multimodal data spatiotemporal alignment and dynamic fusion mechanisms and intelligent decision support, resulting in poor scenario generalization ability, high cost of adapting to new risk types, low recognition reliability, and difficulty in proactive early warning and efficient handling. It can improve the accuracy and generalization ability of coal mine operation safety risk identification, reduce the threshold and full life cycle cost of AI application, and achieve the technical effect of making a leap from passive alarm to proactive early warning and accurate decision support, thus ensuring the safety of coal mine operations.
[0029] Figure 2 A flowchart illustrating another security risk monitoring method provided in this application embodiment includes the following steps: Step 201: The visual large model and the multimodal large model adopt the Transformer architecture to perform self-training based on visual data, text data, and structured data, and fine-tuning based on industry feature data to generate application models with the functions of visual large models and multimodal large models. The application models are then pushed to the edge layer through the cloud-edge collaboration mechanism.
[0030] In some embodiments, both the visual large model and the multimodal large model adopt the Transformer architecture—this architecture, by segmenting data into blocks and transforming it into sequences, can more efficiently learn the correlation characteristics between data, and has advantages in feature extraction accuracy and generalization ability compared to traditional models. The training of the visual large model is based on massive visual data, including more than 1 billion publicly available images on the Internet, more than 100TB of video, and annotated images specific to coal mine scenarios (such as images of scene images such as belt misalignment and personnel violations). It masters general visual representation capabilities through a self-training process. The multimodal large model integrates visual data, text data (such as coal mine safety regulations and accident reports), and structured data (such as equipment sensor time-series logs) for joint self-training, focusing on learning the correlation logic of different modal information (such as the correspondence between the visual feature of "belt smoke" and the text description of "smoke alarm"). After self-training, both types of models are further fine-tuned based on coal mine industry characteristic data. This industry characteristic data encompasses labeled scenario data from coal mines across the city (covering 24 typical coal mine safety scenarios), historical equipment failure data, and expert experience summaries. Fine-tuning ensures the model accurately adapts to the coal mine operating environment, ultimately generating an application model with coal mine safety supervision functions (i.e., a lightweight model adapted to edge layer computing power requirements). After model generation, a cloud-edge collaboration mechanism facilitates push notifications: the cloud layer utilizes a dedicated data transmission channel built on an industrial internet platform to send the lightweight application model to the edge layer's AI inference server in an incremental update manner. During the push process, the model's integrity and compatibility are automatically verified, ensuring direct deployment and use at the edge layer without secondary debugging. The Transformer architecture and industry fine-tuning ensure model adaptation to coal mine scenarios, while cloud-edge collaboration enables efficient model deployment, providing reliable model support for edge layer data analysis.
[0031] Step 202: The multimodal data acquisition equipment includes cameras deployed at key locations in the coal mine, as well as fiber optic acoustic temperature sensing systems, employee vital sign acquisition devices, liquid level / flow sensors, and PLC equipment.
[0032] In some embodiments, cameras can be deployed at key locations in the coal mine, including areas such as the coal conveyor belt discharge port, the mine working face, the pump room, and the intersection of underground roadways. High-definition network cameras are selected to capture clear on-site images and video streams, focusing on collecting visual data such as personnel work behavior, equipment operating status, and environmental conditions. A fiber optic acoustic signature temperature sensing system is laid along the coal conveyor belt, using fiber optic sensing technology to collect acoustic signature signals and surface temperature data in real time during belt operation, accurately capturing sound and temperature changes during abnormal equipment operation. The employee vital sign collection device uses a Kuanghong smart bracelet, worn by underground workers, which can collect vital sign data such as heart rate and blood oxygen in real time. It also integrates positioning functions to obtain the real-time coordinates of personnel underground, ensuring that personnel safety status and location can be monitored. Liquid level / flow sensors are connected to the Kuanghong smart drainage control box and installed in areas such as the pump room and underground water accumulation points to collect liquid level height and water flow data of the drainage system, reflecting the operating efficiency of the drainage equipment. PLC equipment is connected to the main production equipment in the coal mine to collect equipment operating parameters in real time, forming structured equipment operating data. All data acquisition devices were selected according to intrinsically safe standards for underground coal mines, ensuring stable operation in dusty, humid, and explosion-proof environments. Through diverse acquisition equipment, comprehensive acquisition of multi-dimensional data, including visual data, acoustic signatures, temperature data, personnel vital signs, and equipment parameters, is achieved, providing sufficient and complete raw data for subsequent feature extraction and fusion analysis.
[0033] Step 203: Extract visual features from the video stream and extract text semantic features from the text data by applying the model.
[0034] In some embodiments, relying on the application model deployed at the edge layer, visual features and textual semantic features are extracted from video streams and text data, respectively. Specifically, the application model, when processing video streams, selects an appropriate feature extraction direction based on the actual needs of the coal mine operation scenario: During target detection, the model locates the position and category of key targets in the video frame. For example, in an underground working face scenario, it accurately identifies targets such as "personnel not wearing safety helmets" and "illegally placed tools," outputting the target's bounding box and category label. During image segmentation, the model splits video frames according to scene elements. For example, in a coal conveyor belt scenario, it segments different elements such as "belt body," "idlers," and "coal flow" into independent regions for subsequent targeted analysis. During scene understanding, the model comprehensively judges the overall scene state by considering multiple elements in the video frame. For example, based on the image of "expanding water area in the mining pit and drainage pump not starting," it understands the scene semantics of "abnormal drainage in the mining pit." For text data processing, the model first collects various text materials from coal mine scenarios, including equipment operation manuals, safety procedures, and real-time event descriptions. Then, it processes these texts using natural language processing (NLP) techniques—first performing basic preprocessing such as word segmentation and part-of-speech tagging to remove meaningless stop words; then using word embedding technology to convert keywords in the text into low-dimensional vectors; and finally, integrating the word vectors through semantic encoding to generate text semantic feature vectors that reflect the core meaning of the text, ensuring that the machine can accurately understand safety-related information in the text. Targeted extraction of key visual and textual features from coal mine scenarios provides high-quality visual and semantic data support for multi-dimensional feature fusion.
[0035] Step 204: Extract auditory features from audio data using a fiber optic acoustic signature sensing system. The auditory features include acoustic signature patterns, abnormal sounds, or frequency characteristics.
[0036] In some embodiments, audio data acquisition and feature analysis are accomplished using fiber optic acoustic signature sensing systems deployed at the coal mine site. These systems are typically laid along key equipment such as coal conveyor belts and pump room pipelines. Their sensing fibers can sensitively capture various sound signals generated during equipment operation and convert them into analyzable audio data. When extracting auditory features, the system first performs noise reduction processing on the acquired audio data, filtering out background noise from the underground environment such as dust friction and airflow to ensure the clarity of the effective sound signal. Subsequently, the system extracts three core auditory features from the processed audio data: First, voiceprint patterns, which are stable sound characteristics during normal equipment operation. For example, the uniform friction sound produced when a belt conveyor rotates normally exhibits a regular, periodic fluctuation in its voiceprint pattern, serving as a benchmark for judging whether the equipment is functioning properly. Second, abnormal sounds, which are sounds significantly different from normal voiceprint patterns. For instance, when an idler roller jams due to bearing wear, it produces an intermittent "clunking" sound, clearly different from the normal uniform friction sound. The system can quickly identify and mark these abnormal sounds. Third, frequency features, which involve quantifying the frequency of the sound signal. For example, the frequency of normal belt operation is concentrated between 200-500Hz. When the belt deviates and rubs against the frame, the sound frequency rises to 800-1200Hz. The system records this frequency change and converts it into quantifiable frequency feature data. These extracted auditory features are transmitted synchronously with other modal features, providing a basis for subsequent fusion analysis in terms of sound dimensions. By using a professional sensing system and feature extraction logic, abnormal sounds during the operation of coal mining equipment are captured, filling the gaps in the detection of potential equipment faults that are difficult to find by visual means alone, and improving the comprehensiveness of feature extraction.
[0037] Step 205: Extract equipment parameter features from the equipment operating parameters output by the PLC device. The equipment parameter features include current, voltage, temperature, or flow time sequence data.
[0038] In some embodiments, the PLC device, as the control core of the main production and auxiliary equipment in the coal mine, will collect and output various operating parameters of the connected equipment in real time. These parameters cover key indicators such as current, voltage, temperature, and flow rate, and are recorded in the form of time-series data—that is, parameter values are continuously collected at fixed time intervals (such as once per second) to form a parameter sequence that changes over time. When extracting features, targeted processing is carried out for different parameter types: For current parameters, the focus is on analyzing the numerical fluctuations and peak values in the parameter sequence. For example, the current of a belt motor is stable at 30-35A during normal operation, but when coal flow blockage occurs, the current will rise rapidly to over 50A. The system will extract features such as "the magnitude of the abnormal current increase" and "the duration of continuous exceedance of the threshold". For voltage parameters, the stability of the parameter sequence is considered. For example, the normal voltage of a water pump motor is 380V. If the voltage fluctuates frequently by more than ±8V, features such as "voltage fluctuation frequency" and "maximum fluctuation range" will be extracted. For temperature parameters, the trend of parameter sequence changes is tracked. For example, the normal operating temperature of a hydraulic support cylinder is 40-60℃. If the temperature continues to rise at a rate of 5℃ per hour, features such as "temperature rise rate" and "difference between the current temperature and the threshold" will be extracted. For flow parameters, the numerical changes in the parameter sequence are analyzed. For example, the normal drainage flow rate of a drainage pump is 40-50m³ / h. When the pipeline is blocked, the flow rate drops to below 20m³ / h, and features such as "flow rate decrease ratio" and "duration of flow rate stabilizing at low values" will be extracted. These extracted equipment parameter features are presented in the form of quantitative data, which can accurately reflect the real-time operating status of the equipment. By using quantified equipment operating parameter features, the health status of the equipment is presented intuitively, providing an objective and accurate basis for identifying risk events and reducing subjective judgment errors.
[0039] Step 206: Use a time-space synchronization algorithm to control the synchronization error between the video frame timestamp and the device current data timestamp within a preset time threshold.
[0040] In some embodiments, the system establishes a unified time reference system in the coal mine, typically relying on an industrial-grade Network Time Protocol (NTP) server. This provides a unified clock calibration source for high-definition network cameras capturing video frames and PLC devices outputting current data, avoiding initial time deviations caused by clock drift within the devices themselves. Subsequently, the spatiotemporal synchronization algorithm reads the original timestamps of the video frames and device current data respectively. The video frames are automatically timestamped by the camera according to the shooting frame rate (e.g., 25 frames / second commonly used in coal mine scenarios), recording the acquisition time of each frame. The device current data is timestamped by the PLC according to a fixed acquisition cycle (e.g., 10 times per second), marking the specific time of each current sampling. The algorithm compares the two types of original timestamps with the unified NTP clock, calculates their respective time deviation values (e.g., the original timestamp of a certain video frame is 20 milliseconds slower than the unified clock, and the original timestamp of a certain current data is 15 milliseconds faster than the unified clock), and dynamically calibrates the original timestamps based on the deviation values, ensuring that the calibrated video frame timestamps and current data timestamps are strictly aligned with the unified clock. Meanwhile, the algorithm presets a time threshold (typically set to within 100 milliseconds, considering the real-time requirements of coal mine safety supervision). After each calibration, it checks the synchronization error. If the error exceeds the threshold, a second calibration is automatically triggered until the timestamp synchronization error between the two types of data is controlled within the threshold range. For example, when monitoring the operating status of a belt motor, the calibrated video frame showing "abnormal vibration on the belt surface" (timestamp 14:30:00.050) accurately corresponds to the "sudden increase in motor current" data (timestamp 14:30:00.060) at the same moment, with an error of only 10 milliseconds, meeting the preset requirements. This avoids misjudgments caused by time misalignment, such as "the video display is normal but matches abnormal current data," and provides a precise time correlation basis for subsequent multi-feature fusion.
[0041] Step 207: The camera coordinate system and the spatial coordinates of the fiber optic acoustic signature sensing system are transformed into three-dimensional coordinates through feature space mapping technology to ensure that the spatial position error is controlled within the preset spatial range.
[0042] In some embodiments, the system will construct a global three-dimensional coordinate system for underground coal mines, with a fixed underground reference point (such as the center point of the main shaft) as the origin. Combined with the actual mapping data of underground roadways and equipment installation locations, a unified spatial coordinate system including X (horizontal roadway direction), Y (vertical roadway direction), and Z (underground depth direction) axes will be established, and the parameters of this coordinate system will be pre-set into the multimodal information fusion algorithm. Next, the spatial position parameters of the camera and the fiber optic acoustic signature sensor system are obtained respectively: For the camera, its coordinates in the global three-dimensional coordinate system are determined by the underground mapping data of its installation location (e.g., installed 3 meters above the feed outlet of the coal conveyor belt, 200 meters horizontally from the main shaft opening) (e.g., 200, 50, -15, where the negative sign indicates the underground depth). At the same time, the monitoring angle range of the camera (e.g., horizontal angle 60°, vertical angle 45°) is recorded to clarify the spatial area covered by its vision. For the fiber optic acoustic signature sensor system, the coordinates of each sensor node in the global three-dimensional coordinate system are determined based on the node distribution along the belt (e.g., one sensor node every 5 meters, a total of 20 nodes from the start of the belt to the feed outlet) (e.g., 190, 50, -15, 195, 50, -15, etc.) according to its installation location (e.g., one sensor node every 5 meters, a total of 20 nodes from the start of the belt to the feed outlet), to clarify the spatial range of the belt section monitored by each node. The feature space mapping technology generates a coordinate transformation matrix based on the aforementioned parameters. This matrix converts the visual feature positions in the camera coordinate system (local coordinates with the camera lens as the origin) to the node coordinates (global coordinates) of the fiber optic acoustic signature sensing system. For example, the visual region corresponding to the camera's local coordinates (2, 1, 0.5) is mapped to global coordinates (201.8, 50.9, -14.7) using the transformation matrix, ensuring a precise correspondence with the coordinates (202, 50, -15) of the neighboring fiber optic sensing node. Simultaneously, the algorithm presets a spatial error range (typically set within 50 cm, considering the density of coal mine equipment layout), automatically verifying spatial position errors after conversion to ensure that both types of features correspond to the same underground spatial area. This solves the spatial misalignment problem caused by coordinate system differences between different devices, preventing "features of belt segment A monitored by the camera from being mistakenly associated with acoustic signature features of segment B," and providing a reliable spatial association basis for multimodal feature fusion.
[0043] Step 208: Using the application model based on cross-modal association weights, dynamically weight the confidence of visual features, auditory features, device parameter features, and text semantic features.
[0044] In some embodiments, this step relies on the cross-modal association weights built into the application model to dynamically weight the confidence levels of visual features, auditory features, equipment parameter features, and textual semantic features. The cross-modal association weights are key parameters learned by the application model during the training phase (based on fine-tuning of the large industry model in the cloud layer and scenario-based fine-tuning in the edge layer). These weights are preset based on the degree of correlation between different modal features and risk events in coal mine operation scenarios. For example, in the "coal belt blockage" risk scenario, auditory features (friction noise) and equipment parameter features (increased motor current) are more closely associated with the event, and their association weights are preset to 0.4 and 0.4, respectively. Visual features (coal accumulation) have a slightly lower correlation due to their susceptibility to dust obscuring, and their weight is preset to 0.2. Textual semantic features (such as descriptions of historical coal blockage events) serve as an auxiliary reference, and their weight is preset to 0.1. The application model first obtains the real-time confidence scores of each modality after spatiotemporal alignment. For example, at a certain moment, the confidence score of visual features after alignment is 0.6 (the image is not clear enough due to slight dust accumulation), the confidence score of auditory features is 0.9 (the voiceprint signal is free of interference), the confidence score of equipment parameter features is 0.85 (the current data is stable and reliable), and the confidence score of text semantic features is 0.7 (matching historical text descriptions of coal blockage). Then, it calculates each item according to the formula "confidence score of each modality × corresponding cross-modal association weight", that is, 0.6×0.2=0.12, 0.9×0.4=0.36, 0.85×0.4=0.34, 0.7×0.1=0.07. Finally, the results of each item are summed to obtain the preliminary weighted confidence score (0.12+0.36+0.34+0.07=0.89), which provides basic data for subsequent weight adjustment and result determination. By pre-setting contextualized association weights, it is ensured that modal features closely related to risk events can play a core role in the fusion, avoiding interference from irrelevant modalities and improving the relevance of weighted calculation.
[0045] Step 209: Adjust the weights of each modal feature according to the real-time environmental conditions. When the confidence of a certain modal feature is lower than the preset threshold, increase the weights of other high-confidence modal features.
[0046] In some embodiments, the application model presets a confidence threshold for each modal feature (usually set to 0.5, taking into account the complexity of the coal mine environment). This threshold represents the minimum standard at which the modal feature has effective reference value. Meanwhile, real-time environmental conditions mainly refer to factors that may affect the quality of modal features, such as underground light intensity, dust concentration, and equipment operating noise. For example, excessive dust concentration can lead to blurred visual features, low light can reduce image clarity, and excessive background noise from equipment operation can interfere with auditory feature recognition. When the application model detects that the confidence level of a certain modality is below a threshold (e.g., a sudden dust exceedance in the mine, causing the confidence level of visual features to drop to 0.3, below the preset threshold of 0.5), a weight adjustment mechanism will be automatically activated. Assuming the initial weights were 0.2 for visual features, 0.4 for auditory features, 0.4 for equipment parameter features, and 0.1 for text semantic features, the weight of visual features will be reduced to 0.05 after adjustment (to minimize their impact due to low confidence). Simultaneously, the weights of high-confidence modalities will be proportionally increased, such as raising the weights of auditory features and equipment parameter features to 0.45, while maintaining the weight of text semantic features at 0.05. This ensures that the total weight of all modalities remains 1, preventing weight imbalance. The adjustment process requires no manual intervention and is automatically completed by the application model based on real-time environmental data and confidence level changes. This flexibly addresses the interference of complex underground environments on modal features, preventing unreliable fusion results due to the failure of a single modality and ensuring the stability of risk identification.
[0047] Step 210: By comparing the weighted comprehensive confidence level with the preset event threshold, a risk event identification result is generated, wherein the identification result includes event type, location and severity.
[0048] In some embodiments, the weighted overall confidence level is the final value after the preliminary calculation in step 208 and the dynamic adjustment in step 209, which can comprehensively reflect the degree of support of each modal feature for the risk event; the preset event threshold is a judgment standard set according to the hazard level of different risk events in the coal mine (such as the threshold for general risk events is set to 0.6, and the threshold for major risk events is set to 0.8). The higher the threshold, the higher the rigor required for event identification. When the application model compares the overall confidence level with the preset event threshold, if the overall confidence level is lower than the threshold (e.g., 0.5 < 0.6), it is judged as a "risk-free event"; if it is higher than the threshold (e.g., 0.85 > 0.8), it is judged as an "abnormal risk event," and further specific identification results are generated: the event type is determined according to the risk scenario pointed to by each modal feature. For example, if auditory features (friction noise) and equipment parameter features (sudden increase in motor current) both point to "coal blockage on the conveyor belt," then the event type is marked as "coal blockage on the conveyor belt." The event location is determined based on the spatial coordinates after spatiotemporal alignment in step 103, such as the area corresponding to the No. 3 area of the underground coal conveyor belt discharge port (from the spatial coordinate matching results of the camera and the fiber optic sensing system). The severity of the event is judged comprehensively based on the equipment parameter deviation range and the level of feature confidence. For example, if the motor current exceeds the rated value by 20% and the overall confidence level is 0.85, it is judged as "moderate risk." If the current exceeds the rated value by 50% and the overall confidence level is 0.95, it is judged as "severe risk." All identification results will be stored in structured data format to facilitate the generation of subsequent alarms and decision-making solutions. The benefit of this step is that it can clearly and comprehensively output key information about risk events, providing clear basis for dispatch personnel to quickly grasp the situation and take action, avoiding delays caused by ambiguous information.
[0049] Step 211: Perform semantic association analysis between the identification results and the structured rules and unstructured text knowledge in the cloud knowledge base to generate a structured decision-making scheme containing root cause diagnosis and treatment suggestions, and upload it to the cloud industrial internet platform.
[0050] In some embodiments, the cloud-based knowledge base is a comprehensive knowledge system built upon coal mining industry data. The structured rules cover quantitative standards and operational procedures for coal mine safety supervision, such as explicit threshold rules like "if the water level in the pump room exceeds 1.5 meters, the standby pump must be started" and "if the current of the belt motor exceeds the rated value by 30% for 10 consecutive seconds, the machine must be shut down for inspection." These rules are stored in directly callable coded logic for easy matching. The unstructured text knowledge is information extracted from historical coal mine accident reports, equipment operation manuals, and expert experience records. This text has been transformed into semantic vectors that can be associated with the recognition results through semantic processing of a cloud-based multimodal large model, eliminating matching barriers caused by differences in textual expression.
[0051] When performing semantic association analysis, the system first converts the anomaly identification results (such as "Water accumulation in the pump room: current water level 1.8 meters, water accumulation rate 0.3 meters / minute, main pump operating normally") into semantic vectors in a unified format. Then, it compares the semantic vectors with the structured rules and unstructured text knowledge in the cloud knowledge base through a similarity calculation algorithm. For example, the features "water level 1.8 meters > 1.5 meters" and "water accumulation rate 0.3 meters / minute > 0.2 meters / minute" in the identification results match the semantic vector of "high water level + fast water accumulation rate requires dual pump linkage" in the structured rule with a matching degree of 95%. At the same time, it matches the semantic vector of the historical case "main pump normal + water accumulation acceleration = drainage pipe blockage" in the unstructured text with a matching degree of 92%. Based on these highly correlated results, the system conducts root cause diagnosis: combining information that the main water pump is operating normally to rule out pump malfunctions, and based on the high correlation with blockage cases, the root cause of the anomaly is determined to be "blockage of the main drainage pipe in the pump room"; subsequently, it generates handling suggestions, referring to the rules and experience of "blockage handling procedures" in the knowledge base, and sorts them according to execution priority as follows: "Suggestion 1 (Immediate Execution): Start two standby drainage pumps to control the water level below 1.2 meters; Suggestion 2 (Simultaneous Execution): Notify the maintenance team to bring pipe dredging tools to the site to locate the blockage; Suggestion 3 (Post-Handling): After dredging, check the wear of the inner wall of the drainage pipe and update the equipment maintenance file," forming a structured decision-making scheme that includes "root cause conclusion + priority suggestions + operation steps + time limit requirements." Finally, the scheme is uploaded to the cloud-based industrial internet platform through an encrypted data channel. The platform categorizes and archives the scheme according to "equipment type - anomaly type" (e.g., classified as "pump room - water accumulation blockage"), and records the scheme generation time and related knowledge source, which facilitates subsequent traceability and adds new practical cases to the knowledge base, optimizing the accuracy of subsequent correlation analysis.
[0052] With the support of multi-dimensional knowledge from the cloud-based knowledge base, decision-making solutions can not only respond quickly to anomalies, but also accurately pinpoint the root cause and provide actionable steps, avoiding unfounded subjective suggestions. Uploading to the platform enables the accumulation and reuse of knowledge, improving the standardization of anomaly handling in coal mines as a whole.
[0053] Figure 3 This is a schematic diagram of the structure of a security risk monitoring device provided in an embodiment of this application, as shown below. Figure 3 As shown, it includes: a construction module 301, an extraction module 302, an alignment module 303, a first generation module 304, and a second generation module 305.
[0054] Module 301 is configured to build a collaborative architecture of cloud layer, edge layer, and device layer. The cloud layer deploys visual large model and multimodal large model for pre-training and industry fine-tuning. The edge layer deploys application model and multimodal information fusion algorithm for data analysis, data fusion, and event reasoning. The device layer deploys multimodal data acquisition devices to acquire multimodal data. The extraction module 302 is configured to perform cross-modal feature extraction on multimodal data to generate multi-dimensional features including visual features, auditory features, device parameter features and text semantic features; Alignment module 303 is configured to accurately match timestamps and spatial coordinates of multi-dimensional features based on a multi-modal information fusion algorithm to achieve spatiotemporal alignment; The first generation module 304 is configured to dynamically weight and fuse the spatiotemporally aligned multi-dimensional features based on the application model, and adjust the fusion weights according to the confidence of each modality to generate risk event identification results. The second generation module 305 is configured to trigger an alarm and generate a decision plan and push it to the scheduling terminal when the identification result is abnormal.
[0055] In some examples of this embodiment, the construction module 301 is specifically configured to use the Transformer architecture for the visual large model and the multimodal large model. It is used to perform self-training based on visual data, text data, and structured data, and to fine-tune based on industry feature data to generate application models with visual large model and multimodal large model functions. The application models are pushed to the edge layer through the cloud-edge collaboration mechanism. The multimodal data acquisition equipment includes cameras deployed at key points in the coal mine site, as well as fiber optic acoustic temperature sensing system, employee vital sign acquisition device, liquid level / flow sensor and PLC equipment.
[0056] In some examples of this embodiment, the extraction module 302 is specifically configured to extract visual features from the video stream through an application model, extract text semantic features from text data, the visual features including object detection, image segmentation, or scene understanding features, and the text data including equipment operation manuals, safety procedures, or real-time event descriptions, and the text semantic features are generated into vector representations through natural language processing technology; extract auditory features from audio data through a fiber optic acoustic signature sensing system, the auditory features including acoustic signature patterns, abnormal sounds, or frequency features; and extract equipment parameter features from the equipment operating parameters output by the PLC device, the equipment parameter features including current, voltage, temperature, or flow time-series data.
[0057] In some examples of this embodiment, the alignment module 303 is specifically configured to use a spatiotemporal synchronization algorithm to control the synchronization error between the video frame timestamp and the device current data timestamp within a preset time threshold; and to use feature space mapping technology to perform three-dimensional coordinate transformation between the camera coordinate system and the spatial coordinates of the fiber optic acoustic signature sensing system to ensure that the spatial position error is controlled within a preset spatial range.
[0058] In some examples of this embodiment, the first generation module 304 is specifically configured to use the application model to dynamically weight the confidence of visual features, auditory features, device parameter features and text semantic features based on cross-modal association weights; adjust the weight of each modal feature according to real-time environmental conditions; when the confidence of a certain modal feature is lower than a preset threshold, increase the weight of other high-confidence modal features; and generate a risk event identification result by comparing the weighted comprehensive confidence with a preset event threshold, wherein the identification result includes event type, location and severity.
[0059] In some examples of this embodiment, the second generation module 305 is specifically configured to perform semantic association analysis between the recognition results and the structured rules and unstructured text knowledge in the cloud knowledge base, generate a structured decision scheme containing root cause diagnosis and treatment suggestions, and upload it to the cloud industrial internet platform.
[0060] It should be noted that other corresponding descriptions of the functional units involved in the security risk monitoring device provided in this embodiment can be found in [reference needed]. Figure 1 , Figure 2 The corresponding description in [the document] will not be repeated here.
[0061] Based on the above, Figure 1 , Figure 2 The embodiment illustrates a security risk monitoring method. Correspondingly, this embodiment also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the aforementioned method. Figure 1 , Figure 2 This illustrates a method for monitoring security risks.
[0062] Based on the above, Figure 1 , Figure 2 The embodiment illustrates a security risk monitoring method. Correspondingly, this embodiment also provides a computer program product storing a computer program, which, when executed by a processor, implements the aforementioned... Figure 1 , Figure 2 This illustrates a method for monitoring security risks.
[0063] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.
[0064] Based on the above, Figure 1 , Figure 2 A method for monitoring security risks, and Figure 3 To achieve the above objectives, the present application also provides an electronic device, such as a personal computer or a server, in the illustrated virtual device embodiment. This device includes a storage medium and a processor; the storage medium stores a computer program; the processor executes the computer program to implement the above-described virtual device. Figure 1 , Figure 2 This illustrates a method for monitoring security risks.
[0065] In some embodiments, the aforementioned physical device may further include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, an input unit such as a keyboard, etc., and optionally, a USB interface, a card reader interface, etc. In some embodiments, the network interface may include a standard wired interface, a wireless interface (such as a Wi-Fi interface), etc.
[0066] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.
[0067] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0068] The above are merely specific embodiments of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to these embodiments, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for monitoring safety risks, characterized in that, include: A collaborative architecture is constructed, consisting of a cloud layer, an edge layer, and a device layer. The cloud layer deploys large visual models and multimodal models for pre-training and industry-specific fine-tuning. The edge layer deploys application models and multimodal information fusion algorithms for data analysis, data fusion, and event reasoning. The device layer deploys multimodal data acquisition devices to acquire multimodal data. Cross-modal feature extraction is performed on the multimodal data to generate multi-dimensional features that include visual features, auditory features, device parameter features, and text semantic features; Based on the multimodal information fusion algorithm, the timestamps and spatial coordinates of the multi-dimensional features are accurately matched to achieve spatiotemporal alignment. Based on the application model, the multi-dimensional features after spatiotemporal alignment are dynamically weighted and fused, and the fusion weights are adjusted according to the confidence of each modality to generate risk event identification results. When the identification result is abnormal, an alarm is triggered, a decision plan is generated, and it is pushed to the scheduling terminal.
2. The safety risk monitoring method according to claim 1, characterized in that, The construction of the collaborative architecture of cloud layer, edge layer, and device layer includes: The visual large model and the multimodal large model adopt the Transformer architecture, which is used to perform self-training based on visual data, text data, and structured data, and fine-tuning based on industry feature data to generate an application model with the functions of the visual large model and the multimodal large model, and push the application model to the edge layer through the cloud-edge collaboration mechanism. The multimodal data acquisition equipment includes cameras deployed at key locations in the coal mine, as well as a fiber optic acoustic temperature sensing system, an employee vital signs acquisition device, a liquid level / flow sensor, and a PLC device.
3. The safety risk monitoring method according to claim 1, characterized in that, The process of extracting cross-modal features from the multimodal data to generate multi-dimensional features including visual features, auditory features, device parameter features, and textual semantic features includes: The application model extracts visual features from video streams and textual semantic features from text data. The visual features include object detection, image segmentation, or scene understanding features. The text data includes device operation manuals, safety procedures, or real-time event descriptions. The textual semantic features are generated into vector representations using natural language processing techniques. Auditory features are extracted from audio data using a fiber optic acoustic signature sensing system. These auditory features include acoustic signature patterns, abnormal sounds, or frequency characteristics. Extract equipment parameter features from the equipment operating parameters output by the PLC device. These equipment parameter features include time-series data of current, voltage, temperature, or flow rate.
4. The safety risk monitoring method according to claim 1, characterized in that, The step of accurately matching the timestamps and spatial coordinates of the multi-dimensional features based on the multi-modal information fusion algorithm includes: A time-space synchronization algorithm is used to control the synchronization error between the video frame timestamp and the device current data timestamp within a preset time threshold. By using feature space mapping technology, the camera coordinate system and the spatial coordinates of the fiber optic acoustic signature sensing system are transformed into three-dimensional coordinates to ensure that the spatial position error is controlled within a preset spatial range.
5. The safety risk monitoring method according to claim 1, characterized in that, The process of dynamically weighting and fusing multi-dimensional features aligned to spatiotemporal levels based on the application model, and adjusting the fusion weights according to the confidence levels of each modality to generate risk event identification results, includes: The application model is used to dynamically weight the confidence of visual features, auditory features, device parameter features and text semantic features based on cross-modal association weights. The weights of each modality feature are adjusted according to real-time environmental conditions. When the confidence of a certain modality feature is lower than a preset threshold, the weights of other high-confidence modality features are increased. By comparing the weighted overall confidence level with a preset event threshold, a risk event identification result is generated, wherein the identification result includes the event type, location, and severity.
6. The safety risk monitoring method according to claim 1, characterized in that, When the identification result is abnormal, triggering an alarm and generating a decision plan and pushing it to the scheduling terminal includes: The identification results are semantically correlated with the structured rules and unstructured text knowledge in the cloud knowledge base to generate a structured decision scheme containing root cause diagnosis and treatment suggestions, which is then uploaded to the cloud industrial internet platform.
7. A safety risk monitoring device, characterized in that, include: The building module is configured to construct a collaborative architecture of cloud layer, edge layer, and device layer. The cloud layer deploys large visual models and multimodal large models for pre-training and industry fine-tuning. The edge layer deploys application models and multimodal information fusion algorithms for data analysis, data fusion, and event reasoning. The device layer deploys multimodal data acquisition devices to acquire multimodal data. The extraction module is configured to perform cross-modal feature extraction on the multimodal data to generate multi-dimensional features including visual features, auditory features, device parameter features and text semantic features; The alignment module is configured to accurately match the timestamps and spatial coordinates of the multi-dimensional features based on the multi-modal information fusion algorithm to achieve spatiotemporal alignment. The first generation module is configured to dynamically weight and fuse the spatiotemporally aligned multi-dimensional features based on the application model, and adjust the fusion weights according to the confidence of each modality to generate risk event identification results. The second generation module is configured to trigger an alarm and generate a decision scheme and push it to the scheduling terminal when the identification result is abnormal.
8. An electronic device, characterized in that, include: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the security risk monitoring method according to any one of claims 1-6.
9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the security risk monitoring method according to any one of claims 1-6.
10. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the security risk monitoring method according to any one of claims 1-6.