Safety management and control method, device and equipment for industrial and mining enterprises and storage medium

By integrating feature extraction and scene matching models from video, audio, and environmental data, a safety hazard detection report is generated, solving the problems of low efficiency and insufficient accuracy of traditional methods in safety management of industrial and mining enterprises, and achieving highly accurate safety control in complex industrial and mining scenarios.

CN122200481APending Publication Date: 2026-06-12SHANXI ZHIJIAN ENERGY TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANXI ZHIJIAN ENERGY TECHNOLOGY CO LTD
Filing Date
2026-03-02
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

In the safety management of industrial and mining enterprises, existing technologies suffer from low efficiency in manual inspections and the inability of single sensor data to comprehensively and accurately identify safety hazards, resulting in insufficient precision in safety control.

Method used

By employing a multimodal data fusion method, feature extraction is performed on video, audio, and environmental data, and combined with a hazard detection model that matches scene type, a safety hazard detection report is generated, enabling precise control over the regulated area.

Benefits of technology

It improves the accuracy of safety hazard detection and control in industrial and mining enterprises, overcomes the problems of missed and incorrect detection by traditional single sensors, and adapts to the specific risks of different operating scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122200481A_ABST
    Figure CN122200481A_ABST
Patent Text Reader

Abstract

The present disclosure provides a safety management method, device and equipment for industrial and mining enterprises and a storage medium, which comprises the following steps: acquiring video data, audio data, environmental data and scene types of a supervision area; performing feature extraction on the video data, audio data and environmental data respectively to obtain video features, audio features and environmental features; generating a safety hazard detection report of the supervision area by using a pre-constructed safety hazard detection model based on the video features, audio features, environmental features and hazard detection prompt words matched with the scene types; and performing safety management and control on the supervision area based on the safety hazard detection report. The method of the present disclosure can improve the accuracy of safety management and control of industrial and mining enterprises.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a safety management method, apparatus, equipment and storage medium for industrial and mining enterprises. Background Technology

[0002] The working environment in mining and industrial sectors is complex and ever-changing, characterized by darkness, high dust concentrations, and a wide variety of complex work scenarios, posing a severe challenge to safe production. Currently, safety management at mining and industrial sites primarily relies on manual inspections and single-sensor monitoring. These traditional safety management methods have several limitations: First, manual inspections are inefficient, have limited coverage, and are easily influenced by the experience and sense of responsibility of the inspectors, making it difficult to comprehensively and accurately identify safety hazards. Second, a single sensor (such as traditional video surveillance) can only acquire a single type of data, which cannot accurately identify safety hazards in multiple scenarios, easily leading to missed or incorrect diagnoses. Therefore, how to improve the accuracy of safety hazard identification in different work scenarios, and thus enhance the precision of safety management in mining and industrial enterprises, is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0003] In view of this, this disclosure proposes a safety management method, device, equipment and storage medium for industrial and mining enterprises, which can improve the accuracy of safety management in industrial and mining enterprises.

[0004] According to a first aspect of this disclosure, a safety management method for industrial and mining enterprises is provided, comprising: Acquire video data, audio data, environmental data, and scene types within the monitored area; Feature extraction is performed on the video data, audio data, and environmental data respectively to obtain video features, audio features, and environmental features; Based on the video features, audio features, environmental features, and hazard detection prompts matching the scene type, a pre-built safety hazard detection model is used to generate a safety hazard detection report for the monitored area. Based on the aforementioned safety hazard detection report, safety management and control measures are implemented in the supervised area.

[0005] In one possible implementation, when performing feature extraction on the video data to obtain the video features, the process includes: Video segments are extracted from the video data based on a preset time window; Keyframes are extracted from the video segment using a first frame extraction algorithm that matches the scene type. Extract the image features of each key frame, and generate the video features based on the image features of each key frame.

[0006] In one possible implementation, extracting the image features of the keyframe includes: Calculate the acquisition timestamp of the key frame; Target detection is performed on the keyframe to obtain the target detection result of the keyframe; Based on the acquisition timestamp and the target detection results, the image features of the keyframe are generated.

[0007] In one possible implementation, generating the video features based on the image features of each of the keyframes includes: Extract the image features of each of the key frames; Based on the image features of each key frame, determine whether the preset associated frame extraction conditions are met; If the conditions for extracting the associated frames are not met, the image features of each key frame are combined in sequence to generate the video features.

[0008] In one possible implementation, if the associated frame extraction condition is met, the following operations are performed: Invoke the second frame extraction algorithm that matches the associated frame extraction conditions; The second frame extraction algorithm is used to extract the associated frames of each key frame and extract the image features of each associated frame. The video features are obtained by sequentially combining the image features of each key frame and the image features of each associated frame.

[0009] In one possible implementation, the construction of the security hazard detection model includes: Acquire and construct training data based on historical video data, historical audio data, historical environmental data of the regulated area, and corresponding scene types; The training data is used to sequentially train each sub-task of the security hazard detection on the pre-loaded large language model to obtain the security hazard detection model.

[0010] In one possible implementation, the DoRA fine-tuning algorithm is used to sequentially train each subtask of the security hazard detection on a pre-loaded large language model using the training data.

[0011] According to a second aspect of this disclosure, a safety control device for industrial and mining enterprises is provided, comprising: The data acquisition module is used to acquire video data, audio data, environmental data, and scene types of the monitored area; The feature extraction module is used to extract features from the video data, audio data, and environmental data respectively to obtain video features, audio features, and environmental features. The hazard detection module is used to generate a safety hazard detection report for the supervised area based on the video features, the audio features, the environmental features, and hazard detection prompts that match the scene type, using a pre-built safety hazard detection model. The safety management module is used to manage and control the safety of the supervised area based on the safety hazard detection report.

[0012] According to a third aspect of this disclosure, a safety management and control device for industrial and mining enterprises is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the method described in the first aspect of this disclosure.

[0013] According to a fourth aspect of this disclosure, a non-volatile computer-readable storage medium is provided that stores computer program instructions thereon, wherein the computer program instructions, when executed by a processor, implement the method described in the first aspect of this disclosure.

[0014] This disclosure provides a safety management method, device, equipment, and storage medium for industrial and mining enterprises. The method includes: acquiring video data, audio data, environmental data, and scene types of a monitored area; extracting features from the video data, audio data, and environmental data to obtain video features, audio features, and environmental features; generating a safety hazard detection report for the monitored area using a pre-built safety hazard detection model based on the video features, audio features, environmental features, and hazard detection prompts matched with the scene type; and conducting safety management and control of the monitored area based on the safety hazard detection report. This disclosure achieves cross-validation and complementarity of multimodal information by integrating video, audio, and environmental data, overcoming the shortcomings of traditional single-sensor data being one-sided and prone to missed detections. Simultaneously, by combining a prompt-guided model matched with scene types for targeted analysis, hazard identification can adapt to the specific risks of different work scenarios, thereby significantly improving the accuracy of safety hazard detection and the precision of management in complex industrial and mining scenarios.

[0015] Other features and aspects of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0016] The accompanying drawings, which are included in and form part of this specification, illustrate exemplary embodiments, features, and aspects of this disclosure together with the specification and serve to explain the principles of this disclosure.

[0017] Figure 1A flowchart is shown illustrating a safety management method for industrial and mining enterprises according to an embodiment of the present disclosure; Figure 2 A schematic block diagram of a safety control device for industrial and mining enterprises according to an embodiment of the present disclosure is shown. Figure 3 A schematic block diagram of a safety control device for industrial and mining enterprises according to an embodiment of the present disclosure is shown. Detailed Implementation

[0018] Various exemplary embodiments, features, and aspects of this disclosure will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.

[0019] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.

[0020] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.

[0021] <Method Implementation> Figure 1 A flowchart illustrating a safety management method for industrial and mining enterprises according to an embodiment of this disclosure is shown. Figure 1 As shown, the method includes steps S1100-S1400.

[0022] S1100: Obtain video data, audio data, environmental data, and scene type of the monitored area. The video data includes at least one of the following: video ID, storage location ID, and start and end timestamps.

[0023] It should be noted that the mining site is divided into multiple monitoring zones, each equipped with high-definition explosion-proof cameras, microphones, and environmental sensors to support the collection of video, audio, and environmental data. The cameras collect video data to record worker actions, equipment operation status, and material placement. The microphones collect audio data to record conversations between workers and abnormal noises from equipment. The environmental sensors collect environmental data such as temperature, humidity, dust concentration, and toxic gas levels. The number and location of the cameras, microphones, and environmental sensors in each monitoring zone are configured according to the specific work scenario and are not specifically limited here.

[0024] Furthermore, each monitored area corresponds to a scenario type, which reflects the nature of the work performed within the monitored area. For example, this scenario type may include at least one of the following: hot work, coal mining, tunneling, working at height, belt conveyor, wearing protective gear, personnel sleeping on duty, roadway operation, sensor-related work, equipment placement work, and confined space work. The system implementing the method of this disclosure records the mapping relationship between each monitored area and its corresponding scenario type. Thus, when conducting safety management of a monitored area, its corresponding scenario type can be directly obtained by querying the mapping relationship. S1200 extracts features from video data, audio data, and environmental data respectively, to obtain video features, audio features, and environmental features.

[0025] In one possible implementation, the process of extracting features from video data to obtain video features may include the following steps: First, video segments are extracted from the video data based on a preset time window. The size of the time window can be configured according to specific needs; preferably, the time window can be set to 10 seconds, that is, a video segment is cut every 10 seconds.

[0026] Second, keyframes are extracted from video clips using a first frame extraction algorithm that matches the scene type.

[0027] First, it's important to note that the distribution density of key information in video data varies significantly across different scenarios. For example, in some scenarios, operations are primarily completed by automated equipment, whose operating status changes relatively slowly and predictably. Key information (such as abnormal equipment vibration or parts falling off) may only become apparent over extended periods. In such cases, using a high-frequency frame extraction algorithm would result in a large number of redundant frames, increasing the computational load for subsequent feature extraction and analysis, and reducing system processing efficiency. Conversely, in other scenarios, operations are mainly performed by operators, whose actions are often rapid and complex. Key information (such as whether a safety helmet is worn or whether equipment is being operated improperly) may occur instantaneously. If the frame extraction frequency is too low, these critical actions may be missed, leading to undetected safety hazards. Therefore, the system pre-configures a first frame extraction algorithm suitable for different scenario types. Once the scenario type is identified, the corresponding first frame extraction algorithm can be used to extract key frames from video segments. This ensures complete capture of key information while minimizing redundant data, thereby improving the processing efficiency and analysis accuracy of video data by the safety management system.

[0028] In one possible implementation, the first frame extraction algorithm corresponding to each scene type can be a single frame extraction algorithm or a hybrid frame extraction algorithm (i.e., combining multiple frame extraction algorithms for frame extraction). The specific choice between using a single frame extraction algorithm or a hybrid frame extraction algorithm depends on the specific scene.

[0029] Third, extract the image features of each keyframe and generate video features based on these features. Since the image feature extraction method is the same for all keyframes, the following section uses a single keyframe as an example to explain the image feature extraction process in detail.

[0030] In one possible implementation, the process of extracting image features from keyframes may include the following steps: First, the capture timestamps of the keyframes are calculated. Specifically, if the keyframe includes a capture timestamp, the timestamp is directly extracted from the keyframe using an OCR recognition algorithm. If the keyframe does not include a capture time, the capture timestamp for each keyframe is calculated based on the start and end timestamps in the video data and the video frame rate.

[0031] Secondly, target detection is performed on the keyframes to obtain the target detection results. Specifically, the keyframes are input into a pre-loaded target detection model for target detection, and the bounding boxes and entity labels of each entity target included in the keyframes are identified. The target detection results of the keyframes are obtained by sequentially combining the bounding boxes and entity labels of each entity target.

[0032] Finally, based on the acquisition timestamp and object detection results, image features of keyframes are generated. Specifically, the image features of keyframes are obtained by sequentially combining the acquisition timestamp and object detection results.

[0033] In one feasible implementation model, to improve the accuracy of image feature extraction, image preprocessing of the keyframes is required before performing keyframe image feature extraction. Specifically, image preprocessing of keyframes includes the following: performing local adaptive histogram equalization on the keyframes to enhance contrast in dark areas; performing dehazing / dust removal on keyframes in areas with severe fog or dust to effectively reduce image noise and improve image clarity; and applying at least one correction algorithm from white balance correction, gamma correction, and tone correction to improve the accuracy of subsequent image feature extraction.

[0034] Following the method described above, image features of each keyframe can be extracted. After obtaining the image features of the keyframes, video features can be generated based on these features. Specifically, this can include the following steps: First, extract the image features of each keyframe.

[0035] Secondly, based on the image features of each keyframe, it is determined whether the preset conditions for extracting related frames are met.

[0036] It's important to note that meeting the association frame extraction conditions indicates that the currently extracted keyframes are insufficient to support the identification of certain specific security vulnerabilities in that scene type, while not meeting the conditions indicates that the currently extracted keyframes can support the identification of certain specific security vulnerabilities in that scene type. Different scene types have different association frame extraction conditions. Therefore, the system pre-configures different association frame extraction conditions for each scene type. After extracting the image features of each keyframe, the system can obtain the association frame extraction conditions matching its scene type, and then determine whether the currently extracted keyframes meet these conditions based on the image features of each keyframe. These keyframe extraction conditions are set based on the entity labels of each entity target identified in the keyframe and the association between the bounding boxes of each entity target.

[0037] In embodiments where the scenario types include hot work, coal mining, tunneling, high-altitude work, belt conveyor, wearing protective gear, personnel sleeping on duty, roadway driving, sensor-related, and equipment placement, the first frame extraction algorithm and the corresponding associated frame extraction condition information pre-configured for each scenario type are shown in Table 1. The associated frame extraction condition information includes the number of associated frame extraction conditions under that scenario type and the specific associated frame extraction conditions.

[0038] Table 1 Finally, if the conditions for extracting associated frames are not met, the image features of each key frame are combined in sequence to generate video features.

[0039] However, if the conditions for extracting associated frames are met, the following associated frame extraction operation still needs to be performed, as detailed below: First, the second frame extraction algorithm that matches the associated frame extraction conditions is retrieved. Specifically, for each pre-set associated frame extraction condition, a corresponding second frame extraction algorithm is configured, so that the corresponding second frame extraction algorithm can be retrieved when the associated frame extraction conditions are met.

[0040] The corresponding second frame extraction algorithm configured for each associated frame extraction condition under each scene type in the above embodiments can be shown in Table 2.

[0041] Table 2 Second, a second frame extraction algorithm is used to extract the associated frames of each keyframe and extract the image features of each associated frame. The method for extracting the image features of each associated frame is the same as that for the keyframes, and will not be described again here.

[0042] Third, the image features of each keyframe and the image features of each associated frame are combined sequentially to obtain video features.

[0043] The extraction of the aforementioned associated frames can effectively compensate for the lack of key frame information when identifying specific security risks, thereby improving the accuracy of identifying specific security risks.

[0044] In one possible implementation, when extracting features from audio data to obtain audio features, the following steps may be included: First, the waveform signal of the audio data is converted from the time domain to the frequency domain; then, the Mel-frequency cepstral coefficients are calculated based on the frequency domain representation to obtain the MFCC features represented in the form of a continuous, dense floating-point vector; finally, the extracted MFCC features are used as the audio features corresponding to the audio data.

[0045] In one possible implementation, when extracting features from environmental data to obtain environmental features, the following steps may be included: combining the names and values ​​of each environmental parameter (such as gas concentration: [numerical value]) in sequence to obtain the environmental features.

[0046] In one possible implementation, after video feature extraction, the acquisition time corresponding to each frame of the video feature is calculated based on its timestamp. Subsequently, based on the acquisition time of the video feature, the acquisition time of the video feature, audio feature, and environmental feature is synchronized. This ensures that the video feature, audio feature, and environmental feature reflect features within the same time period, thereby guaranteeing high consistency of multimodal data in the temporal dimension. This provides accurately matched input data support for the subsequent fusion analysis of the safety hazard detection model, avoids detection errors caused by temporal misalignment of different modal features, and ultimately enhances the model's comprehensive identification efficiency for safety hazards in complex industrial and mining scenarios.

[0047] The S1300, based on video features, audio features, environmental features, and hazard detection prompts matched with scene types, uses a pre-built safety hazard detection model to generate a safety hazard detection report for the regulated area.

[0048] First, it's important to note that the safety hazard inspection tasks and detection criteria differ across scenarios. Therefore, the system pre-configures different hazard detection prompts for different scenario types. Each scenario's prompt includes at least one of the following: the definition of the safety hazard detection task to be performed, the safety regulations and operating procedures to be followed, the specific detection tasks to be performed, the judgment criteria for each task, and the relevant fields, output format requirements, and examples of the detection results. Thus, after obtaining the scenario type of the regulated area, the system can automatically retrieve the hazard detection prompts matching that scenario type. Furthermore, the specific safety hazard detection tasks, the corresponding safety regulations, operating procedures, judgment criteria, and relevant fields of the detection results can all be manually adjusted (including adding, deleting, or modifying) through the system's configuration interface.

[0049] By configuring different hazard detection prompts for different scenario types, the safety hazard detection model can accurately focus on key inspection tasks in specific scenarios, thereby significantly improving the accuracy and scenario adaptability of safety hazard detection and meeting the differentiated safety management needs of different industrial and mining enterprises.

[0050] In another possible implementation, the safety detection tasks to be performed in each scenario, the safety regulations and operating procedures to be followed, the specific detection tasks to be performed, and the judgment criteria corresponding to each detection task can be integrated. Simultaneously, the relevant fields and output format requirements of the detection results of each safety monitoring task are unified, resulting in a comprehensive hazard detection prompt that can adapt to multiple scenario types. This allows for the detection of safety hazards in various different scenario types using a single comprehensive hazard detection prompt, eliminating the need to configure prompts separately for each scenario, simplifying the system configuration process, and reducing manual maintenance costs. This comprehensive hazard detection prompt can automatically match and activate the corresponding scenario's detection logic based on the input scenario type information, including the relevant safety hazard detection task definitions, safety regulations and operating procedures, and specific judgment criteria. It also ensures that the detection results of all scenarios are generated according to a unified output format requirement, guaranteeing both the comprehensiveness and accuracy of the detection, improving the consistency and comparability of detection results across different scenarios, and further enhancing the versatility and practicality of the safety hazard detection model in complex industrial and mining enterprise environments.

[0051] In one specific embodiment, a comprehensive hazard detection prompt in Python format adapted to multiple scenarios can be as follows: You are an intelligent safety inspector for an industrial setting. Please strictly adhere to the provided video, audio, and environmental features, and follow pre-configured safety regulations and operating procedures to perform the following identification and judgment tasks. Guessing is prohibited: ### I. User Configuration Fields: Prior Knowledge Input (Specific constraints entered by the user during online requests, including task definitions for each adaptation scenario and the security regulations and operating procedures to be followed) - Task definitions: [{user_defined_criteria}, {user_defined_task}] - Safety regulations and operating procedures: [{rules_1}, {rules_2}] ### II. Identification of Key Scenarios in Industrial and Mining Operations (If multiple scenarios are identified, output the safety hazard detection results for all scenarios) 1. On-site personnel safety helmet and reflective vest identification - Judgment criteria: If there are personnel in the picture, they must be wearing safety helmets and reflective clothing. - Task: Identify the location and clothing status of all personnel.

[0052] 2. hoisting operation scene recognition - Judgment criteria: The crane, hook, gantry, pulley block and other lifting equipment are directly visible in the image. - Task: Determine whether it is a hoisting operation, and identify the work area, personnel location, safety helmet status, and whether there is any unauthorized approach.

[0053] 3. Tunneling operation scene recognition - Judgment criteria: The image shows equipment such as tunneling machines, drilling rigs, and crushers breaking rocks or removing slag. - Task: Determine whether it is a tunneling operation, mark the positions of the driver and surrounding personnel, and identify the operation status.

[0054] 4. Identification of violations of safety distance during hot work - Judgment criteria: The image shows open flame sources such as electric welding, gas cutting, and blowtorches, as well as flammable materials such as cables, wood, and oil drums. - Task: If a hot work source and flammable materials are present at the same time, the relative distance between them must be calculated based on visual positioning to determine whether it violates Article 328 of the Coal Mine Safety Regulations (safe distance ≥ 5 meters).

[0055] 5. User-defined scene recognition - Judgment criteria: {user_defined_criteria}. - Task: {user_defined_task}.

[0056] ### III. Safety Hazard Detection, Judgment, and Handling Recommendations (Final generated template, strictly including the following fields related to the detection results) { "The image contains the scene": "", "Whether there are any behaviors that pose a risk to workplace safety": "", "Violation type and violation description": "", "Referencing specific rule entries": "", Risk Level: "", "Rectification Plan": "", "Rectification Measures": "", Solution: "", "Location information of the work area": ​​"", "Is there a risk of insufficient safe distance?": "", "Distance estimation between targets and the distance to be maintained": "", } ### IV. Output Format Requirements - You must reply in the form of a JSON object only; - Strictly adhere to the above-mentioned fields related to the test results; - All judgments must be based on clear visual evidence, i.e., the image contains the scene. Fill in "yes", "no" or "uncertain"; - Location coordinates explanation: [x1, y1, x2, y2] are the 0-1 normalized coordinates of the image, where x1 and y1 represent the top left corner of the target, and x2 and y2 represent the bottom right corner of the target.

[0057] ### V. Output Structure Example { "Image contains the scene": "Yes / No / Uncertain", "Are there any safety production risks?": "Yes", "Type and description of violation": "Welding was not carried out within 20 meters of the wellhead or underground; open flames were not used in flammable and explosive locations underground or on the surface." "Specific regulation cited": "Article 197 of the Coal Mine Safety Regulations: Electric welding, gas welding, and blowtorch welding are prohibited underground and in the mine shaft. If electric welding, gas welding, and blowtorch welding must be carried out in the main underground chambers, main intake airways, and mine shaft, safety measures must be formulated each time, approved by the mine's chief engineer, and relevant regulations must be followed." Risk Level: High "Immediate Rectification Plan": "Immediately cease all hot work operations and investigate potential safety hazards on site." "Rectification Measures": [ 1. Immediately cease all hot work operations, both underground and on the surface. 2. Conduct a comprehensive safety hazard inspection of the site to ensure that there are no flammable or explosive materials. 3. Develop and approve specific safety measures, including but not limited to fire isolation measures, ventilation measures, and the provision of fire extinguishing equipment. 4. Organize safety training for relevant personnel, emphasizing safety regulations and operating procedures for hot work. 5. Arrange for professional personnel to conduct regular inspections of the hot work area to ensure that safety measures are implemented effectively. 6. No hot work of any kind is permitted without taking effective safety measures. ], "Handling Plan": "Implement internal criticism of the violation, educate and punish those responsible, and ensure that similar incidents do not occur again." "Location information of the working face area": ​​"Within 20 meters of the wellhead, underground and surface flammable and explosive locations", "Is there a risk of insufficient safe distance?": "Yes", "Distance estimation between targets and required distances": "According to the 'Coal Mine Safety Regulations,' sufficient safety distances should be maintained in flammable and explosive areas underground and on the surface. The specific distance needs to be determined based on the actual site conditions and safety measures, but it should at least meet fire isolation requirements. It is generally recommended to maintain a safety distance of at least 20 meters." } After obtaining the hazard detection prompts that match the scene type, the extracted video features, audio features, environmental features, and hazard detection prompts that match the scene type can be input into the pre-built safety hazard detection model. The safety hazard model can then generate a safety hazard detection report for the regulatory area according to the requirements of the hazard prompts.

[0058] It should be further noted that the safety hazard detection model needs to be built in advance before performing this step. The construction of the safety hazard detection model may include the following steps: First, acquire and utilize multiple historical data points from the regulated area, and construct training data based on these historical data. Each historical data point includes historical video data, historical audio data, historical environmental data, and the corresponding scene type. For each historical data point: extract video features, audio features, and video and audio features as described above; obtain hazard detection prompts matching the scene type; pre-label the corresponding safety hazard detection reports; use the video features, audio features, and hazard detection prompts as input, and the labeled safety hazard detection reports as the standard output, thus forming a training data point. See the method described above for processing each historical data point to obtain the training data.

[0059] Second, the pre-loaded large language model is used to train each sub-task of the security hazard detection in sequence using training data to obtain the security hazard detection model.

[0060] It should be noted that, in order to output the aforementioned safety hazard detection report, the pre-loaded large language model is configured with the following sub-tasks: visual basic perception (used to filter out visual features with potential safety hazards from the input visual features), distance perception calculation (used to calculate the relative distance between entities involved in the visual features with potential safety hazards), rule knowledge injection (used to automatically match relevant regulations for visual features with potential safety hazards to provide accurate basis for hazard judgment), contextual comprehensive reasoning (used to jointly reason with the visual features with potential safety hazards, the relative distance between entities related to the safety hazards, the input audio features and environmental features, and the injected rule knowledge to obtain the relevant field values ​​of the detection results specified in the prompt words), and result output (used to output the final safety hazard detection report according to the output format requirements and examples in the prompt words). When training the pre-loaded large language model, the execution capabilities of each sub-task are trained sequentially using the training data constructed above, thereby obtaining the final safety hazard detection model. This gradual fine-tuning allows for a progression from simple to complex, from perception to reasoning, guiding the model to learn steadily and avoiding early interference from complex tasks, thus improving the model's convergence stability and final performance.

[0061] For example, in an embodiment where the large model includes sub-tasks such as basic visual perception, distance perception calculation, rule knowledge injection, contextual reasoning, and result output, the basic visual perception task of the large model is first trained using training data. Training of the basic visual perception task is stopped when the loss corresponding to this task decreases to a preset expected value. Next, the distance perception calculation task is trained using training data. Training of the distance perception calculation task is terminated when the loss corresponding to this task decreases to a preset expected value. Then, the rule knowledge injection task is trained using training data until the loss corresponding to this task decreases to a preset expected value, at which point training of this task is terminated. This process is repeated until the training of the subsequent contextual reasoning and result output tasks is completed, thus obtaining the final security vulnerability detection model.

[0062] In one possible implementation, the specific calculation steps for distance perception calculation of the large model are as follows: An object of known size (such as a standard-sized fire extinguisher or equipment sign) is fixed in the scene as a scale; target pixel distance calculation: the large model identifies target A (e.g., a ignition point), target B (e.g., an oil drum), and a reference object; finally, by calculating the pixel distance between targets A and B and comparing it with the pixel size of the reference object, the relative distance between the two is estimated. The formula simplifies to: Actual distance ≈ (AB pixel distance / reference object pixel width) The actual width of the reference object.

[0063] In one possible implementation, when training the pre-loaded large language model sequentially to perform each sub-task of security vulnerability detection using training data, a staged DoRA fine-tuning algorithm is employed. Specifically, for each network layer in the large model that performs the sub-task, its corresponding pre-trained weights W0 are decomposed into magnitude m and direction V. The input-output expressions of this network layer after decomposition are shown below: In the formula, Let h be the input to the network layer, h be the output of the network layer, W0 be the frozen pre-trained weights of the network layer, and m be the magnitude of the pre-trained weights W0 after decomposition. V represents the direction after the pre-trained weights W0 are decomposed. Let A be the product of two learnable low-rank matrices B and A. This indicates that each row of the matrix is ​​independently normalized in the column direction (specifically, L2 normalization).

[0064] After obtaining the input-output expressions of the above network layer, the DoRA fine-tuning of the network layer parameters using training data is divided into three stages: Stage 1, only direction training is performed, that is, m is frozen (i.e., m is set to an all-1 vector) and only optimization is performed. Phase 2 involves amplitude training only, essentially freezing the data after convergence in Phase 1. The value is optimized only for the magnitude m. Stage 3: Joint training, using a minimal learning rate (smaller than both Stage 1 and Stage 2 learning rate settings), after Stage 1 and Stage 2 have converged. The values ​​of m and m are used for joint training, thereby completing the DoRA fine-tuning training of the above network layers.

[0065] The training of network layers for each subtask can be achieved by referring to the DoRA fine-tuning training method described above.

[0066] Coal mines are subject to various disturbances, including equipment vibration, rapid personnel movement, and sudden changes in light. By performing phased DoRA fine-tuning training on the network layers of each sub-task, the model can adapt to the complex environment of multiple disturbances underground, layer by layer and component by component, thereby improving the robustness of the final model.

[0067] S1400 is used to manage safety in the regulated area based on safety hazard detection reports.

[0068] This disclosure provides a safety management method for industrial and mining enterprises, including: acquiring video data, audio data, environmental data, and scene types of the monitored area; extracting features from the video data, audio data, and environmental data respectively to obtain video features, audio features, and environmental features; generating a safety hazard detection report for the monitored area using a pre-built safety hazard detection model based on the video features, audio features, environmental features, and hazard detection prompts matched with the scene type; and conducting safety management and control of the monitored area based on the safety hazard detection report. This disclosure achieves cross-validation and complementarity of multimodal information by integrating video, audio, and environmental data, overcoming the shortcomings of traditional single-sensor data being one-sided and prone to missed detections. Simultaneously, by combining a prompt-guided model matched with scene types for targeted analysis, hazard identification can adapt to the specific risks of different work scenarios, thereby significantly improving the accuracy of safety hazard detection and the precision of management in complex industrial and mining scenarios.

[0069] <Device Embodiment> Figure 2 A schematic block diagram of a safety control device for industrial and mining enterprises according to an embodiment of the present disclosure is shown. Figure 2 As shown, the device 100 includes: Data acquisition module 110 is used to acquire video data, audio data, environmental data, and scene types of the monitored area; The feature extraction module 120 is used to extract features from the video data, audio data and environmental data respectively to obtain video features, audio features and environmental features; The hazard detection module 130 is used to generate a safety hazard detection report for the supervised area based on the video features, the audio features, the environmental features, and hazard detection prompts that match the scene type, using a pre-built safety hazard detection model. The safety management module 140 is used to manage and control the safety of the supervised area based on the safety hazard detection report.

[0070] <Equipment Example> Figure 3 A schematic block diagram of a safety control device for industrial and mining enterprises according to an embodiment of the present disclosure is shown. Figure 3 As shown, the safety management and control device 200 for industrial and mining enterprises includes a processor 210 and a memory 220 for storing executable instructions of the processor 210. The processor 210 is configured to implement any of the aforementioned safety management and control methods for industrial and mining enterprises when executing the executable instructions.

[0071] It should be noted here that the number of processors 210 can be one or more. Furthermore, the safety management and control equipment 200 for industrial and mining enterprises in this embodiment may also include an input device 230 and an output device 240. The processors 210, memory 220, input device 230, and output device 240 can be connected via a bus or other means, without specific limitations here.

[0072] The memory 220, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and various modules, such as the program or module corresponding to the safety management and control method for industrial and mining enterprises in this embodiment of the present disclosure. The processor 210 executes various functional applications and data processing of the safety management and control equipment 200 for industrial and mining enterprises by running the software programs or modules stored in the memory 220.

[0073] Input device 230 can be used to receive input digital numbers or signals. These signals may include key signals related to user settings and function control of the device / terminal / server. Output device 240 may include a display device such as a screen.

[0074] <Storage Medium Examples> According to a fourth aspect of this disclosure, a non-volatile computer-readable storage medium is also provided, on which computer program instructions are stored, which, when executed by processor 210, implement any of the aforementioned safety management methods for industrial and mining enterprises.

[0075] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technical improvements to the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A safety management and control method for industrial and mining enterprises, characterized in that, include: Acquire video data, audio data, environmental data, and scene types within the monitored area; Feature extraction is performed on the video data, audio data, and environmental data respectively to obtain video features, audio features, and environmental features; Based on the video features, audio features, environmental features, and hazard detection prompts matching the scene type, a pre-built safety hazard detection model is used to generate a safety hazard detection report for the monitored area. Based on the aforementioned safety hazard detection report, safety management and control measures are implemented in the supervised area.

2. The method according to claim 1, characterized in that, When extracting features from the video data to obtain the video features, the process includes: Video segments are extracted from the video data based on a preset time window; Keyframes are extracted from the video segment using a first frame extraction algorithm that matches the scene type. Extract the image features of each key frame, and generate the video features based on the image features of each key frame.

3. The method according to claim 2, characterized in that, Extracting image features from the keyframes includes: Calculate the acquisition timestamp of the key frame; Target detection is performed on the keyframe to obtain the target detection result of the keyframe; Based on the acquisition timestamp and the target detection results, the image features of the keyframe are generated.

4. The method according to claim 2, characterized in that, When generating the video features based on the image features of each of the keyframes, the process includes: Extract the image features of each of the key frames; Based on the image features of each key frame, determine whether the preset associated frame extraction conditions are met; If the conditions for extracting the associated frames are not met, the image features of each key frame are combined in sequence to generate the video features.

5. The method according to claim 4, characterized in that, If the associated frame extraction condition is met, perform the following operations: Invoke the second frame extraction algorithm that matches the associated frame extraction conditions; The second frame extraction algorithm is used to extract the associated frames of each key frame and extract the image features of each associated frame. The video features are obtained by sequentially combining the image features of each key frame and the image features of each associated frame.

6. The method according to claim 1, characterized in that, The construction of the aforementioned safety hazard detection model includes: Acquire and construct training data based on historical video data, historical audio data, historical environmental data of the regulated area, and corresponding scene types; The training data is used to sequentially train each sub-task of the security hazard detection on the pre-loaded large language model to obtain the security hazard detection model.

7. The method according to claim 6, characterized in that, When training the pre-loaded large language model sequentially to implement each sub-task of the security hazard detection using the training data, a phased DoRA fine-tuning algorithm is implemented.

8. A safety control device for industrial and mining enterprises, characterized in that, include: The data acquisition module is used to acquire video data, audio data, environmental data, and scene types of the monitored area; The feature extraction module is used to extract features from the video data, audio data, and environmental data respectively to obtain video features, audio features, and environmental features. The hazard detection module is used to generate a safety hazard detection report for the supervised area based on the video features, the audio features, the environmental features, and hazard detection prompts that match the scene type, using a pre-built safety hazard detection model. The safety management module is used to manage and control the safety of the supervised area based on the safety hazard detection report.

9. A safety management and control device for industrial and mining enterprises, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to implement the method of any one of claims 1 to 7 when executing the executable instructions.

10. A non-volatile computer-readable storage medium storing computer program instructions thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.