A security early warning and perception method based on video AI and industry knowledge big data model

By using video AI and industry knowledge big data models to provide safety early warning and perception methods, and dynamically adjusting video parameters, the problem of missed detections caused by environmental changes in mine safety monitoring has been solved, and efficient and continuous risk identification and analysis have been achieved.

CN120766192BActive Publication Date: 2025-12-02四川省地质大数据中心
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511295075.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2025-12-02
Estimated Expiration
2045-09-11

Smart Images

  • Figure CN120766192B_ABST
    Figure CN120766192B_ABST
Patent Text Reader

Abstract

This application relates to the field of safety early warning technology, and in particular to a safety early warning perception method based on video AI and an industry knowledge big data model. The method includes: acquiring current video data; analyzing the current video data based on an industry knowledge big data model to determine the current ambient lighting conditions and target object type; determining monitoring requirements for different objects based on the current ambient lighting conditions and target object type; determining video parameter adjustment instructions based on the monitoring requirements; using the video parameter adjustment instructions and acquiring the adjusted real-time video stream; executing a video AI algorithm to analyze the real-time video stream and determine video risks. It optimizes customized strategies for video analysis, reducing the loss of details caused by general rules. It achieves automated adjustment of video parameters, eliminating the need for manual intervention. By optimizing the video stream, it improves analysis accuracy and real-time performance, meeting the intelligent requirements of the dual-control mechanism for hazard identification, and ensuring the timeliness and reliability of safety early warnings.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of security early warning technology, and in particular to a security early warning perception method based on video AI and industry knowledge big data model. Background Technology

[0002] In the field of mine safety production, video surveillance systems are widely used for safety risk control and hidden danger investigation and management, so as to realize real-time monitoring of equipment operation status, environmental anomalies and personnel behavior, thereby preventing accidents from occurring.

[0003] Existing mine safety monitoring methods rely on fixed video parameters or require manual intervention to adjust parameters remotely or on-site. This can lead to key risk points being missed due to parameter mismatch, posing significant hidden dangers in scenarios requiring continuous safety assurance. Summary of the Invention

[0004] This application provides a security early warning and perception method based on video AI and industry knowledge big data model to solve the above problems.

[0005] This application provides a security early warning and perception method based on video AI and industry knowledge big data models, the method comprising:

[0006] Acquire current video data; based on the industry knowledge model, analyze the current video data to determine the current ambient lighting conditions and the type of target object;

[0007] Based on the current ambient lighting conditions and the type of the target object, determine the monitoring requirements for different objects;

[0008] Based on the monitoring requirements, determine the video parameter adjustment instructions;

[0009] The video parameter adjustment instructions are used to obtain the adjusted real-time video stream; a video AI algorithm is executed to analyze the real-time video stream and determine video risks.

[0010] This solution acquires current video data, avoiding analysis interruptions due to data loss and supporting continuous monitoring. Based on a large industry knowledge model, it analyzes current video data to determine ambient lighting conditions and target object types, enabling intelligent identification of the environment and objects, and leveraging industry knowledge to improve analysis accuracy. According to current ambient lighting conditions and target object types, it identifies monitoring needs for different objects, optimizes customized video analysis strategies, ensures monitoring needs match the actual mine scenario, improves the targeting of video acquisition, and reduces the loss of details due to general rules. Based on monitoring needs, it determines video parameter adjustment instructions, enabling automated adjustment of video parameters, eliminating the need for manual intervention, ensuring real-time response to environmental changes, and maintaining monitoring continuity. Using video parameter adjustment instructions and acquiring the adjusted real-time video stream improves video stream quality, provides clear video input, lays the foundation for the risk analysis phase, ensures efficient processing of high-quality data by the video AI algorithm, and improves overall robustness. The video AI algorithm is executed to analyze the real-time video stream, identify video risks, and improve analysis accuracy and real-time performance by optimizing the video stream, meeting the intelligent requirements of the dual-control mechanism for hazard identification, and ensuring the timeliness and reliability of safety warnings.

[0011] Optionally, the step of analyzing the current video data based on the industry knowledge model to determine the current ambient lighting conditions and target object type includes:

[0012] Analyze the current video data to determine the keyframe sequence;

[0013] Based on the pre-set mine optical feature library in the industry knowledge big model, the key frame sequence is analyzed to determine the light intensity level and brightness distribution.

[0014] Based on the brightness distribution, the uniformity of illumination is determined;

[0015] The current ambient lighting conditions are determined based on the light intensity level and the light uniformity.

[0016] Based on the current ambient lighting conditions, the current video data is analyzed to determine the object outline;

[0017] Based on the industry knowledge model, the object outline is matched to determine the target object type.

[0018] This solution analyzes current video data, determines keyframe sequences, improves processing efficiency, and ensures continuous processing of real-time video streams. Based on a pre-built mine optical feature library within the industry knowledge model, it analyzes keyframe sequences to determine light intensity levels and brightness distribution, providing precise quantitative indicators for environmental lighting assessment. Based on brightness distribution, it determines illumination uniformity to avoid video analysis distortion caused by uneven illumination. Based on light intensity levels and uniformity, it determines current environmental lighting conditions and clearly defines lighting environment labels, providing a basis for dynamic adjustment of video parameters. Based on current environmental lighting conditions, it analyzes current video data to determine object contours, ensuring accurate contour extraction and reducing the impact of lighting changes on shape recognition. Based on the industry knowledge model, it matches object contours, determines target object types, distinguishes easily confused targets, and meets the need for focused monitoring of accident-prone areas.

[0019] Optionally, the step of analyzing the current video data based on the current ambient lighting conditions to determine the object outline includes:

[0020] The edge detection gradient threshold is dynamically adjusted based on the current ambient lighting conditions.

[0021] Based on the adjusted edge detection gradient threshold, the current video data is analyzed to determine the object outline.

[0022] This solution dynamically adjusts the edge detection gradient threshold based on current ambient lighting conditions, reducing processing latency and meeting the timeliness requirements of mine safety monitoring. Based on the adjusted edge detection gradient threshold, the current video data is analyzed to determine object contours, reflecting the true edges of objects, supporting efficient matching of target object types, and improving the overall output quality of the video analysis process.

[0023] Optionally, the step of matching the object contour based on the industry knowledge model to determine the target object type includes:

[0024] Based on the pre-built equipment motion feature library in the industry knowledge big model, predict the trajectory of the object's posture change;

[0025] Generate a dynamic projection template based on the object's posture change trajectory;

[0026] The type of the target object is determined based on the object outline and the dynamic projection template.

[0027] This solution predicts object posture change trajectories based on a pre-built equipment motion feature library within an industry knowledge model, avoiding full library traversal and improving real-time performance. Based on the object posture change trajectory, a dynamic projection template is generated, integrating the spatiotemporal characteristics of the object's posture trajectory to provide a dynamic reference standard for actual contour comparison. Based on the object contour and the dynamic projection template, the target object type is determined, enabling differentiated processing for different objects.

[0028] Optionally, determining the monitoring requirements for different objects based on the current ambient lighting conditions and the type of the target object includes:

[0029] Based on the target object type and the pre-set risk location map in the industry knowledge big data model, the key components that need enhanced monitoring are identified;

[0030] Calculate the minimum identifiable contrast of the key component based on the current ambient lighting conditions;

[0031] Based on the minimum identifiable contrast, the monitoring requirements for different objects are determined.

[0032] This solution identifies key components requiring enhanced monitoring based on the target object type and pre-defined risk location maps within an industry knowledge model, avoiding resource waste caused by generic detection. It calculates the minimum identifiable contrast ratio of these key components based on current ambient lighting conditions, providing an actionable quantitative basis for monitoring needs and preventing detail blurring due to undefined thresholds. Based on this minimum identifiable contrast ratio, it determines monitoring requirements for different objects, responding to optimization needs tailored to specific object monitoring scenarios and improving risk identification accuracy.

[0033] Optionally, dynamically adjusting the edge detection gradient threshold based on the current ambient lighting conditions includes:

[0034] The current video data is analyzed to obtain the inter-frame difference matrix;

[0035] Analyze the inter-frame difference matrix to determine the changes in consecutive frames;

[0036] The object's acceleration is determined based on the changes in the continuous frames.

[0037] The edge detection gradient threshold is dynamically adjusted based on the object's moving acceleration and the current ambient lighting conditions.

[0038] This solution analyzes current video data to obtain an inter-frame difference matrix, effectively capturing pixel-level differences caused by sudden changes in lighting conditions or object movement. Analyzing the inter-frame difference matrix determines continuous frame changes and quantifies the global motion intensity of the video stream, replacing subjective judgment based on manual observation. Based on continuous frame changes, the acceleration of object movement is determined, enabling prediction of object motion trends and overcoming the lag of responding only to instantaneous speeds. Based on object movement acceleration and current ambient lighting conditions, a dynamically adjusted edge detection gradient threshold is determined, improving the robustness of the video AI algorithm in dynamic environments and ensuring the accuracy of key component recognition.

[0039] Optionally, the construction of the industry knowledge model includes:

[0040] Obtain historical data on mine safety production;

[0041] Based on the historical data and the preset general model, the parameters of the knowledge enhancement layer are determined;

[0042] Based on the parameters of the knowledge enhancement layer, a large industry knowledge model is trained.

[0043] This solution acquires historical data on mine safety production, ensuring data comprehensiveness and industry relevance while avoiding the limitations of generic rules. Based on historical data and a pre-set general model, it determines the parameters of the knowledge enhancement layer, transforming the general model into a mine-customized framework to address issues related to insufficient real-time performance and robustness. Using the knowledge enhancement layer parameters, it trains an industry-specific knowledge model, ensuring the accuracy of video AI analysis and the real-time nature of early warnings, thus meeting the intelligent requirements for mine safety early warning.

[0044] Optionally, determining the knowledge enhancement layer parameters based on the historical data and the preset general model includes:

[0045] The historical data on mine safety production is structured to extract equipment operation feature sets, environmental status feature sets, and accident correlation feature sets.

[0046] Input the device operation feature set and the environmental state feature set into a preset general large model to generate a general feature vector;

[0047] Based on the accident-related feature set, a trainable weight matrix is ​​obtained;

[0048] Calculate the feature space difference between the accident-related feature set and the general feature vector;

[0049] Based on the difference in the feature space, the trainable weight matrix is ​​dynamically adjusted to obtain the parameters of the knowledge enhancement layer.

[0050] This solution structures historical mine safety production data, extracting equipment operation feature sets, environmental state feature sets, and accident-related feature sets. This reduces the need for manual intervention and improves the model's adaptability to complex mine scenarios. The equipment operation feature sets and environmental state feature sets are input into a pre-defined general model to generate general feature vectors, addressing robustness issues, reducing video quality fluctuations, and ensuring the model can adapt to diverse target objects. Based on the accident-related feature sets, a trainable weight matrix is ​​obtained, optimizing video parameters to meet the monitoring needs of different objects. The feature space difference between the accident-related feature sets and the general feature vectors is calculated, addressing real-time performance limitations and providing a basis for adjustment to predict environmental changes or object behavior, ensuring the model can respond to risks promptly and improving the robustness of the video AI algorithm. Based on the feature space difference, the trainable weight matrix is ​​dynamically adjusted to obtain knowledge enhancement layer parameters, addressing the issues of high manual intervention requirements and lack of adaptability. This enables dynamic optimization of video parameters, thereby improving the real-time performance of early warnings and enhancing the accuracy of video stream acquisition.

[0051] Optionally, training the industry knowledge model based on the parameters of the knowledge enhancement layer includes:

[0052] The general feature vector is input into the hidden layer of a preset general large model;

[0053] The knowledge enhancement layer parameters are superimposed on the output of the hidden layer to generate an enhanced feature representation.

[0054] Based on the accident association feature set, construct supervisory labels;

[0055] Calculate the loss function between the enhanced feature representation and the supervision label;

[0056] Based on the loss function, the trainable weight matrix is ​​back-optimized and iteratively updated until convergence, resulting in a large industry knowledge model.

[0057] This solution inputs general feature vectors into the hidden layer of a pre-defined general model, ensuring that feature representations are more easily integrated with industry knowledge. Knowledge enhancement layer parameters are superimposed on the output of the hidden layer to generate enhanced feature representations, improving their relevance and accuracy. Based on the accident-related feature set, supervisory labels are constructed to ensure the training direction aligns with mine safety early warning needs, avoiding deviations. The loss function between the enhanced feature representations and the supervisory labels is calculated to identify mismatches between the feature representations and mine accident knowledge, providing a basis for back-optimization and driving the model towards more accurate mine industry knowledge alignment. Based on the loss function, the trainable weight matrix is ​​back-optimized, iteratively updated until convergence, resulting in a large industry knowledge model. This enables intelligent dynamic adjustment of video parameters, improving the analytical accuracy and real-time performance of video AI algorithms, and addressing the issues of high demand for manual intervention, insufficient adaptability, and poor real-time performance.

[0058] Optionally, calculating the loss function between the enhanced feature representation and the supervision label includes:

[0059] The enhanced feature representation is input into the learnable mapping layer of the preset general large model, and the prediction matrix is ​​output.

[0060] Calculate the element-wise difference between the prediction matrix and the supervision label matrix;

[0061] Based on the pre-set accident weight allocation table in the industry knowledge big data model, the element-by-element difference is weighted and summed to generate a loss function.

[0062] This solution inputs enhanced feature representations into a learnable mapping layer of a pre-defined general model, outputting a prediction matrix. This avoids comparison errors caused by dimensionality mismatch, achieving a feature-to-prediction conversion and reflecting the feature mapping effect after industry knowledge enhancement. Element-wise dissimilarity between the prediction matrix and the supervision label matrix is ​​calculated to capture specific errors in the prediction, rather than just the overall error, thus providing high-resolution dissimilarity information that reflects the accuracy of the prediction at the micro level and avoids the loss of detail caused by global averaging. Based on the pre-defined accident weight allocation table in the industry knowledge model, the element-wise dissimilarity is weighted and summed to generate a loss function, optimizing the prediction accuracy of key safety events and thereby improving the targeting of video parameter adjustments and the precision of early warnings. Attached Figure Description

[0063] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0064] Figure 1 This is a schematic diagram illustrating an application scenario provided in one embodiment of this application;

[0065] Figure 2 A flowchart illustrating a security early warning and perception method based on video AI and a large industry knowledge model, provided as an embodiment of this application. Detailed Implementation

[0066] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0067] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.

[0068] The embodiments of this application will now be described in further detail with reference to the accompanying drawings.

[0069] Existing mine safety monitoring methods rely on fixed video parameters or require manual intervention to adjust parameters remotely or on-site. This can lead to key risk points being missed due to parameter mismatch, posing significant hidden dangers in scenarios requiring continuous safety assurance.

[0070] When ambient lighting conditions change abruptly (e.g., from a bright area to a dimly lit tunnel) or when the target object moves rapidly, the system cannot automatically adjust video parameters. Operators must pause the inspection process and manually reconfigure the parameters, which not only disrupts the continuity of real-time monitoring but also reduces overall efficiency. Furthermore, video parameter adjustments are based only on general rules, rather than being optimized for the specific monitoring needs of different objects. For example, the outline of high-speed rotating mining machinery is easily distorted under dynamic lighting, but the existing system cannot predict posture changes based on equipment type, leading to inaccurate edge detection. Similarly, for high-risk components requiring focused monitoring (such as conveyor belt joints), the system does not consider minimum identifiable contrast requirements, resulting in lost details.

[0071] Based on this, this application provides a safety early warning and perception method based on video AI and an industry knowledge big data model. This method acquires current video data, avoids analysis interruptions due to data loss, and supports continuous monitoring. Based on the industry knowledge big data model, it analyzes current video data to determine current ambient lighting conditions and target object types, achieving intelligent identification of the environment and objects, and improving analysis accuracy using industry knowledge. According to the current ambient lighting conditions and target object types, it determines monitoring needs for different objects, optimizes customized video analysis strategies, ensures that monitoring needs match the actual mining scenario, improves the targeting of video acquisition, and reduces the loss of details caused by general rules. Based on monitoring needs, it determines video parameter adjustment instructions, achieving automated adjustment of video parameters, eliminating the need for manual intervention, ensuring real-time response to environmental changes, and maintaining monitoring continuity. Using the video parameter adjustment instructions and acquiring the adjusted real-time video stream improves video stream quality, provides clear video input, lays the foundation for the risk analysis stage, ensures that the video AI algorithm efficiently processes high-quality data, and improves overall robustness. By executing video AI algorithms, analyzing real-time video streams, identifying video risks, and optimizing video streams to improve analysis accuracy and real-time performance, the system meets the intelligent requirements of the dual-control mechanism for hazard identification and ensures the timeliness and reliability of safety warnings.

[0072] Figure 1 This is a schematic diagram of an application scenario provided by this application, in which the method provided by this application is applied when performing security early warning perception.

[0073] Specifically, the method provided in this application can be applied to any server, which interacts with the mine video monitoring system to acquire current video data in real time. Based on an industry knowledge model, the current video data is analyzed to determine the current ambient lighting conditions and target object types. Based on the current ambient lighting conditions and target object types, monitoring requirements for different objects are determined. Based on the monitoring requirements, video parameter adjustment instructions are determined. The video parameter adjustment instructions are used to acquire the adjusted real-time video stream, improving video stream quality, providing clear video input, laying the foundation for the risk analysis phase, ensuring that the video AI algorithm efficiently processes high-quality data, and improving overall robustness. The video AI algorithm is executed to analyze the real-time video stream, determine video risks, and improve analysis accuracy and real-time performance by optimizing the video stream, meeting the intelligent requirements of the dual-control mechanism for hazard identification, and ensuring the timeliness and reliability of safety warnings.

[0074] For specific implementation details, please refer to the following examples.

[0075] Figure 2 This is a flowchart illustrating a security early warning and perception method based on video AI and a large industry knowledge model, provided as an embodiment of this application. The method of this embodiment can be applied to servers in the above scenarios. Figure 2As shown, the method includes:

[0076] S201. Obtain current video data; based on the industry knowledge model, analyze the current video data to determine the current ambient lighting conditions and target object type;

[0077] The current video data can be a real-time video stream of a mining scene.

[0078] Industry knowledge big models can be pre-trained large-scale artificial intelligence models used to analyze current video data to output ambient lighting conditions and target object types.

[0079] The current ambient lighting conditions can be the lighting conditions in the video scene.

[0080] The target object type can be the object category identified in the video.

[0081] Specifically, real-time video data is acquired through the mine video monitoring system. By integrating object feature data from the mining industry (such as the structural outline of mining equipment, the movement trajectory of conveying machinery, and typical behavior patterns of personnel), environmental lighting data (such as the distribution of light intensity in underground areas, equipment shadow patterns, and high / low contrast scene samples), and a risk case database (such as equipment failure images and accident videos), the data is input into a pre-trained neural network for multimodal training to build a large-scale industry knowledge model.

[0082] The current video data is input into the industry knowledge big model; the industry knowledge big model extracts the brightness features (such as overall light intensity, shadow distribution and regional contrast) in the video frames (single static image frames in the current video data) of the current video data, and determines the current conditions (for example, bright areas indicate good lighting, and dense shadow areas indicate low uniformity) through feature matching.

[0083] At the same time, the system uses an industry knowledge big data model to identify the outlines of objects (the outer boundary lines of objects in the current video data) and motion patterns (such as the rotational characteristics of mining equipment, the linear motion of conveying machinery, or the walking trajectory of personnel) in the current video data, and determines the type of target object (such as mining equipment, conveying machinery, personnel, etc.).

[0084] S202. Based on the current ambient lighting conditions and the type of target object, determine the monitoring requirements for different objects;

[0085] Different objects can be multiple target object types.

[0086] Monitoring requirements can be optimized for different object types under different ambient lighting conditions.

[0087] Specifically, a pre-set monitoring requirement rule base is established based on risk analysis of the mining industry (such as historical data of equipment failures, accident cases, etc.) and expert experience (such as minimum identifiable contrast requirements, abnormal behavior identification standards, etc.) to determine the monitoring requirements of different objects.

[0088] Then, based on the current ambient lighting conditions and the type of target object, the system queries the preset monitoring requirement rule base to determine the monitoring requirements for different objects. For example, if the target object is mining equipment and the current ambient lighting conditions are dim, the monitoring requirement is to increase exposure and contrast to capture mechanical details; if the target object is a person and the current ambient lighting conditions are high dynamic range shadows, the monitoring requirement is to increase the frame rate to reduce motion blur.

[0089] S203. Determine the video parameter adjustment instructions based on monitoring requirements;

[0090] Video parameter adjustment commands can be commands that control the adjustment of parameters of video acquisition devices.

[0091] Specifically, based on the mapping logic, the monitoring requirements are converted into video parameter adjustment instructions. For example, when the monitoring requirements include high contrast, instructions are generated to increase the contrast value (such as increasing the difference between the brightness and darkness of the image); when the monitoring requirements include low light enhancement, instructions are generated to increase the exposure time (such as extending the light-gathering time of the image sensor).

[0092] S204. Use video parameter adjustment instructions and obtain the adjusted real-time video stream; execute video AI algorithms to analyze the real-time video stream and determine video risks.

[0093] Real-time video streams can be modified video data streams.

[0094] Video AI algorithms can be pre-defined artificial intelligence algorithms used for video risk identification.

[0095] Video risks can be security vulnerabilities identified in real-time video streams.

[0096] Specifically, the video parameter adjustment command is sent to the front-end camera; after receiving the command, the front-end camera adjusts its internal parameters in real time (such as modifying sensor exposure settings or image processor contrast) and re-captures video to obtain the adjusted real-time video stream.

[0097] By analyzing real-time video streams frame by frame using video AI algorithms, features (such as motion trajectories or texture changes) can be extracted to determine video risks. For example, for mining equipment, this can detect whether key components are malfunctioning; for the environment, it can identify potential hazards caused by sudden changes in lighting.

[0098] This solution acquires current video data, avoiding analysis interruptions due to data loss and supporting continuous monitoring. Based on a large industry knowledge model, it analyzes current video data to determine ambient lighting conditions and target object types, enabling intelligent identification of the environment and objects, and leveraging industry knowledge to improve analysis accuracy. According to current ambient lighting conditions and target object types, it identifies monitoring needs for different objects, optimizes customized video analysis strategies, ensures monitoring needs match the actual mine scenario, improves the targeting of video acquisition, and reduces the loss of details due to general rules. Based on monitoring needs, it determines video parameter adjustment instructions, enabling automated adjustment of video parameters, eliminating the need for manual intervention, ensuring real-time response to environmental changes, and maintaining monitoring continuity. Using video parameter adjustment instructions and acquiring the adjusted real-time video stream improves video stream quality, provides clear video input, lays the foundation for the risk analysis phase, ensures efficient processing of high-quality data by the video AI algorithm, and improves overall robustness. The video AI algorithm is executed to analyze the real-time video stream, identify video risks, and improve analysis accuracy and real-time performance by optimizing the video stream, meeting the intelligent requirements of the dual-control mechanism for hazard identification, and ensuring the timeliness and reliability of safety warnings.

[0099] In some embodiments, the current video data is parsed to determine the keyframe sequence; the keyframe sequence is analyzed based on the pre-built mine optical feature library in the industry knowledge big model to determine the light intensity level and brightness distribution; the light uniformity is determined based on the brightness distribution; the current ambient lighting conditions are determined based on the light intensity level and light uniformity; the current video data is analyzed based on the current ambient lighting conditions to determine the object outline; and the object outline is matched based on the industry knowledge big model to determine the target object type.

[0100] A keyframe sequence can be a sequence of frames extracted from the current video data that reflects the lighting conditions.

[0101] The mine optical feature library can be a pre-built library in the industry knowledge big model used to analyze key frame sequences to quantify illumination characteristics.

[0102] Light intensity level can be a classification level of ambient light intensity.

[0103] The brightness distribution can be the distribution characteristics of brightness in the video frame space.

[0104] Illumination uniformity can be an indicator of the uniformity of illumination in space.

[0105] The object outline can be the edge outline of the object extracted from the current video data.

[0106] Specifically, a time sampling method is used to extract representative frames from the current video data at fixed time intervals. The extracted representative frames are then sorted by timestamps to generate a keyframe sequence. The industry knowledge model's pre-built mine optical feature library (which stores typical lighting reference data for mine scenes) is invoked to analyze the keyframe sequence, calculate its average brightness value (e.g., the average pixel value of a grayscale image), and map the lighting intensity level (including low, medium, and high intensity) according to the preset thresholds in the mine optical feature library (used to map the calculated average brightness value to discrete levels). The keyframe sequence is then divided into a grid, and the average brightness value of each grid (the arithmetic mean of pixel brightness values ​​within the grid area) is calculated. The distribution characteristics of the grid's average brightness value (e.g., maximum, minimum, and variance) are statistically analyzed to determine the brightness distribution.

[0107] Calculate the overall variance of the brightness distribution; map the light uniformity level (including uniform, relatively uniform, and non-uniform) based on the variance. Based on the light intensity level and light uniformity, determine the current ambient lighting conditions using conditional mapping rules. For example, if the light intensity level is low and the light uniformity is non-uniform, the current ambient lighting conditions are dim and uneven; if the light intensity level is high and the light uniformity is uniform, the current ambient lighting conditions are bright and uniform.

[0108] Adjusting image processing parameters based on current ambient lighting conditions (adjustable settings in contour detection algorithms used to control the object contour extraction process), different algorithms are applied to the current video data for contour detection to determine the object contour. For example, under uniform strong light conditions, the Canny edge detection algorithm is applied to extract the object contour; under dim and uneven conditions, an adaptive threshold segmentation algorithm is used to enhance the low-light area and then extract the object contour.

[0109] The system calls upon the object shape template library in the industry knowledge big data model (pre-defined standard outlines of target objects in the mine, such as the geometry of mining equipment); it matches the extracted object outlines with the object shape template library and calculates the outline similarity; when the outline similarity exceeds the similarity threshold based on experimental statistics (used to determine whether the match is successful), the match is successful; based on the matching result, the target object type is determined, for example, matching the mining equipment template will identify it as a crusher, and matching the personnel template will identify it as a worker.

[0110] This solution analyzes current video data, determines keyframe sequences, improves processing efficiency, and ensures continuous processing of real-time video streams. Based on a pre-built mine optical feature library within the industry knowledge model, it analyzes keyframe sequences to determine light intensity levels and brightness distribution, providing precise quantitative indicators for environmental lighting assessment. Based on brightness distribution, it determines illumination uniformity to avoid video analysis distortion caused by uneven illumination. Based on light intensity levels and uniformity, it determines current environmental lighting conditions and clearly defines lighting environment labels, providing a basis for dynamic adjustment of video parameters. Based on current environmental lighting conditions, it analyzes current video data to determine object contours, ensuring accurate contour extraction and reducing the impact of lighting changes on shape recognition. Based on the industry knowledge model, it matches object contours, determines target object types, distinguishes easily confused targets, and meets the need for focused monitoring of accident-prone areas.

[0111] In some embodiments, the edge detection gradient threshold is dynamically adjusted according to the current ambient lighting conditions; based on the adjusted edge detection gradient threshold, the current video data is analyzed to determine the object outline.

[0112] Edge detection gradient thresholds can be threshold values ​​for the intensity of grayscale changes used to determine whether a pixel in an image belongs to the edge of an object. They include high gradient thresholds and low gradient thresholds.

[0113] Specifically, the edge detection gradient threshold is set based on the experience data of mining scenarios from the industry knowledge big data model. The edge detection gradient threshold is dynamically adjusted according to the current ambient lighting conditions. For example, if the current ambient lighting conditions are bright and uniform (such as high intensity and uniformity of light), the high gradient threshold and low gradient threshold of the Canny edge detection algorithm are selected. If the current ambient lighting conditions are dim and uneven (such as low intensity and unevenness of light), the Canny algorithm is abandoned and the adaptive threshold segmentation algorithm is switched to, and the parameters are set.

[0114] Under bright and uniform conditions, Gaussian filtering is applied to the current video data for noise reduction based on the adjusted edge detection gradient threshold. The pixel gradient magnitude (the intensity of the gray value change at each pixel in the current video data) and direction (the direction of the gray value change at the pixel) are calculated. Then, non-maximum suppression is applied to refine the edges. Dual threshold detection is performed using a high gradient threshold and a low gradient threshold, and edge pixels are connected to generate candidate contours.

[0115] Under bright and uniform conditions, the current video data is divided into multiple neighborhood blocks (used to divide the local regions of the current video data); a local threshold is independently calculated for each neighborhood block (adapting to local lighting conditions), and connected regions are extracted; then, after filtering out noise, candidate contours are generated.

[0116] Perform morphological operations on candidate contours (basic operations for modifying shape or structure, including dilation (expanding the area), erosion (shrinking the area), opening (erosion followed by dilation) or closing (dilation followed by erosion, such as closing to fill small holes or opening to remove isolated points), and filter contours with too small an area; use the filtered candidate contours as the final object contour.

[0117] This solution dynamically adjusts the edge detection gradient threshold based on current ambient lighting conditions, reducing processing latency and meeting the timeliness requirements of mine safety monitoring. Based on the adjusted edge detection gradient threshold, the current video data is analyzed to determine object contours, reflecting the true edges of objects, supporting efficient matching of target object types, and improving the overall output quality of the video analysis process.

[0118] In some embodiments, the trajectory of an object's posture change is predicted based on a pre-built equipment motion feature library in an industry knowledge big data model; a dynamic projection template is generated based on the trajectory of the object's posture change; and the type of the target object is determined based on the object's outline and the dynamic projection template.

[0119] The equipment motion feature library can be a set of motion patterns of typical mining equipment pre-set in the industry knowledge big model.

[0120] The trajectory of an object's attitude change can be a parameterized path of the object's motion in a short time window in the future, predicted by a device motion feature library.

[0121] A dynamic projection template can be a set of temporal contour sequences generated based on the trajectory of an object's posture change.

[0122] Specifically, it accesses the pre-set equipment motion feature library in the industry knowledge big model (which stores the motion patterns of typical mining equipment, such as the rotation cycle of mining machinery, the linear displacement of conveyor belts, and the random movement patterns of personnel); based on the geometric attributes of the object's outline (such as the aspect ratio of the outline's circumscribed rectangle and the position of the centroid), it matches the motion patterns of similar equipment in the pre-set equipment motion feature library, thereby predicting the trajectory of the object's posture change, such as the horizontal displacement vector of the conveyor belt joint and the rotation angle sequence of the mining drill arm.

[0123] Based on the posture change trajectory, calculate the geometric deformation path of the object (such as the contour coordinate set of the conveyor belt joint after horizontal translation according to the trajectory; the contour frame vertices of the mining machinery according to the rotation angle); convert the geometric deformation path into a temporal projection sequence, and generate a set of dynamic projection templates that evolve according to the trajectory. For example, the translation template of the conveyor belt joint is a sequence of rectangular contour positions for 10 consecutive frames; the rotation template of the mining drill arm is a set of polygon contours at 15° intervals.

[0124] The object outline is geometrically matched with the dynamic projection template; if the matching degree exceeds the threshold preset based on industry experience knowledge, it is determined to be the target object type.

[0125] This solution predicts object posture change trajectories based on a pre-built equipment motion feature library within an industry knowledge model, avoiding full library traversal and improving real-time performance. Based on the object posture change trajectory, a dynamic projection template is generated, integrating the spatiotemporal characteristics of the object's posture trajectory to provide a dynamic reference standard for actual contour comparison. Based on the object contour and the dynamic projection template, the target object type is determined, enabling differentiated processing for different objects.

[0126] In some embodiments, based on the target object type and the pre-set risk location map in the industry knowledge big data model, the key components that need enhanced monitoring are identified; based on the current ambient lighting conditions, the minimum identifiable contrast of the key components is calculated; and based on the minimum identifiable contrast, the monitoring requirements for different objects are determined.

[0127] Risk location map can be a structured database pre-built into an industry knowledge big model, storing risk location information for different equipment types.

[0128] Key components can be specific sub-components that require enhanced monitoring, extracted from the risk location map.

[0129] Minimum identifiable contrast can be the lowest contrast threshold that ensures critical components can be clearly identified in a video stream under ambient lighting conditions.

[0130] Specifically, the target object type is matched with the pre-set risk location map in the industry knowledge big model (which stores risk location information for different equipment types, such as the risk locations of a mining drill arm including the rotary joint and the drill bit, and the risk locations of a conveyor belt joint including the connection point) to determine the set of key components that need enhanced monitoring (such as the rotary joint). For example, if the target object type is a mining drill arm, the risk location map returns its list of key components: the rotary joint and the drill bit.

[0131] Based on the current ambient lighting conditions, the minimum identifiable contrast of key components is calculated using the pre-set calculation rules in the industry knowledge model (considering the impact of lighting conditions on component visibility). For example, in low-brightness environments (such as dimly lit areas underground), the minimum identifiable contrast value of key components (such as conveyor belt joints) is higher; while in high-brightness environments (such as equipment lighting areas), the minimum identifiable contrast value is lower.

[0132] The minimum identifiable contrast is mapped to the adjustment requirements of video acquisition parameters (dynamically adjustable physical imaging control parameters in front-end video surveillance equipment (such as underground cameras) used to determine the monitoring requirements of key components under different objects), thereby determining the monitoring requirements of key components under different objects. For example, if the minimum identifiable contrast of the key component is high, the monitoring requirements include increasing the video contrast parameter or increasing exposure compensation; if the minimum identifiable contrast is low, the monitoring requirements may be to maintain the default contrast setting.

[0133] This solution identifies key components requiring enhanced monitoring based on the target object type and pre-defined risk location maps within an industry knowledge model, avoiding resource waste caused by generic detection. It calculates the minimum identifiable contrast ratio of these key components based on current ambient lighting conditions, providing an actionable quantitative basis for monitoring needs and preventing detail blurring due to undefined thresholds. Based on this minimum identifiable contrast ratio, it determines monitoring requirements for different objects, responding to optimization needs tailored to specific object monitoring scenarios and improving risk identification accuracy.

[0134] In some embodiments, the current video data is parsed to obtain an inter-frame difference matrix; the inter-frame difference matrix is ​​analyzed to determine the changes in consecutive frames; the object's motion acceleration is determined based on the changes in consecutive frames; and the edge detection gradient threshold is dynamically adjusted based on the object's motion acceleration and the current ambient lighting conditions.

[0135] An inter-frame difference matrix can be a two-dimensional matrix that stores the brightness difference values ​​of corresponding pixel positions between consecutive video frames.

[0136] Frame-to-frame variation can be a scalar value representing the overall level of dynamic change between consecutive frames.

[0137] The acceleration of an object can be a scalar value representing the rate of change of the velocity of the target object.

[0138] Specifically, a continuous sequence of video frames (including the current frame and its previous frame) is obtained from the current video data; each frame is converted into a grayscale image to simplify processing; then the inter-frame differences are calculated to generate an inter-frame difference matrix.

[0139] A global statistical analysis is performed on the inter-frame difference matrix to calculate the average absolute value of all elements in the inter-frame difference matrix (i.e., the average inter-frame difference value), which is used as the continuous frame variation. High continuous frame variation indicates significant motion or object displacement (such as rapid device movement), while low continuous frame variation indicates relative stillness or small changes (such as gradual changes in ambient light).

[0140] Based on the changes in consecutive frames, the object's moving speed is deduced (e.g., high consecutive frame changes correspond to high object moving speed); by analyzing the object's moving speed at multiple consecutive time points, the object's moving acceleration is determined. For example, when the object's moving acceleration is positive, it indicates that it is accelerating, and when the object's moving acceleration is negative, it indicates that it is decelerating.

[0141] Based on the object's acceleration and the current ambient lighting conditions, the edge detection gradient threshold is dynamically adjusted using predefined rules. For example, if the object's acceleration is high and the current ambient lighting conditions are poor (low brightness), the edge detection gradient threshold is significantly reduced to capture blurry edges; if the object's acceleration is low and the current ambient lighting conditions are good (high brightness), the edge detection gradient threshold is moderately increased to reduce false noise detections.

[0142] This solution analyzes current video data to obtain an inter-frame difference matrix, effectively capturing pixel-level differences caused by sudden changes in lighting conditions or object movement. Analyzing the inter-frame difference matrix determines continuous frame changes and quantifies the global motion intensity of the video stream, replacing subjective judgment based on manual observation. Based on continuous frame changes, the acceleration of object movement is determined, enabling prediction of object motion trends and overcoming the lag of responding only to instantaneous speeds. Based on object movement acceleration and current ambient lighting conditions, a dynamically adjusted edge detection gradient threshold is determined, improving the robustness of the video AI algorithm in dynamic environments and ensuring the accuracy of key component recognition.

[0143] In some embodiments, historical data on mine safety production is acquired; based on the historical data and a preset general model, the parameters of the knowledge enhancement layer are determined; and based on the parameters of the knowledge enhancement layer, an industry knowledge model is trained.

[0144] Historical data on mine safety production can be a collection of historical records related to safety production obtained from the mine safety monitoring system, including equipment operation status records, environmental anomaly records, personnel behavior records, accident history data, etc.

[0145] The preset general-purpose large model can be a pre-trained general artificial intelligence basic model for extracting mine feature vectors from historical data through the embedding layer. It is stored in advance on the server and called when needed.

[0146] The knowledge enhancement layer parameters can be model layer parameters obtained by combining and optimizing historical data of mine safety production with a preset general model, which are used to enhance the model's ability to integrate knowledge of the mining industry.

[0147] Specifically, historical data on mine safety production is retrieved from the mine safety monitoring system's mine safety production database (used to store historical data on mine safety production).

[0148] Historical data is cleaned and standardized (e.g., noise removal and missing value filling); features are extracted from historical data using the embedding layer of a pre-defined general model to generate mine feature vectors (e.g., equipment movement patterns or environmental risk patterns); based on the mine feature vectors, optimization algorithms are applied to determine the parameters of the knowledge enhancement layer.

[0149] Historical data on mine safety production is divided into training and validation sets. Through iterative training (such as forward and backward propagation), the parameters of the knowledge enhancement layer are applied to adjust the internal layers of the model (such as fully connected layers or convolutional layers), the training loss (such as cross-entropy or mean squared error) is monitored, and the performance (such as accuracy or recall) is evaluated on the validation set. When the model converges (loss stabilizes), the trained industry knowledge model is output.

[0150] This solution acquires historical data on mine safety production, ensuring data comprehensiveness and industry relevance while avoiding the limitations of generic rules. Based on historical data and a pre-set general model, it determines the parameters of the knowledge enhancement layer, transforming the general model into a mine-customized framework to address issues related to insufficient real-time performance and robustness. Using the knowledge enhancement layer parameters, it trains an industry-specific knowledge model, ensuring the accuracy of video AI analysis and the real-time nature of early warnings, thus meeting the intelligent requirements for mine safety early warning.

[0151] In some embodiments, historical data on mine safety production is structured to extract equipment operation feature sets, environmental state feature sets, and accident-related feature sets; the equipment operation feature sets and environmental state feature sets are input into a preset general large model to generate general feature vectors; a trainable weight matrix is ​​obtained based on the accident-related feature sets; the feature space difference between the accident-related feature sets and the general feature vectors is calculated; and the trainable weight matrix is ​​dynamically adjusted based on the feature space difference to obtain the knowledge enhancement layer parameters.

[0152] Equipment operation feature set can be a set of features extracted from historical data on mine safety production to characterize the dynamic behavior of mine equipment.

[0153] The environmental state feature set can be a set of features extracted in a structured manner from historical data on mine safety production to characterize the dynamic changes in the mine environment.

[0154] Accident-related feature sets can be structured features extracted from historical data on mine safety production for the purpose of integrating mine accident knowledge.

[0155] A general feature vector can be a vector that represents an abstract general pattern of the operating characteristics of the fusion device and the environmental state characteristics.

[0156] The trainable weight matrix can be a matrix initialized based on the accident-related feature set to carry different knowledge of the mining industry.

[0157] The feature space dissimilarity can be a scalar value representing the degree of mismatch between mine accident knowledge and general features.

[0158] Specifically, the historical data on mine safety production is processed in a structured manner, converted into a structured format, and analyzed to extract numerical features (such as vibration amplitude, temperature value, and equipment start-up and shutdown time) from the equipment operation status records (such as sensor data and operation logs of mining equipment and conveying machinery) to form an equipment operation feature set; environmental anomaly records (such as sudden changes in lighting conditions and shadow distribution data) are analyzed to extract numerical features (such as changes in light intensity and shadow area coordinates) to form an environmental status feature set; and accident history data and personnel behavior records (such as accident reports and video clips of violations) are analyzed to extract classification features (such as accident type, location of occurrence, and category of violation) to form an accident association feature set.

[0159] By using an embedding layer of a pre-defined general model, the device operation feature set and the environmental state feature set are mapped to a high-dimensional feature space, generating a general feature vector. Based on the accident-related feature set, a trainable weight matrix is ​​initialized.

[0160] Align the accident-related feature set with the general feature vectors to the same feature space (a high-dimensional vector space used to represent feature vectors and determine the feature space dissimilarity), and calculate the feature space dissimilarity. Using the feature space dissimilarity as the optimization objective, apply an optimization algorithm to calculate the gradient of the trainable weight matrix (e.g., by taking the derivative of the weight matrix through backpropagation); then dynamically adjust the trainable weight matrix; iterate the adjustment process until the dissimilarity converges, and obtain the parameters of the knowledge enhancement layer.

[0161] This solution structures historical mine safety production data, extracting equipment operation feature sets, environmental state feature sets, and accident-related feature sets. This reduces the need for manual intervention and improves the model's adaptability to complex mine scenarios. The equipment operation feature sets and environmental state feature sets are input into a pre-defined general model to generate general feature vectors, addressing robustness issues, reducing video quality fluctuations, and ensuring the model can adapt to diverse target objects. Based on the accident-related feature sets, a trainable weight matrix is ​​obtained, optimizing video parameters to meet the monitoring needs of different objects. The feature space difference between the accident-related feature sets and the general feature vectors is calculated, addressing real-time performance limitations and providing a basis for adjustment to predict environmental changes or object behavior, ensuring the model can respond to risks promptly and improving the robustness of the video AI algorithm. Based on the feature space difference, the trainable weight matrix is ​​dynamically adjusted to obtain knowledge enhancement layer parameters, addressing the issues of high manual intervention requirements and lack of adaptability. This enables dynamic optimization of video parameters, thereby improving the real-time performance of early warnings and enhancing the accuracy of video stream acquisition.

[0162] In some embodiments, a general feature vector is input into the hidden layer of a preset general large model; knowledge enhancement layer parameters are superimposed on the output of the hidden layer to generate an enhanced feature representation; a supervision label is constructed based on the accident association feature set; the loss function between the enhanced feature representation and the supervision label is calculated; the trainable weight matrix is ​​back-optimized based on the loss function, and iteratively updated until convergence is obtained to obtain the industry knowledge large model.

[0163] Hidden layers can be internal processing layers of a pre-defined general-purpose large model.

[0164] The output port of the hidden layer can be the output location or the output result after the hidden layer processing is completed.

[0165] Enhanced feature representations can be output vectors generated by superimposing knowledge enhancement layer parameters at the output of the hidden layer.

[0166] The supervisory label can be a target output constructed from the accident-related feature set for loss calculation with the enhanced feature representation.

[0167] The loss function can be a function used to quantify the difference between the augmented feature representation and the supervisory label.

[0168] Specifically, the general feature vector is used as input data and passed to the hidden layer of a pre-defined general model. At the output of the hidden layer, the parameters of the knowledge enhancement layer are multiplied with the general feature vector to generate an enhanced feature representation (an output vector that integrates the general features of the pre-defined general model (such as equipment operation or environmental status features) and mining industry knowledge). The categorical features (such as accident type or violation category) in the accident-related feature set are encoded into numerical vectors to form supervision labels.

[0169] The enhanced feature representation is compared with the supervised labels, and the loss function is calculated. For example, a large difference (high loss function) indicates that the enhanced features have failed to effectively capture accident-related features; a small difference (low loss function) indicates that the enhanced features are highly aligned with accident knowledge. An optimization algorithm is used to calculate the gradient of the loss function with respect to the trainable weight matrix (representing the instantaneous rate of change of the loss function with respect to each parameter of the trainable weight matrix). Based on the gradient value, the parameters of the trainable weight matrix are dynamically adjusted (e.g., by updating the weight values ​​using gradient descent). This process is then repeated iteratively, involving inputting a general feature vector, overlaying knowledge enhancement layer parameters to generate enhanced feature representations, constructing supervised labels, calculating the loss function, and inversely optimizing the trainable weight matrix.

[0170] During the iteration process, the changes in the loss function are monitored; when the loss function stabilizes or reaches the maximum number of iterations, the iteration is stopped and convergence is determined; finally, the optimized trainable weight matrix is ​​combined with the preset general large model to form an industry knowledge large model.

[0171] This solution inputs general feature vectors into the hidden layer of a pre-defined general model, ensuring that feature representations are more easily integrated with industry knowledge. Knowledge enhancement layer parameters are superimposed on the output of the hidden layer to generate enhanced feature representations, improving their relevance and accuracy. Based on the accident-related feature set, supervisory labels are constructed to ensure the training direction aligns with mine safety early warning needs, avoiding deviations. The loss function between the enhanced feature representations and the supervisory labels is calculated to identify mismatches between the feature representations and mine accident knowledge, providing a basis for back-optimization and driving the model towards more accurate mine industry knowledge alignment. Based on the loss function, the trainable weight matrix is ​​back-optimized, iteratively updated until convergence, resulting in a large industry knowledge model. This enables intelligent dynamic adjustment of video parameters, improving the analytical accuracy and real-time performance of video AI algorithms, and addressing the issues of high demand for manual intervention, insufficient adaptability, and poor real-time performance.

[0172] In some embodiments, the enhanced feature representation is input into the learnable mapping layer of a preset general large model, and the prediction matrix is ​​output; the element-wise difference between the prediction matrix and the supervision label matrix is ​​calculated; and the element-wise difference is weighted and summed according to the pre-set accident weight allocation table in the industry knowledge large model to generate a loss function.

[0173] The learnable mapping layer can be a component of a pre-defined general large model used to perform feature mapping operations.

[0174] The prediction matrix can be a structured data representation of the prediction results of mine safety events through a representation model output by a learnable mapping layer.

[0175] The supervision label matrix can be a numerical matrix generated from the encoding of the accident association feature set.

[0176] Element-wise dissimilarity can be an indicator generated by calculating the difference between elements at the same position in the prediction matrix and the supervision label matrix.

[0177] The accident weight allocation table can be a pre-built data structure in the industry knowledge big data model that stores the weight values ​​of different accident types or risk levels.

[0178] Specifically, the generated enhanced feature representation is used as input data and passed to the learnable mapping layer of the pre-defined general large model; through the linear transformation of the learnable mapping layer, the vector form of the enhanced feature representation is converted into the output matrix, i.e., the prediction matrix.

[0179] A supervision label matrix is ​​generated by encoding the accident association feature set; the prediction matrix and the supervision label matrix are compared element by element, and the difference value, i.e., the element-by-element difference degree, is calculated for each element at the same position in the prediction matrix and the supervision label matrix.

[0180] Based on the pre-set accident weight allocation table in the industry knowledge big data model (which stores weight values ​​for different accident types or risk levels, such as equipment failure having a higher weight than environmental anomalies), the corresponding weight value of each element's difference degree in the accident weight allocation table is found (e.g., the difference degree of high-risk components is given a higher weight); then, the difference degree of each element and its corresponding weight value are weighted and summed to generate a scalar value as the loss function.

[0181] This solution inputs enhanced feature representations into a learnable mapping layer of a pre-defined general model, outputting a prediction matrix. This avoids comparison errors caused by dimensionality mismatch, achieving a feature-to-prediction conversion and reflecting the feature mapping effect after industry knowledge enhancement. Element-wise dissimilarity between the prediction matrix and the supervision label matrix is ​​calculated to capture specific errors in the prediction, rather than just the overall error, thus providing high-resolution dissimilarity information that reflects the accuracy of the prediction at the micro level and avoids the loss of detail caused by global averaging. Based on the pre-defined accident weight allocation table in the industry knowledge model, the element-wise dissimilarity is weighted and summed to generate a loss function, optimizing the prediction accuracy of key safety events and thereby improving the targeting of video parameter adjustments and the precision of early warnings.

Claims

1. A security early warning and perception method based on video AI and industry knowledge big data models, characterized in that, include: Get the current video data; Based on the industry knowledge model, the current video data is analyzed to determine the current ambient lighting conditions and the type of the target object; Based on the current ambient lighting conditions and the type of the target object, determine the monitoring requirements for different objects; Based on the monitoring requirements, determine the video parameter adjustment instructions; Use the video parameter adjustment instructions to obtain the adjusted real-time video stream; execute the video AI algorithm to analyze the real-time video stream and determine video risks; Analyze the current video data to determine the keyframe sequence; Based on the pre-set mine optical feature library in the industry knowledge big model, the key frame sequence is analyzed to determine the light intensity level and brightness distribution. Based on the brightness distribution, the uniformity of illumination is determined; The current ambient lighting conditions are determined based on the light intensity level and the light uniformity. Based on the current ambient lighting conditions, the current video data is analyzed to determine the object outline; Based on the industry knowledge model, the object outline is matched to determine the target object type.

2. The method according to claim 1, characterized in that, The step of analyzing the current video data based on the current ambient lighting conditions to determine the object outline includes: The edge detection gradient threshold is dynamically adjusted based on the current ambient lighting conditions. Based on the adjusted edge detection gradient threshold, the current video data is analyzed to determine the object outline.

3. The method according to claim 1, characterized in that, The process of matching the object contour based on the industry knowledge model to determine the target object type includes: Based on the pre-built equipment motion feature library in the industry knowledge big model, predict the trajectory of the object's posture change; Generate a dynamic projection template based on the object's posture change trajectory; The type of the target object is determined based on the object outline and the dynamic projection template.

4. The method according to claim 1, characterized in that, The step of determining monitoring requirements for different objects based on the current ambient lighting conditions and the type of the target object includes: Based on the target object type and the pre-set risk location map in the industry knowledge big data model, the key components that need enhanced monitoring are identified; Calculate the minimum identifiable contrast of the key component based on the current ambient lighting conditions; Based on the minimum identifiable contrast, the monitoring requirements for different objects are determined.

5. The method according to claim 2, characterized in that, The step of dynamically adjusting the edge detection gradient threshold based on the current ambient lighting conditions includes: The current video data is analyzed to obtain the inter-frame difference matrix; Analyze the inter-frame difference matrix to determine the changes in consecutive frames; The object's acceleration is determined based on the changes in the continuous frames. The edge detection gradient threshold is dynamically adjusted based on the object's moving acceleration and the current ambient lighting conditions.

6. The method according to claim 1, characterized in that, The construction of the aforementioned industry knowledge model includes: Obtain historical data on mine safety production; Based on the historical data and the preset general model, the parameters of the knowledge enhancement layer are determined; Based on the parameters of the knowledge enhancement layer, a large industry knowledge model is trained.

7. The method according to claim 6, characterized in that, The step of determining the knowledge enhancement layer parameters based on the historical data and the preset general model includes: The historical data on mine safety production is structured to extract equipment operation feature sets, environmental status feature sets, and accident correlation feature sets. Input the device operation feature set and the environmental state feature set into a preset general large model to generate a general feature vector; Based on the accident-related feature set, a trainable weight matrix is ​​obtained; Calculate the feature space difference between the accident-related feature set and the general feature vector; Based on the difference in the feature space, the trainable weight matrix is ​​dynamically adjusted to obtain the parameters of the knowledge enhancement layer.

8. The method according to claim 7, characterized in that, The process of training a large industry knowledge model based on the parameters of the knowledge enhancement layer includes: The general feature vector is input into the hidden layer of a preset general large model; The knowledge enhancement layer parameters are superimposed on the output of the hidden layer to generate an enhanced feature representation. Based on the accident association feature set, construct supervisory labels; Calculate the loss function between the enhanced feature representation and the supervision label; Based on the loss function, the trainable weight matrix is ​​back-optimized and iteratively updated until convergence, resulting in a large industry knowledge model.

9. The method according to claim 8, characterized in that, The calculation of the loss function between the enhanced feature representation and the supervisory label includes: The enhanced feature representation is input into the learnable mapping layer of the preset general large model, and the prediction matrix is ​​output. Calculate the element-wise difference between the prediction matrix and the supervision label matrix; Based on the pre-set accident weight allocation table in the industry knowledge big data model, the element-by-element difference is weighted and summed to generate a loss function.

Citation Information

Patent Citations

  • Intelligent human shape trajectory prediction and alarm system and method based on multi-modal video analysis

    CN120047897A

  • Coal mine underground video monitoring method

    CN120416443A