Method and system for intelligently analyzing and identifying potential safety hazards based on picture AI
Through multi-scale feature extraction and fusion technology and unsupervised comparative learning, the limitations of single-scale analysis in existing technologies are overcome, efficient and reliable safety hazard identification is achieved, adapting to different shooting conditions and scene requirements, and reducing training costs and time.
Patent Information
- Application Number
- CN202510726414.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies for identifying safety hazards use single-scale feature analysis, which cannot take into account both subtle and macro safety hazards. Reliance on manually labeled data leads to high training costs, long training cycles, weak model generalization capabilities, and poor adaptability, making it difficult to quickly adapt to different shooting conditions and scene requirements.
Multi-scale feature extraction and fusion technology is adopted, combined with an unsupervised comparative learning training model, to identify subtle and macro safety hazards through multi-scale feature extraction, and to build a positive and negative sample optimization model using unlabeled data, and to introduce a confidence verification mechanism for secondary analysis.
It significantly improves the accuracy and comprehensiveness of safety hazard identification in complex scenarios, reduces training costs and time, enhances the model's recognition capabilities in unlabeled scenarios, and improves recognition reliability and cross-scenario adaptability.
Smart Images

Figure CN120635644A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition technology, and in particular to a method and system for identifying safety hazards based on AI intelligent analysis of images. Background Art
[0002] Image-based intelligent analysis technology has been widely used in the field of safety hazard identification. Currently, mainstream methods mostly rely on single-scale feature extraction and analysis, using manually annotated image data to train supervised learning models to identify safety hazards. For example, in industrial scenarios, some systems only extract mid-level image features for analysis; in construction scenarios, models rely on manually annotated tens of thousands of images to train models to identify hazards such as helmet wearing and scaffolding construction. Furthermore, existing systems lack effective mechanisms for processing low-quality images and are difficult to quickly adapt to the needs of different industry scenarios, limiting the widespread application of this technology.
[0003] However, existing technologies have many shortcomings. First, single-scale feature analysis cannot take into account both subtle and macro safety hazards, resulting in low recognition accuracy in complex scenarios and prone to missed detections. Second, it is highly dependent on manually labeled data, with high training costs and long cycles, and the model has weak generalization capabilities, making it difficult to quickly adapt to new scenarios and new types of hazards. Third, it has poor adaptability to images taken under different shooting conditions, and the recognition effect is significantly reduced in environments such as insufficient lighting and complex angles. Therefore, a method and system for identifying safety hazards based on image AI intelligent analysis are proposed. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, the embodiments of the present invention provide a method and system for identifying safety hazards based on image AI intelligent analysis, which solves the problems of single scale, labeling dependence, poor adaptability and low efficiency in traditional technologies.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A method for identifying safety hazards based on AI intelligent analysis of images, comprising the following steps:
[0007] S1: Obtain the safety hazard image to be analyzed;
[0008] S2: Perform multi-scale feature extraction on the image to generate feature data at different levels;
[0009] S3: Identify subtle security risks in images using low-level feature data and identify macro security risks in images using high-level feature data.
[0010] S4: Use feature fusion algorithm to fuse feature data at different levels to obtain comprehensive feature data;
[0011] S5: Based on the comprehensive feature data, safety hazard identification is performed through an image processing model.
[0012] Preferably, it also includes: using unlabeled safety hazard image data, defining different shooting condition images of the same hazard scene as positive samples, and defining other scene images as negative samples; based on the positive samples and negative samples, performing unsupervised comparative learning training on the image processing model.
[0013] Preferably, the different shooting conditions include at least one of shooting angle, light parameters and shooting distance.
[0014] Preferably, the minor safety hazards include damaged appearance of fire-fighting facilities and blurred escape signs; the macro safety hazards include blocked evacuation passages and abnormal structures of large equipment.
[0015] Preferably, after the safety hazard is identified through the image processing model, the method further includes: extracting key features and performing enhancement processing on the images whose recognition result confidence is lower than a preset threshold; and inputting the enhanced images into the image processing model for secondary analysis.
[0016] Preferably, the safety hazard pictures are from industrial production scenes, construction scenes or public place monitoring scenes.
[0017] Preferably, a system for identifying safety hazards based on intelligent image analysis includes:
[0018] An image acquisition unit, used to acquire images of safety hazards to be analyzed;
[0019] A multi-scale feature extraction unit, configured to perform multi-scale feature extraction on the image to generate feature data at different levels;
[0020] A layered recognition unit, which identifies subtle safety hazards in images using low-level feature data and macro safety hazards in images using high-level feature data;
[0021] A feature fusion unit is used to fuse feature data of different levels using a feature fusion algorithm to obtain comprehensive feature data;
[0022] An image processing module is used to identify safety hazards based on the comprehensive feature data.
[0023] Preferably, it also includes: a sample construction unit, which is used to use unlabeled safety hazard image data to define different shooting condition images of the same hazard scene as positive samples, and other scene images as negative samples; a model training unit, which is used to perform unsupervised comparative learning training on the image processing module based on the positive samples and negative samples.
[0024] Preferably, it also includes: a confidence verification unit, which is used to extract key features and perform enhancement processing on pictures whose recognition result confidence is lower than a preset threshold, and feed the enhanced pictures back to the image processing module for secondary analysis.
[0025] Preferably, the safety hazard pictures are from industrial production scenes, construction scenes or public place monitoring scenes.
[0026] The technical effects and advantages of the method and system for identifying safety hazards based on AI intelligent analysis of images in the present invention are as follows:
[0027] 1. This invention, high-precision hidden danger identification: adopts multi-scale feature extraction and fusion technology, hierarchically processes image feature data, and takes into account both subtle and macro safety hazard identification. Compared with traditional single-scale analysis, it significantly improves the comprehensiveness and accuracy of hidden danger identification in complex scenarios.
[0028] 2. This invention has low data dependence: it uses an unsupervised comparative learning training scheme to build a positive and negative sample optimization model based on unlabeled data, which greatly reduces the need for manual labeling, reduces training costs and time, and enables the model to quickly adapt to new scenarios and hidden danger types.
[0029] 3. This invention has strong generalization ability: by training samples under different shooting angles, lighting conditions and other conditions, the model can grasp the invariance of hidden danger characteristics in complex environments, effectively enhancing the model's recognition ability in unlabeled scenes.
[0030] 4. This invention provides reliable early warning: the confidence verification mechanism automatically performs feature enhancement and secondary analysis on low-confidence recognition results, significantly reducing the misjudgment rate caused by image quality and improving the reliability of safety hazard early warning.
[0031] 5. This invention is adaptable across scenarios: for different industry scenarios such as industrial production and construction, the recognition strategy can be quickly optimized by expanding the feature library, adjusting parameters, etc., and has good scenario versatility.
[0032] 6. This invention, efficient system processing: The system modular design realizes multi-functional collaborative operation, combines edge computing and cloud processing, greatly improves the processing speed of a single image, supports large-scale concurrent analysis, and meets real-time monitoring needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 This is a flowchart of a method and system proposed by the present invention for identifying safety hazards based on image AI intelligent analysis. DETAILED DESCRIPTION
[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0035] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or apparatus that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or apparatus. In the absence of further restrictions, the elements defined by the sentence "include..." do not exclude the existence of other identical elements in the process, method, article or apparatus that includes the elements.
[0036] Example 1
[0037] refer to Figure 1 This embodiment provides a method and system for identifying safety hazards based on AI-powered image analysis, which is used to identify hidden dangers in industrial equipment. The specific implementation content includes:
[0038] Implementation purpose:
[0039] Verify the effectiveness of multi-scale feature fusion in identifying different levels of safety hazards in industrial scenarios, and solve the problem that traditional single-scale analysis is difficult to take into account both subtle and macro hazards.
[0040] Implementation method:
[0041] Image collection:
[0042] (1) Equipment: Hikvision DS-2CD2T47G0-I5 industrial camera (1080P, 1 / 1.8" CMOS, minimum illumination 0.001 Lux).
[0043] (2) Parameter configuration: frame rate 25fps, exposure time 1 / 50s, automatic white balance.
[0044] (3) Collection area: chemical plant reactor area (covering 5 reactors, 3 transportation pipelines, and 2 operating platforms).
[0045] Multi-scale feature extraction:
[0046] (1) Bottom-layer features (512×512): Use the Conv2_x layer of ResNet50 to extract edge and texture features.
[0047] Convolution kernel configuration: 3×3 convolution, stride 1, padding 1.
[0048] Output channels: 256.
[0049] Identification targets: loose bolts (pixel-level displacement > 0.5mm), pipe surface cracks (width > 1mm).
[0050] (2) Middle-layer features (256×256): The Conv3_x layer of ResNet50 is used to extract structural features.
[0051] Convolution kernel configuration: 3×3 convolution, stride 2, padding 1.
[0052] Output channels: 512.
[0053] Identification targets: abnormal valve opening (angle deviation > 5°), instrument pointer misalignment (scale deviation > 2 grids).
[0054] (3) High-level features (64×64): The Conv5_x layer of ResNet50 is used to extract global features.
[0055] Convolution kernel configuration: 3×3 convolution, stride 2, padding 1.
[0056] Output channels: 2048.
[0057] Identification targets: excessive material stacking on the operating platform (area share > 30%), and missing safety fences (continuous missing > 2 meters).
[0058] Feature fusion:
[0059] (1) Algorithm: Feature Pyramid Network (FPN).
[0060] (2) Fusion method: high-level features are upsampled (bilinear interpolation) to the middle-level size, and added element-by-element with the middle-level features. The merged features are further upsampled to the bottom-level size, and added element-by-element with the bottom-level features, and finally the comprehensive feature map (512×512×256) is output.
[0061] Hazard identification:
[0062] (1) Classifier: fully connected layer (256→64→12).
[0063] (2) Activation function: Softmax.
[0064] (3) Output: probability distribution of 12 types of hidden dangers (including 6 types of minor hidden dangers and 6 types of macro hidden dangers).
[0065] (4) Decision rule: select the category with the highest probability and the confidence level > 0.7 to determine that there is a hidden danger.
[0066] Implementation Effect
[0067] Accuracy: The comprehensive recognition accuracy of the test set (containing 1,000 industrial images) reached 89.2%.
[0068] Missed detection rate: The missed detection rate for minor hidden dangers is 7.3%, and the missed detection rate for macro hidden dangers is 5.2%.
[0069] Processing speed: Single image processing time (including feature extraction, fusion, and classification) is 230ms.
[0070] Compared with traditional methods: Under the same test conditions, the traditional single-scale method (using only mid-level features) has an accuracy rate of 72.3% and a missed detection rate of 27.7%.
[0071] Example 2
[0072] This embodiment provides a method and system for identifying safety hazards based on AI-powered image analysis for unsupervised comparative learning training. Specific implementation details include:
[0073] Implementation purpose: To verify the effectiveness of unsupervised contrastive learning in reducing the reliance on training data annotation and improving the generalization ability of the model, and to solve the problem that traditional supervised learning requires a large amount of manual annotation.
[0074] Implementation Method
[0075] (1) Data construction:
[0076] Data source: 3 months of historical surveillance video of an industrial park (approximately 500,000 frames).
[0077] Filtering rules:
[0078] Positive samples: images of the same device taken from different viewing angles (angle difference ≥ 15°), different lighting conditions (brightness difference ≥ 20%), and different distances (object distance difference ≥ 1 meter).
[0079] Negative samples: randomly select images of devices that are not of the same type.
[0080] Data augmentation: rotation (±30°), scaling (0.8-1.2 times), and brightness adjustment (±30%) are applied to the original images.
[0081] Final sample set: 80,000 pairs of positive samples and 20,000 pairs of negative samples.
[0082] (2) Model training:
[0083] Network architecture: Siamese Network.
[0084] Feature extractor: ResNet50 (pre-trained on ImageNet).
[0085] Projection head: MLP (2048 → 512 → 128).
[0086] Loss function: NT-Xent loss where z i 、z j is the feature vector of the positive sample pair, is the cosine similarity, T=0.5 is the temperature parameter, and N=128 is the batch size.
[0087] Optimizer: Adam (learning rate 1e-4, weight decay 1e-6).
[0088] Training configuration:
[0089] Training period: 100 epochs.
[0090] Batch size: 256 (128 pairs of samples).
[0091] Hardware: NVIDIA V100 GPU × 4.
[0092] Training time: approximately 36 hours.
[0093] (3) Model evaluation
[0094] Evaluation Metrics:
[0095] Feature space clustering purity: the average cosine similarity of similar hidden danger features in the feature space.
[0096] Zero-shot transfer accuracy: the recognition accuracy in new unlabeled scenes.
[0097] Evaluation results:
[0098] Cluster purity: increased from 0.68 in supervised learning to 0.92.
[0099] Zero-shot transfer accuracy: reached 81.5% on the new plant test set (supervised learning model only 53.2%).
[0100] Implementation effect:
[0101] Labeling cost: Reduce labeling workload by 80% compared to traditional supervised learning.
[0102] Generalization capability: The recognition accuracy of unseen equipment types (such as new reactor models) has increased by 42%.
[0103] Training efficiency: The training time required to achieve the same accuracy is reduced by 30%.
[0104] Storage requirements: Model parameters are reduced by 15% (by optimizing feature representation through contrastive learning).
[0105] Example 3
[0106] This embodiment provides a method and system for identifying safety hazards based on AI-powered image analysis, which is suitable for construction scene adaptation. Specific implementation details include:
[0107] Implementation purpose: To verify the adaptability of the technical solution of the present invention in construction scenarios, and to optimize the recognition effect based on the characteristics of the scenario (high-altitude work, dynamic construction, and complex lighting).
[0108] Implementation steps (1) Scenario analysis
[0109] Typical hidden danger types:
[0110] Minor hidden dangers: The safety helmet is not worn properly (the chin strap is not buckled) and the seat belt buckle is damaged.
[0111] Macro-level hidden dangers: external scaffolding wall connections are missing and material loading exceeds the safety line.
[0112] Environmental Challenges:
[0113] High-altitude aerial photography (shooting height 50-100 meters).
[0114] Dynamic construction (frequent movement of personnel and equipment).
[0115] Complex lighting (shadows, reflections).
[0116] (2) Technical adjustments
[0117] Feature extraction optimization:
[0118] Bottom-layer feature enhancement: Add LBP (local binary pattern) texture features to strengthen the recognition of helmet logos and damaged seat belt webbing.
[0119] High-level feature optimization: Use dilated convolution (expansion rate = 2) to replace some pooling layers to maintain the resolution of the feature map while expanding the receptive field.
[0120] Model training enhancements:
[0121] MixUp data enhancement technology is introduced to synthesize mixed samples of standard / irregular helmet wearing.
[0122] Adjust the weight of the loss function: assign 3 times the weight to the hidden dangers related to high-altitude work (such as not wearing a safety belt).
[0123] Real-time processing optimization:
[0124] The inter-frame difference method is used to detect dynamic areas, giving priority to areas with frequent human activities.
[0125] Implement model quantization (INT8) to compress single-frame processing time to 150ms.
[0126] (3) System deployment
[0127] Hardware configuration:
[0128] DJI M300RTK drone (with Zenmuse H20N camera).
[0129] Edge computing unit: NVIDIA Jetson AGX Orin (64TOPS computing power).
[0130] Network architecture:
[0131] 5G backhaul (upload processed feature data, not the original image).
[0132] Deploy the multi-scale fusion model on the cloud (8-card V100 GPU cluster).
[0133] Implementation effect:
[0134] Scenario adaptability: In a test at a 30-story residential construction site, the overall recognition accuracy reached 93.7%.
[0135] Dynamic response: The response time to new violations (such as entering the work area without wearing a safety helmet) is less than 2 seconds.
[0136] Anti-interference ability: Under strong backlight conditions (contrast ratio > 1000:1), the misjudgment rate is controlled within 5%.
[0137] Improved practicality: Through drone inspections, the daily coverage area can be increased by 400% compared to fixed cameras.
[0138] Example 4
[0139] This embodiment provides a method and system for identifying safety hazards based on AI-powered image analysis for confidence verification. Specific implementation details include:
[0140] Implementation purpose: To verify the effectiveness of the confidence verification mechanism in improving recognition reliability and reducing the false positive rate, and to solve the false positive problem caused by low-quality images (such as low light and blur).
[0141] Implementation steps (1) Confidence assessment model
[0142] Input features:
[0143] Softmax probability distribution of the classifier output.
[0144] The entropy of the feature map.
[0145] Image clarity index (Laplacian operator variance).
[0146] Evaluation function: Confidence = w1·MaxProb + w2·(1-Entropy) + w3·Sharpness, where w1 = 0.6, w2 = 0.2, and w3 = 0.2 are weight relationships, and the threshold setting is 0.7 (which can be dynamically adjusted according to scenario risks).
[0147] (2) Low-confidence image processing process
[0148] Feature Enhancement Module:
[0149] Low light processing: adaptive histogram equalization (CLAHE).
[0150] Blur restoration: Image deblurring based on guided filtering.
[0151] Detail enhancement: non-local means filtering (NLM).
[0152] (3) Secondary analysis strategy:
[0153] The enhanced image is re-input into the original recognition model.
[0154] If the secondary confidence level is greater than 0.8, the hidden danger is confirmed; otherwise, it is marked as "manual review required".
[0155] Experimental verification:
[0156] Test set: Contains 2,000 low-quality images (40% low-light, 30% blur, and 30% occlusion).
[0157] Comparison plan:
[0158] Baseline solution: directly output the first recognition result.
[0159] Improvement plan: Apply confidence verification mechanism.
[0160] Implementation effect:
[0161] Typical case: In the fire hydrant obstruction detection in the dark environment of the underground garage, the misjudgment rate was reduced from 21% to 4%.
[0162] Performance overhead: Adds approximately 0.7 seconds of processing latency (mainly from feature enhancement).
[0163] Improved reliability: During continuous operation tests in high-risk scenarios (such as chemical plant areas), no false alarms due to misjudgment occurred.
[0164] index Baseline scenario Improvement plan Improvement False positive rate 12.5% 2.9% 76.8% Manual review rate 35.2% 18.7% 46.9% Average response time 1.8 seconds 2.5 seconds +38.9%
[0165] Example 5
[0166] This embodiment provides a method and system for identifying security risks based on AI-powered image analysis, which is used for system software and hardware deployment. Specific implementation details include:
[0167] Implementation purpose: To verify the engineering feasibility of the technical solution of the present invention in an actual industrial environment and to achieve full-link optimization from front-end collection to cloud analysis.
[0168] Implementation method:
[0169] (1) Edge layer design
[0170] Hardware selection:
[0171] Camera: Hikvision DS-2CD2T47G0-I5 (1080P, supports wide dynamic range 120dB).
[0172] Edge computing unit: NVIDIA Jetson AGX Orin (32-core ARM Cortex-A78AE CPU, 64TOPS computing power).
[0173] Functional division:
[0174] Original image acquisition: 1920×1080@25fps.
[0175] Preprocessing: Gaussian filtering for noise reduction and ROI cropping (reducing data volume by 50%).
[0176] Feature extraction: Complete the extraction of underlying features (512×512).
[0177] Data compression: Use JPEG 80% quality compression (size reduced from 3MB to about 300KB).
[0178] (2) Transport layer optimization
[0179] Communication protocol: MQTT over TLS.
[0180] Data Strategy:
[0181] Normal status: Feature data is uploaded every 5 seconds.
[0182] Abnormal status: Immediately upload full-resolution images and trigger an alarm.
[0183] Bandwidth requirements:
[0184] Normal status: about 150kbps / channel.
[0185] Abnormal status: about 8Mbps / channel (lasts 30 seconds).
[0186] (3) Cloud-layer architecture
[0187] Hardware configuration:
[0188] Server: Dell PowerEdge R750xa (2 x Intel Xeon Platinum 8380 CPU, 512 GB RAM).
[0189] GPU cluster: 8 NVIDIA A100 cards (80GB HBM2 per card).
[0190] Software modules:
[0191] Multi-scale feature extraction and fusion: built on TensorFlow 2.8.
[0192] Distributed reasoning: Horovod is used to achieve model parallelism.
[0193] Database: MongoDB (stores historical hidden danger data).
[0194] Visualization platform: developed based on Django and ECharts.
[0195] Performance indicators:
[0196] Single cluster processing capacity: 2000 concurrent analyses.
[0197] Average response time: <1 second (from image acquisition to alarm generation).
[0198] (4) Application Cases
[0199] Deployment scenario: A large chemical plant (covering three production workshops and 12 storage tank areas).
[0200] Terminal configuration:
[0201] Management center: 55-inch 4K LCD splicing screen.
[0202] Mobile terminal: explosion-proof smartphone (supports vibration, sound and light alarms).
[0203] Alarm mechanism, three-level alarm:
[0204] Level 1 (red): Deal with it immediately (such as a pipeline leak).
[0205] Level 2 (yellow): Handle within 24 hours (such as safety signs are damaged).
[0206] Level 3 (blue): Process within 72 hours (such as rust on the equipment surface).
[0207] Alarm method: SMS, APP push, and broadcast system linkage.
[0208] Implementation Effect
[0209] System stability: Continuous operation for 30 days without failure, availability > 99.9%.
[0210] Processing efficiency: End-to-end processing time for a single image is less than 800ms.
[0211] Resource utilization: Edge CPU utilization is less than 60%, and GPU utilization is less than 50%.
[0212] Actual value: Within three months of deployment, three potential safety hazards that could have led to production stoppages were identified and avoided, reducing economic losses by approximately 5 million yuan.
[0213] Comparative Example 1
[0214] This comparative example provides a single-scale traditional approach, with specific implementation details including:
[0215] Implementation purpose: To compare the performance differences between the traditional single-scale analysis method and the multi-scale fusion method of the present invention, and to verify the necessity of multi-scale analysis.
[0216] Implementation steps:
[0217] (1) Technical solution
[0218] Feature extraction: Only the Conv3_x layer of ResNet50 (256×256 pixel features) is used.
[0219] Classifier: fully connected layer (512→128→12).
[0220] Training data: the same annotated set as in Example 1 (5000 images).
[0221] Training method: Supervised learning (cross entropy loss).
[0222] (2) Experimental configuration
[0223] Test set: 1000 industrial images, the same as in Example 1.
[0224] Evaluation indicators: accuracy, missed detection rate, and false positive rate.
[0225] Implementation effect:
[0226] There is a serious lack of recognition of minor hidden dangers such as loose bolts (accounting for 62% of the total number of missed detections).
[0227] There is insufficient identification of macro-level hidden dangers such as local blockage of channels (accounting for 38% of the total number of missed detections).
[0228] Performance bottleneck:
[0229] Insufficient feature resolution: 256×256 pixels cannot capture subtle defects smaller than 5 pixels.
[0230] Weak global perception capabilities: The lack of high-level features leads to inaccurate judgments of scene-level hidden dangers.
[0231] Comparative Example 2
[0232] This comparison provides a traditional fully supervised model, and the specific implementation includes:
[0233] Implementation purpose: To compare the performance differences between the traditional fully supervised learning method and the unsupervised comparative learning method of this invention, and to verify the advantages of unsupervised learning in reducing labeling dependence.
[0234] Implementation steps:
[0235] (1) Technical solution
[0236] Model architecture: ResNet50+MLP, same as in Example 2.
[0237] Training data:
[0238] Annotation set: 50,000 sheets (≥5,000 sheets for each type of hidden danger annotation).
[0239] Labeling cost: approximately RMB 250,000 (calculated at RMB 5 per sheet).
[0240] Training methods:
[0241] Loss function: cross entropy loss.
[0242] Optimizer: Adam (learning rate 1e-4).
[0243] Training period: 100 epochs.
[0244] (2) Experimental configuration
[0245] Test set 1: the same validation set as Example 2.
[0246] Test set 2: unlabeled new factory images (1000 pictures).
[0247] Evaluation indicators: accuracy, training cost, and zero-shot transfer capability.
[0248] Implementation effect:
[0249] High data dependence: The labeling cost is too high, making it difficult to quickly adapt to new scenarios.
[0250] Weak generalization ability: poor ability to identify new unlabeled hidden danger types (such as new types of violations).
[0251] High resource consumption: The model has a large number of parameters and is difficult to deploy on edge devices.
[0252] Experimental Example 1
[0253] This experimental example is used to test the experiment and compare the results. The specific contents include:
[0254] Experimental design
[0255] (1) Dataset
[0256] Training set: 5,000 annotated images (Examples 1-4) vs 50,000 annotated images (Comparative Example 2).
[0257] Test set: 2,000 industrial scene images (including 12 categories of hidden dangers, approximately 167 images per category).
[0258] Evaluation Metrics:
[0259] Accuracy rate, missed detection rate, and misjudgment rate.
[0260] F1 score (taking into account both precision and recall).
[0261] Training efficiency (number of labeled images / unit accuracy improvement).
[0262] (2) Experimental environment
[0263] Hardware: NVIDIA V100 GPU × 4, Intel Xeon Silver 4214 CPU, 256GB RAM.
[0264] Software: Ubuntu 20.04, TensorFlow 2.8, PyTorch 1.12.
[0265] Compared with Examples 1-5 and Comparative Examples 1-2, the embodiments of the present invention have multi-scale advantages:
[0266] Compared with Comparative Example 1, the accuracy of Example 1 is improved by 23.4%, and the missed detection rate is reduced by 64.6%.
[0267] It is proved that multi-scale feature fusion effectively solves the limitations of single-scale analysis.
[0268] Unsupervised learning value:
[0269] The training data of Example 2 is reduced by 96% compared with Comparative Example 2, but the accuracy is improved by 11.2%.
[0270] The significant advantages of unsupervised contrastive learning in reducing labeling dependence are verified.
[0271] Confidence verification effect:
[0272] The misjudgment rate of Example 3 is reduced by 55.4% compared with Example 1.
[0273] It proves that the secondary verification mechanism effectively improves the system reliability.
[0274] Scene adaptability:
[0275] The accuracy of Example 4 in the architectural scene (93.7%) is higher than that in the general scene (89.2% in Example 1).
[0276] The effectiveness of scene customization optimization was verified.
[0277] The present invention is superior to traditional methods in terms of recognition accuracy, data efficiency, scene adaptability and reliability through the three core technologies of multi-scale feature fusion, unsupervised comparative learning, and confidence verification mechanism. The progressive optimization of Examples 1-5 proves that multi-scale analysis is the basis for the identification of hidden dangers in complex scenes, unsupervised learning greatly reduces the application threshold, and scene customization and reliability enhancement further enhance the value of technology implementation. In contrast, due to technical limitations, Examples 1 and 2 will face the bottleneck of high cost and low efficiency in the application of industrial Internet, smart construction sites and other fields. The present invention provides a more efficient, more economical and more reliable solution for the intelligent identification of safety hazards, and has significant technical innovation and industrial application value.
[0278] The above embodiments may be implemented in whole or in part through software, hardware, firmware or any other combination. When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product.
[0279] Those skilled in the art will appreciate that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented with electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0280] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0281] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited to this. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
[0282] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for identifying safety hazards based on AI intelligent analysis of images, characterized in that: The following steps are involved: S1: Obtain the safety hazard image to be analyzed; S2: Perform multi-scale feature extraction on the image to generate feature data at different levels; S3: Identify subtle security risks in images using low-level feature data and identify macro security risks in images using high-level feature data. S4: Use feature fusion algorithm to fuse feature data at different levels to obtain comprehensive feature data; S5: Based on the comprehensive feature data, safety hazard identification is performed through an image processing model.
2. The method for identifying safety hazards based on image AI intelligent analysis as claimed in claim 1, characterized in that: It also includes: using unlabeled safety hazard image data, defining different shooting condition images of the same hazard scene as positive samples, and defining other scene images as negative samples; based on the positive samples and negative samples, performing unsupervised comparative learning training on the image processing model.
3. The method for identifying safety hazards based on image AI intelligent analysis as claimed in claim 1, characterized in that: The different shooting conditions include at least one of a shooting angle, a light parameter, and a shooting distance.
4. The method for identifying safety hazards based on image AI intelligent analysis as claimed in claim 1, characterized in that: The minor safety hazards include damaged appearance of fire-fighting facilities and blurred escape signs; the macro safety hazards include blocked evacuation passages and abnormal structures of large equipment.
5. The method for identifying safety hazards based on image AI intelligent analysis as claimed in claim 1, characterized in that: After the safety hazard is identified through the image processing model, the method further includes: extracting key features and performing enhancement processing on the images whose recognition result confidence is lower than a preset threshold; and inputting the enhanced images into the image processing model for secondary analysis.
6. The method for identifying safety hazards based on image AI intelligent analysis as claimed in claim 1, characterized in that: The safety hazard pictures are from industrial production scenes, construction scenes or public place monitoring scenes.
7. A system for identifying safety hazards based on intelligent image analysis, characterized in that: include: An image acquisition unit, used to acquire images of safety hazards to be analyzed; A multi-scale feature extraction unit, configured to perform multi-scale feature extraction on the image to generate feature data at different levels; A layered recognition unit, which identifies subtle safety hazards in images using low-level feature data and macro safety hazards in images using high-level feature data; A feature fusion unit is used to fuse feature data of different levels using a feature fusion algorithm to obtain comprehensive feature data; An image processing module is used to identify safety hazards based on the comprehensive feature data.
8. The system for identifying safety hazards based on image AI intelligent analysis as described in claim 7, characterized in that: Also includes: A sample construction unit is used to use unlabeled safety hazard image data to define images of the same hazard scene under different shooting conditions as positive samples, and images of other scenes as negative samples; A model training unit is used to perform unsupervised contrastive learning training on the image processing module based on the positive samples and negative samples.
9. The system for identifying safety hazards based on AI intelligent analysis of images as described in claim 7, characterized in that: It also includes: a confidence verification unit, which is used to extract key features and perform enhancement processing on pictures whose recognition result confidence is lower than a preset threshold, and feed the enhanced pictures back to the image processing module for secondary analysis.
10. The system for identifying safety hazards based on image AI intelligent analysis as claimed in claim 7, characterized in that: The safety hazard pictures are from industrial production scenes, construction scenes or public place monitoring scenes.