New energy construction small-target full-link safety management and control method based on cross-view feature collaboration and dynamic anti-interference closed loop
By using multi-source data fusion and dynamic calibration technology, the problem of identifying and controlling small targets in new energy construction has been solved, and accurate identification and full-chain safety management have been achieved in dynamic environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-02
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies struggle to effectively identify and manage small targets, such as safety rope clips and photovoltaic panel mounting bolts, during new energy construction. This is especially true in multi-view and dynamic environments where feature loss is severe, making it impossible to achieve end-to-end safety management.
By deploying various heterogeneous sensing devices, collecting multi-source data, performing coordinate mapping and feature decomposition and fusion, and combining dynamic sensing data for calibration, we can achieve accurate identification of small targets and risk level determination, generate rectification instructions, and complete the closed loop of safety management and control.
It enables accurate identification and risk assessment of small targets in complex construction environments, suppresses interference from dynamic lighting and equipment vibration, constructs a full-link safety management process, and improves the robustness and real-time performance of detection.
Smart Images

Figure CN121811010A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of new energy construction safety management, and in particular to a new energy construction small target full-link safety management and control method based on cross-view angle feature cooperation and dynamic anti-disturbance closed loop. BACKGROUND
[0002] With the rapid development of global new energy industry, new energy construction projects such as photovoltaic power stations and wind power plants are often distributed in complex terrains such as low-latitude high-altitude mountainous areas and desertified lands. During the construction process, a large number of small targets such as safety rope buckles, photovoltaic panel installation bolts and wind power tower connection bolts need to be managed and controlled. These small targets only occupy tens or even a few pixels in high-resolution images. They have simple structures and scarce texture features, and have very low contrast with backgrounds such as gray supports and reflective photovoltaic panels. They are easily ignored or misjudged as noise by neural networks.
[0003] To solve the safety management and control problem of power operation scenes, the prior art discloses CN120220241A "Power operation personnel behavior recognition and early warning system and method based on video analysis". This technical solution realizes multi-source video acquisition through fixed monitoring cameras, wearable camera equipment and unmanned aerial vehicle camera equipment. YOLOv8 / EfficientDet deep learning target detection algorithm, OpenPose / HRNet human pose estimation algorithm and TSM / SlowFast time sequence behavior analysis algorithm are adopted. Multi-modal data processing technology combining Kalman filtering and Bayesian fusion is used to identify the illegal behavior of power operation personnel and give hierarchical warning. At the same time, a three-dimensional virtual operation environment is constructed through digital twinning technology to realize remote monitoring. Although this scheme has certain effect on the identification of conventional targets in power operation, it mainly aims at the power operation scene and only realizes regional coverage through fixed, wearable and unmanned aerial vehicle equipment. It does not construct exclusive feature coding logic for the "local structure feature" of new energy construction small targets. Small targets are severely compressed into textureless dots or short lines in the top view / perspective view. Small targets in the first view may appear at any angle and scale and are accompanied by severe motion blur. In addition, there are lens distortion, resolution difference and time synchronization error of different devices. No view angle coordinate mapping or feature alignment mechanism is designed for the coordinate dislocation and feature compression of small targets in different view angles such as top view dots and side view short lines. It is impossible to synergistically enhance the features of multi-view small targets. Instead, the dislocation in time and space aggravates the loss of small target features. In addition, dynamic light and device vibration further destroy the weak signal of small targets. Strong light can erase the texture and distort the color of small targets. That is, a small target with a special marker color may become a black silhouette in backlight, and all its color and texture features are lost. A small target of 10x10 pixels can be blurred into a mass by slight vibration, completely losing its edge and shape features. SUMMARY
[0004] In view of the deficiencies of the prior art, the technical problem to be solved by the present application is to provide a new energy construction small target full-link safety management and control method based on cross-view feature cooperation and dynamic anti-interference closed loop, which can accurately identify small construction targets such as bolts and buckles, solve the problem of cross-view space dislocation of multi-source equipment, suppress dynamic light and equipment vibration interference, and realize full-link management and control from data acquisition, feature enhancement, precision calibration to rectification closed loop.
[0005] To solve the above technical problems, the technical scheme adopted by the present application is as follows: The present application provides a new energy construction small target full-link safety management and control method based on cross-view feature cooperation and dynamic anti-interference closed loop, comprising the following steps: S1, deploying multiple heterogeneous sensing devices to collect multi-source data of small targets in the construction area; S2, according to the multi-source data collected in step S1, through coordinate mapping and feature decomposition fusion, the small target features of different views are unified to the global coordinate system of the construction area, and the aligned small target feature map is output; S3, according to the aligned small target feature map output in step S2, according to the scale and texture characteristics of the small target, adaptively perform multi-scale feature enhancement and lightweight processing, and output the initial identification result of the small target; S4, according to the initial identification result of step S3, when the confidence is insufficient, call the dynamic perception data to calibrate the feature map, correct the identification confidence and output the final identification result with three-dimensional position and risk level; S5, according to the final identification result output in step S4, generate and execute the hazard rectification instruction, and complete the safety management and control closed loop.
[0006] In the preferred scheme, S1, the multi-source data of small targets in the construction area is collected, and the multi-source data of small targets includes event stream data, local texture image data and thermal contour data, which are as follows: Deploy an event camera array, a tracked mobile robot and an infrared thermal imaging sensor to collect data according to the "device and small target" matching rule; the event camera array is arranged at equal intervals in the high-risk area of the foundation pit and the tower crane, and a luminance change threshold is set When the luminance change of the small target is detected to be greater than the luminance change threshold , the "time-pixel-luminance change" event stream data of the small target is output; Wherein, the value range of the interval of the event camera array is 4~6m, and preferably 5m; the value range of the luminance change threshold is 0.08~0.12, i.e. 8%~12%; the data format of the event stream data is <timestamp (x coordinate (pixel), y coordinate (pixel), luminance polarity> ; The tracked mobile robot is equipped with a high-resolution macro camera, and when receiving the GPS inspection path instruction of the construction area, the local texture image data of the construction area is collected. Among them, the camera resolution of the macro camera is 3840x2160~4096x2160, preferably 4K resolution. Set the illumination threshold When the infrared thermal imaging sensor detects that the ambient light intensity exceeds the illumination threshold , the thermal profile data of the small target device is output; Among them, the value range of the illumination threshold is 6000~8000lux, preferably 7000lux; the temperature resolution , The value range is 0.05~0.15℃, and the default is 0.1℃, which is used to supplement the lost profile features under visible light.
[0007] In the preferred scheme, S21, the installation angle of the event camera in step S1, the GPS coordinates of the tracked mobile robot, set the coordinate deviation threshold When the calculated coordinate deviation of the small target collected by different devices exceeds the coordinate deviation threshold Pixel, construct a perspective coordinate mapping matrix; Among them, the accuracy of the installation angle is ±0.05 ~±0.15 , preferably ±0.1 ; the accuracy of the GPS coordinates of the tracked mobile robot is ±0.3~±0.7m, preferably ±0.5m; the value range of the coordinate deviation threshold is 0.8~1.2 pixels, preferably 1 pixel, based on the coordinate comparison algorithm output; the perspective coordinate mapping matrix is a 3x3 rotation matrix And a 3x1 translation vector .
[0008] S22, using the adapted lightweight cross-layer feature fusion algorithm (CLFM), performing db4 wavelet decomposition on the event stream data, local texture image data, and thermal profile data of the small target device in step S1, decomposing the image into Layers, separate high-frequency texture features and low-frequency profile features; Among them, is the number of decomposition layers, the value range is 3~5 layers, preferably 4 layers; S23, based on the perspective coordinate mapping matrix, when the analysis shows that the feature scale difference of different perspectives exceeds the scale difference threshold At this time, the features are unified to the global coordinate system of the construction area through a dynamic alignment and scale refinement algorithm (DASR); A new energy small target feature dictionary is obtained, and a similarity threshold is set When the cosine similarity between the high-frequency features after wavelet decomposition and the new energy small target feature dictionary is calculated to be less than , the embedded feature dictionary guide module focuses on the local structure features of the small target, and outputs the aligned small target feature map; The scale difference threshold is in the range of 4% to 6%, preferably 5%, and is output based on a feature scale analysis algorithm; the similarity threshold is in the range of 0.75 to 0.85, preferably 0.8, and is output based on a feature matching algorithm; the channel number of the small target feature map is in the range of 200 to 300, preferably 256.
[0009] In the preferred scheme, S31, the small target feature map aligned in step S23 is input into a spatial pyramid multi-scale general convolution algorithm (SPMCC) for analysis, and a small target box area threshold When the analysis shows that the area of the small target candidate box is pixels², a convolution kernel is used; When the area of the candidate box is pixels², a convolution kernel is used, and only the local feature enhancement is performed on the small target candidate box area to obtain the small target feature map after local enhancement; The small target box area threshold is in the range of 14 to 18 pixels², preferably 16 pixels², and is output based on a candidate box detection algorithm; The receptive field of the convolution kernel is , in the range of 4 to 6, preferably 5; The receptive field of the convolution kernel is , in the range of 6 to 8, preferably 7; the convolution step length of the local feature enhancement is in the range of 1 to 2, preferably 1; and Padding=1.
[0010] S32, according to the small target feature map after local enhancement, a dynamic upsampling operator Dysample is used for analysis, and a contrast threshold When the analysis shows that the small target texture contrast is less than the contrast threshold , a 2X upsampling rate is used; when the contrast is greater than or equal to the contrast threshold , a 1X rate is used, and a lightweight small target feature map and an initial recognition result are output; The value range of the contrast threshold is 0.25-0.35, preferably 0.3, based on the texture analysis algorithm output; the value range of the channel number of the lightweight small target feature map is 100-150, preferably 128; the initial recognition result includes position <x, y, w, h>, category and confidence <0-1>; In the preferred scheme, S41, a confidence threshold is set, and when the initial recognition result confidence of the judgment step S32 is less than the confidence threshold, the event stream data of step S1 is called, and the dynamic edge features of the small target are extracted according to the preset sampling frequency; The value range of the confidence threshold is 0.65-0.75, preferably 0.7, based on the confidence determination algorithm output; The value range of the sampling window frame is 0.8-1.2 MHz, preferably 1 MHz, and the sampling window frame is 8-12 frames, preferably 10 frames.
[0011] S42, according to the extracted dynamic edge features, the mapping relationship between image features and event stream features is constructed, and the dynamic edge features are converted into image domain edge enhancement weight , and local calibration is performed on the lightweight small target feature map of step S32; The calculation formula of the image domain edge enhancement weight is: Cont is the dynamic edge continuity, and the value range is 0-1; is a weight coefficient, and the value range is 0.08-0.12, preferably 0.1.
[0012] S43, according to the calibrated lightweight small target feature map, the small target recognition confidence is corrected, and the final recognition result is output, including: small target 3D position , , and risk level; The risk level is divided into: confidence is low risk, confidence is medium risk, and confidence is high risk; The value range of the confidence threshold is 0.75-0.85, preferably 0.8; The value range of the confidence threshold
[0013] In the preferred embodiment, step S5 further includes the following steps: S51. Construct a digital twin model of the construction area. Based on the final identification results of S4, map the 3D location and risk level of the small target to the virtual model in real time and mark the specific construction location of the small target. S52. Automatically match rectification plans based on risk levels: If the risk level is high, then the confidence level is < This triggers the generation of an "immediate shutdown" command, calls the wind power construction torque standard library, and outputs the calibration parameters for small target fasteners. The calibration parameters for small target fasteners include bolt torque, clip tightening force, and connector preload. Model bolt torque value The value range is 35~45 N·m, preferably 40 N·m; Model bolt torque value The value range is 55~65 N·m, preferably 60 N·m; If it is medium risk, then Confidence level < ,trigger" The "re-inspect within 24 hours" instruction is associated with the re-inspection path of the tracked mobile robot; in, The time range is 0.8 to 1.2 hours, with 1 hour being preferred; the re-inspection path is planned using a digital twin model to avoid obstacles; The system calls upon the skeletal feature data of the workers in the digital twin model to extract the 3D coordinates of key human body parts such as the shoulders, waist, and chest; obtains the spatial coordinates of the safety belt suspension point through a safety belt recognition algorithm; compares the 3D coordinates of the key human body parts with the spatial coordinates of the safety belt suspension point according to the preset human safety suspension area standard, generates comparison results, triggers a "safety belt suspension violation" judgment based on the comparison results, and outputs a rectification instruction to "adjust the suspension point to the human safety suspension area". The comparison results include: Not suspended: No valid safety belt suspension point coordinates were detected; Improper suspension point position: The suspension point coordinates fall outside the human safety suspension area. In the preferred solution, the worker's waist safety belt suspension point coordinates were detected. Below waist coordinates , , The value range is -5 to 5cm, with 0 being the preferred value. This triggers a "low-hanging, high-use" violation check, outputting "Adjust the suspension point to above waist level (i.e.,...") , The rectification instruction was: "Value range 8~12cm, default 10cm".
[0014] S53, after completing rectification according to the rectification instruction, generating rectification record and rectification confirmation signal, controlling the tracked mobile robot to collect the target small target again according to the rectification confirmation signal, generating a secondary identification result, repeating steps S2 to S4, and when it is judged that the confidence of the secondary identification result is greater than or equal to , the hidden danger closed loop archiving is completed.
[0015] In the preferred scheme, the specific steps of the "adapted cross-layer feature fusion module (CLFM)" in step S22 are as follows: Delete the general cross-modal alignment logic in the original CLFM module of COXNet, and add a small target feature matching layer; Set the similarity threshold value When the calculated cosine similarity between the high-frequency feature after wavelet decomposition and the new energy small target feature dictionary is less than the similarity threshold value , trigger the GeoShape similarity measurement label assignment strategy, calculate the spatial shape similarity between the feature and the template, and the formula is as shown in the following formula (1): (1); In formula (1), represents the spatial shape similarity between the high-frequency feature and the template feature, and the value range is 0~1; represents the shape parameter of the high-frequency feature after wavelet decomposition in step S22; represents the template shape parameter in the new energy small target feature dictionary; represents the dimension number of the feature parameter; the value range of the similarity threshold value is 0.75~0.85, and the preferred value is 0.8, based on the feature matching algorithm output of step S22.
[0016] When it is determined that , the feature is retained, otherwise it is filtered as background noise, and it is ensured that the small target feature is not misjudged.
[0017] In the preferred scheme, the specific steps of the spatial pyramid multi-scale general convolution (SPMCC) in step S31 are as follows: According to the small target feature map aligned in step S23, the feature map is executed by the spatial pyramid multi-scale general convolution algorithm: divided into 1X, 2X, and 4X scale layers; wherein, the original resolution of 1X is 640×640; the original resolution of 2X is 1280×1280; the original resolution of 4X is 2560×2560; When it is analyzed that the area of the small target candidate box is pixels², call the convolution kernel to extract local texture features, the convolution step is , and the weight coefficient is ; When the analysis result is a small target candidate box area pixel², call the convolution kernel to extract the contour feature, the convolution step size , the weight coefficient ; The convolution results of the three scale layers are weighted and fused, and the weighted fusion formula is shown in the following formula (2): (2); In formula (2), denotes the small target feature map after weighted fusion; denotes the weight coefficient of the convolution result; denotes the output feature of the convolution kernel; denotes the weight coefficient of the convolution result; denotes the output feature of the convolution kernel; The output is the small target feature map after local enhancement.
[0018] In the preferred scheme, the specific steps of confidence correction in step S43 are as follows: When the initial recognition result confidence of step S32 is determined , trigger event flow data call, extract the dynamic edge feature of the small target, calculate the time continuity of the edge feature, and the time continuity calculation formula is shown in the following formula (3): (3); In formula (3), Cont denotes the time continuity of the small target dynamic edge, and the value range is 0~1; denotes the time window length, and the value range is 8~12 frames, preferably 10 frames; denotes the edge feature value at the t time in step S41; denotes the edge feature value at the t-1 time in step S41; denotes the frame number in the time window, and the value range is 1~ ; Let the continuity threshold be , when judging , construct the image feature and event flow feature mapping relationship, and convert the dynamic edge feature into image domain edge enhancement weight ; Wherein, the value range of continuity threshold is 0.45~0.55, preferably 0.5, based on the continuity determination result; The value range is 0.08~0.12, preferably 0.1. confidence of the initial recognition result output according to step S32 and the image domain edge enhancement weight , to obtain the modified small target recognition confidence The confidence correction formula is shown in the following formula (4): (4); In formula (4), is the modified small target recognition confidence, with a value range of 0-1; is the initial recognition confidence output according to step S32; is the edge enhancement weight calculated based on dynamic edge features; When it is determined that , it is determined as low risk; when , it is determined as medium risk; when , it is determined as high risk, triggering a high-priority rectification instruction, i.e., only SMS is pushed to the on-site person in charge.
[0019] In the preferred scheme, step S1 further includes "device priority scheduling and edge computing optimization", and the specific steps are as follows: When concurrent collection of event cameras, tracked mobile robots, and infrared thermal imaging sensors is monitored, small target data in high-risk areas is preferentially processed, and the processing priority is event cameras > tracked mobile robots > infrared thermal imaging sensors; Let the CPU occupancy threshold be When the edge end CPU occupancy rate is monitored to be greater than , trigger the Camera-LiDAR multi-sensor fusion layer to generate a shared feature map as a mask , and the shared feature mask calculation formula is shown in the following formula (5): (5); In formula (5), is the shared feature mask, with a value range of 0-1, 0 being the background and 1 being the small target area; is the Sigmoid activation function, used to map the convolution result to the range of 0-1; is the convolution operation; is the camera image feature collected in step S1; is the LiDAR point cloud feature carried by the tracked mobile robot; the value range of the CPU occupancy threshold is 65%-75%, with a default of 70%, and is output based on system resource monitoring; When the small target recognition frequency of the low-risk area is monitored to be greater than frames / second, the collection interval is adjusted to Reduce computing load at the edge by seconds / frame.
[0020] in, Frames per second The value range is 28~32 frames / second, preferably 30 frames / second, based on the identification frequency monitoring output; The time for 1 frame The value range is 1.8 to 2.2 seconds, preferably 2 seconds; In the preferred embodiment, the model training phase in step S31 also includes the calculation of scale-based dynamic loss (SD Loss), with the specific steps as follows: When analysis reveals that small bounding box areas exist in the training samples, When SD Loss calculation is triggered, the dynamic weight coefficients of SD Loss are calculated. The calculation formula is shown in equation (6) below: (6); In equation (6), This is represented as the dynamic weighting coefficient of SDLows, with a value range of 0.5 to 1.2. It is represented as the weighting benchmark coefficient, with a value range of 1.1 to 1.3, preferably 1.2; Represented as the area of a small bounding box in the training samples, satisfying , The value ranges from 3 to 5 pixels², with 4 pixels² being preferred. When the analysis shows that the area of the small target box is ≥ hour, (Using conventional cross-entropy loss;) The total loss function is shown in equation (7) below: (7); In equation (7), This is expressed as the total loss value during model training; Represented as the dynamic weighting coefficient of SDLoss; This is represented as the small target bounding box loss, using IoU loss, with a value range of 0 to 1; This is represented as the small object classification loss, using cross-entropy loss, with a value range of 0 to +∞; The value ranges from 0.75 to 0.85, with 0.8 being the preferred value; When the number of model training iterations is detected = When it's time, The learning rate is reduced to the initial value. This is to avoid model overfitting; in, For the number of iterations, The value range is 90-110 rounds, preferably 100 rounds, and the output is monitored based on the training progress; , The value range is 0.0008-0.0012, preferably 0.001; The value range is 0.45-0.55, preferably 0.5.
[0021] In the preferred scheme, the specific steps of the digital twin model mapping process in step S51 are as follows: According to the small target 3D position in step S43, the small target 3D coordinates are mapped to the corresponding components of the BIM model through the coordinate conversion matrix The coordinate conversion matrix formula is shown in formula (8) as follows: (8); In formula (8), , , represents the 3D coordinates of the small target in the BIM model; represents a 3×3 coordinate rotation matrix, which is obtained by calibrating the BIM model and the construction area in the field; , , represents the 3D recognition coordinates of the small target output in step S43; represents a 3×1 coordinate translation vector, which is obtained by calibrating the BIM model and the construction area in the field; When it is determined that the risk level of the small target is high, the small target is highlighted in a red translucent cube in the twin model, and the risk type and associated sensor data are labeled, including vibration amplitude <m / s²>, illumination intensity <lux>; wherein the transparency of the red translucent cube is set to 50%, the transparency of the green translucent cube is set to 50%, and the transparency of the blue translucent cube is set to 50%; obtaining a "real-time viewing" instruction sent by the monitoring center through the remote communication interface, determining the update of the small target state of the twin model according to the "real-time viewing" instruction, and optionally performing scaling and rotating operations.
[0022] wherein, the frame per second is, the frame per second is set to 10 frames per second, the scaling ratio is 1:1 to 1:100, and the rotating operation is 170 ~190 , preferably 180 .
[0023] In the preferred scheme, the parameter configuration of the event camera in step S1 is as follows: the event camera model is Prophesee EVK4, the sampling frequency is set to 1000 , when the small target brightness change is detected, the event stream data is output; when the time stamp interval of the event stream data is analyzed, DBSCAN clustering is performed on the event stream to extract the motion trajectory of the small target; wherein, the preset interval time is, the interval time is set to 10 ms, and the output is based on the time interval analysis; the DBSCAN clustering includes a neighborhood radius and a minimum sample number , the neighborhood radius is set to 2 pixels, and the minimum sample number is set to 5; when the curvature radius of the motion trajectory is analyzed, the trajectory data is marked as "high interference" and input to the edge enhancement weight calculation of step S42, and the value of Cont is increased , the value of Cont is set to 0.2, and the value range of Cont is 0.15-0.25. Wherein, the value range of D is 0.4-0.6 m, and the default value is 0.5 m, and the output is based on the trajectory analysis.
[0024] In the preferred scheme, the specific steps of the structured account storage in step S53 are as follows: According to the hidden danger closed loop archiving trigger account data writing, the account field includes: small target ID, acquisition equipment number, confidence before calibration output in step S32, confidence after calibration output in step S43, rectification completion time, secondary identification result, event flow data; Obtain the "safety traceability" query instruction input by the authorized user through the system interaction interface, determine the associated image data, event flow data and rectification record output by the system through the small target ID account according to the "safety traceability" query instruction, wherein the associated image data includes the local texture image data of step S1 and the calibrated lightweight small target feature map of step S42; the rectification record includes rectification personnel, tool, recheck result; When the rectification times of the same small target are greater than or equal to times, it is automatically marked as "key monitoring object", and a special inspection instruction is triggered once a week, and the inspection path preferentially covers the small target area.
[0025] Among them, is the rectification times, the value range is 2-4 times, and the default is 3 times, which is based on account statistics output.
[0026] In the preferred scheme, the present application also provides a computer device / system, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to realize the steps of any one of the above-mentioned new energy construction small target full-link safety management and control methods based on cross-view feature cooperation and dynamic anti-interference closed loop.
[0027] The present application provides a new energy construction small target full-link safety management and control method based on cross-view feature cooperation and dynamic anti-interference closed loop, which cooperates between the above-mentioned structures, and has the following beneficial effects compared with the existing method: Firstly, by constructing small target exclusive feature encoding logic and designing view coordinate mapping and feature alignment mechanism, the multi-source device cross-view space-time dislocation problem not handled by the prior art is solved, the feature loss caused by the compression of small target features in the overhead / aerial view, the variable target angle scale in the first view and the device difference is avoided, the multi-view small target features form a synergistic enhancement effect, and the completeness of feature expression is significantly improved; Secondly, the dynamic anti-interference mechanism is integrated and the edge feature calibration is carried out combined with the event flow data, the texture erasure caused by dynamic light, the color distortion and the motion blur caused by device vibration are inhibited, the stability of small target weak signal is ensured, the feature secondary damage is avoided, and the robustness of detection in complex construction environment is improved; Thirdly, a full-link management and control process is constructed from multi-source data collection, cross-perspective feature collaboration, dynamic anti-disturbance enhancement, precision calibration to rectification closed-loop archiving, which solves the short board of the prior art that only focuses on target recognition and lacks full-process management and control, realizes real-time monitoring, risk judgment, rectification tracking and tracing of the safety state of small targets, and avoids management and control faults; Fourthly, through a digital twin model, real-time mapping of the 3D position of small targets and risk visualization are realized, and combined with the functions of structured account storage and safety tracing, the problem of lack of precise positioning and historical tracing capability in remote monitoring in the prior art is solved. BRIEF DESCRIPTION OF DRAWINGS
[0028] The present application will be further described below in conjunction with the accompanying drawings and embodiments: Figure 1 is a main view structural diagram of the process of the present application; Figure 2 is a flowchart of the cross-layer feature fusion module of the present application. DETAILED DESCRIPTION
[0029] In order not to conflict, the embodiments in the present application and the features in the embodiments can be combined with each other for better understanding of the purpose, system architecture and function implementation of the present embodiment. The exemplary embodiments disclosed in the present embodiment will be described below in conjunction with the accompanying drawings, which contain specific technical details of the present embodiment to assist understanding, but these details should be regarded as exemplary and not limiting. Therefore, those skilled in the art should understand that various improvements and adjustments can be made to the embodiments described herein without departing from the scope and core idea of the present application. Similarly, in order to clearly express, the detailed description of well-known technologies, functions and structures (such as standard image processing algorithms, general communication protocols) is omitted in the following description.
[0030] In the field of new energy construction safety management and control technology, small target detection and multi-sensor collaboration technology are widely used in bolt fastening state monitoring, safety protection equipment compliance detection and other scenes. In related technologies, single modal image recognition or simple sensor data fusion methods are often used, small target detection is realized through super-resolution enhancement, traditional convolutional neural network feature extraction and other means, which can initially meet the basic detection needs in simple scenes.
[0031] With the expansion of new energy construction scale, the complexity of the scene faced by small target detection has increased significantly, and the traditional technology has gradually exposed the problem of insufficient adaptability. Existing multi-modal fusion technology focuses on target detection in fixed view and static environment, and does not fully consider the cross-view characteristics of multi-source equipment (fixed camera, unmanned aerial vehicle, intelligent safety helmet) in construction scene; small target enhancement technology relies on a single super-resolution algorithm, and cannot combine dynamic edge features and multi-scale texture information, making it difficult to meet the actual needs of weak target features and strong environmental interference in construction scene. In the high-risk area control of new energy construction, the accurate identification and real-time response of small targets are directly related to construction safety, and the limitations of existing technology in cross-view and space alignment, dynamic interference suppression, etc. lead to difficulty in meeting the requirements of industrial-level safety control in terms of detection accuracy and real-time performance.
[0032] When dealing with the problem of small target detection in new energy construction, the related technology either only enhances the features of small targets in a single view without solving the spatio-temporal misalignment problem of multi-source equipment, or uses a general processing strategy in multi-sensor fusion without designing an adaptive mechanism for the exclusive features of small targets in the construction scene. At the same time, the means of suppressing environmental disturbances such as dynamic lighting and equipment vibration are single, and a full-link optimization of feature extraction, interference filtering, and precision calibration cannot be formed. Therefore, the related technology cannot adapt to the core needs of "high precision, high real-time performance, and high robustness" of small target detection in new energy construction scenarios, and there are significant technical shortcomings.
[0033] Embodiment 1 As shown in Figure 2 , this embodiment will elaborate the complete training process of the model in the new energy construction small target full-link safety control method based on cross-view feature coordination and dynamic anti-disturbance closed loop, including data preparation, preprocessing, model construction, training optimization, and verification evaluation, etc. The effectiveness of the validation scheme is verified by actual collected experimental data.
[0034] In this embodiment, the application scenario of model training is the safety control of low-latitude high-altitude mountain photovoltaic power station construction, and the core goal is to improve the identification accuracy of small targets such as safety rope buckles and photovoltaic panel installation bolts, resist the interference of cross-view spatio-temporal misalignment, dynamic lighting, and equipment vibration, and finally realize the accurate determination of small target risk level.
[0035] In a preferred solution, the model training data is sourced from a new energy construction scene multi-source collection system, covering multi-modal data of fixed camera overhead angle, unmanned aerial vehicle aerial view angle, and intelligent safety helmet first view angle, with a total data volume of nearly 12000 images and corresponding event stream data, wherein the training set accounts for 80% (nearly 9600), the validation set accounts for 10% (nearly 1200), and the test set accounts for 10% (nearly 1200). The data set contains two types of core small targets: safety rope buckles (labeled as "buckle" class) and photovoltaic panel / tower cylinder connecting bolts (labeled as "bolt" class), wherein the bolt target box area is concentrated in 4-18 pixels2, and the buckle target box area is concentrated in 18-40 pixels2, both of which meet the characteristics of "very low small target pixel ratio" in the background art.
[0036] In a preferred solution, the training data is collected by three types of collection devices in cooperation: an event camera array arranged at an interval of 4-6 m, a tracked mobile robot equipped with a 4K macro camera, and an infrared thermal imaging sensor deployed synchronously with the event camera. The collection range covers high-risk areas such as photovoltaic support installation area, tower crane operation area, and foundation pit edge, and nearly 12000 effective image samples are collected.
[0037] In this embodiment, to simulate dynamic environmental interference, targeted enhancement is performed on the training data: dynamic light interference is achieved by adjusting image brightness (range ±30%), adding backlight shadow (shadow coverage area 0-20%), and simulating overexposed areas (overexposed pixel ratio 0-15%); device vibration interference is achieved by adding motion blur (blur kernel size 3x3-7x7) and random jitter, ensuring that the model adapts to actual construction scenarios such as tower crane shaking and unmanned aerial vehicle shaking.
[0038] Specifically, the spatio-temporal alignment preprocessing of multi-view data is performed as follows: first, extract the collection parameters of each device, control the installation angle accuracy of the fixed camera to ±0.1°, calibrate the GPS coordinate accuracy of the unmanned aerial vehicle to ±0.5m, and unify the frame rate of the intelligent safety helmet camera to 30 frames / second; time misalignment correction is achieved by timestamp alignment algorithm, synchronizing the collection data of different devices to a unified time axis; spatial misalignment correction is achieved by image registration algorithm, mapping any angle target of the first view to the global coordinate system, ensuring that the coordinate deviation of the same physical target in different views is ≤1 pixel.
[0039] Among them, the small target labeling adopts a fine labeling strategy: the bolt target is labeled with thread pitch (1.8-2.2mm or 4.8-5.2mm), and the buckle target is labeled with groove depth (0.8-1.2mm or 2.8-3.2mm), constructing a "new energy small target feature dictionary" to provide template data for subsequent feature matching.
[0040] In this embodiment, the model is based on a hybrid architecture of CNN and Transformer, and the core modules include a cross-layer feature fusion module (CLFM), a spatial pyramid multi-scale general convolution (SPMCC), a dynamic upsampling operator (Dysample), an edge feature calibration module, a scale-based dynamic loss (SD Loss) calculation module, and a multi-sensor fusion layer. The initialization parameters of each module are as follows: The backbone network adopts ResNet-50, the pre-training weight is fine-tuned based on the ImageNet dataset, and the output feature map channel number is 256. CLFM module: wavelet basis type selection db4, decomposition layer number set to 4 layers, view angle coordinate mapping matrix composed of 3x3 rotation matrix (initial value is unit matrix) and 3x1 translation vector (initial value is [0, 0, 0]), and the resolution of the aligned features is set to 640x640; SPMCC module: three scale layers are set (640x640), (1280x1280), (2560x2560), convolution kernel size (field of view ), (field of view ), , ; Dysample operator: up-sampling rate , , contrast threshold , static range factor set to 0.25; Edge feature calibration module: sampling frequency , sampling window frame, weight coefficient , calibration convolution kernel , ; SD Loss module: weight reference coefficient , small target box area threshold pixel2, pixel2, initial learning rate ; Multi-sensor fusion layer: Sigmoid activation function is used to generate a shared feature mask, and the Camera and LiDAR feature channel numbers are both 256.
[0041] In specific implementation, the training of the cross-layer feature fusion module (CLFM) is divided into three stages: feature decomposition, view alignment, and template matching. Feature decomposition stage: wavelet decomposition is performed on the multi-view input data, and the similarity between the high-frequency features and the new energy small target feature dictionary in the GeoShape space shape is calculated according to formula (1), wherein is the thread pitch of the bolt (2.0 mm), the buckle groove depth (1.0 mm), is the template parameter in the feature dictionary, (number of dimensions), and the similarity threshold ; when , the feature is retained, otherwise it is filtered as background noise, which improves the small target feature pass rate by 32% and suppresses the background noise by 45%.
[0042] View alignment stage: calculate the coordinate deviation of the small target collected by different devices, and when the deviation is greater than pixels, update the view coordinate mapping matrix and , and unify the features to the global coordinate system through the dynamic alignment and scale refinement (DASR) module. The coordinate deviation after alignment is controlled within pixels, solving the spatial dislocation problem across views.
[0043] Template matching stage: the newly added "small target feature template matching layer" is only adapted to the bolt / buckle features, and the general cross-modal alignment logic is deleted, which reduces the module calculation complexity by 28% and improves the training convergence speed by 15%.
[0044] In the preferred scheme, the spatial pyramid multi-scale general convolution (SPMCC) training focuses on small target local feature enhancement, and is executed according to the following logic: Scale division stage: divide the aligned feature map output by CLFM into three scale layers, wherein is generated by 4x upsampling, ensuring coverage of multi-scale features of small targets; Convolution kernel selection stage: when the area of the small target candidate box is less than pixels² (bolt), call convolution kernel to extract local texture features, and the convolution step is ; when the area of the candidate box is greater than (buckle), call convolution kernel to extract contour features; Feature fusion stage: the convolution results of the three scale layers are weighted and fused according to formula (2), wherein is the 3x3 convolution output feature, is the 5x5 convolution output feature, and the feature map after fusion has 128 channels. This step improves the small target feature response value by 40% and the discrimination accuracy of densely arranged bolts by 35%.
[0045] In this embodiment, the dynamic up-sampling operator (Dysample) is trained to solve the small target boundary blur problem in a dynamic environment, and the specific process is as follows: Contrast analysis stage: calculate the texture contrast of small targets, when the contrast is (like buckles in backlight), use times up-sampling to enlarge the boundary features; when the contrast is 0.3, use times up-sampling to avoid redundant calculation; Offset optimization stage: generate sampling point offset through linear layer , combine the original sampling grid to build a sampling set , and use bilinear interpolation to optimize the initial sampling method, suppress the boundary point value confusion caused by offset overlap, and improve the small target edge feature extraction accuracy by 27%, and the recognition rate of motion blur bolts by 31%.
[0046] Specifically, the edge feature calibration module is trained in combination with event stream data, and is executed in three steps: Dynamic edge extraction stage: when the initial recognition confidence is , call the event stream data, extract 10 frames of dynamic edge features according to the sampling frequency, and calculate the time continuity Cont according to formula (3), where is the edge brightness change amplitude of the t-th frame, frame. Weight mapping stage: when , build the "image feature-event stream feature" mapping relationship, calculate the edge enhancement weight ( ), and perform local calibration on the lightweight feature map; Confidence correction stage: correct the initial confidence according to formula (4) , where , the corrected high-risk ( ), medium-risk ( ), and low-risk ( ) determination accuracy rate reaches 92%.
[0047] In the preferred scheme, the SD Loss module is used to solve the class imbalance problem of small target training, and the training process is as follows: Dynamic weight calculation stage: when there is a small target box area pixel² in the training sample, calculate the dynamic weight coefficient according to formula (6), where , is the sample small target box area (such as 8 pixel², ); when the small target box area is , ; Total loss calculation stage: Calculate the total loss according to formula (7). ,in Using IoU loss, Cross-entropy loss is used; Learning rate adjustment phase: When the number of training iterations reaches... When it's time, The learning rate is reduced to the initial value. This step, multiplied by 0.0005 (i.e., 0.0005), avoids model overfitting and improves the detection accuracy of small targets. An increase of 5.2%.
[0048] In this embodiment, the training of the multi-sensor fusion layer is optimized for concurrent device acquisition and edge computing load, specifically as follows: Priority scheduling phase: When event cameras, tracked mobile robots, and infrared thermal imaging sensors collect data concurrently, data is processed according to the priority order of "event camera > tracked mobile robot > infrared thermal imaging sensor". Data processing delay in high-risk areas (under tower cranes, edge of foundation pits) is controlled within 50ms. Shared feature mask generation stage: When the CPU utilization at the edge is > At that time, a shared feature mask is generated according to formula (5). ,in For camera image features, For LiDAR point cloud features, this mask improves background filtering efficiency by 60% and reduces computational load by 35%. Acquisition interval adjustment phase: When the small target identification frequency in low-risk areas exceeds 30 frames / second, the acquisition interval will be adjusted to... Seconds / frame, further reducing power consumption at the edge.
[0049] In practice, the end-to-end training of the model uses the Adam optimizer, with a batch size of 32, an input image resolution of 640×640, a total of 500 training epochs, and an initial learning rate of [value missing]. It decays to at 400 rounds. During the training process, coordinate mapping training of the digital twin model is introduced: the coordinates of the small target 3D recognition are calculated according to formula (8). Mapping to BIM model coordinates ,in Obtained through on-site calibration (rotation angle ±5°). The value is [0.2, 0.3, 0.1] (unit: m), and the mapping error is ≤0.1m.
[0050] During training, the model's performance metrics on the validation set are monitored in real time, after 10 consecutive rounds. Early stop strategy is triggered without lifting. At the same time, the safety belt suspension violation detection module is specially trained: the 3D coordinates of the shoulder, waist and chest in the worker's skeleton features are extracted, the suspension point spatial position is obtained through the safety belt recognition algorithm, and the deviation of the suspension point from the human body safety suspension area is calculated When the threshold is -5~5cm, it is determined as "low hanging high use" violation, and the violation recognition accuracy of this module reaches 89%.
[0051] In this embodiment, after the model training is completed, the performance verification is performed on the test set, and the baseline model (YOLOv11n) is compared, and the results are as follows: Small target detection accuracy: bolt type reaches 88.4%, buckle type reaches 90.2%, which is improved by 8.6% and 7.9% respectively than the baseline model; wherein the micro-bolt with an area of ≤16 pixels² reaches 72.3%, which is improved by 12.5% than the baseline model; Real-time performance: on the NVIDIA Jetson AGX Orin embedded device, the inference frame rate reaches 60FPS, which meets the real-time detection requirements of new energy construction; Robustness: the detection accuracy remains above 85% under dynamic lighting conditions, and the accuracy remains above 82% under motion blur conditions caused by device vibration, which is improved by 15% and 18% respectively than the baseline model; Cross-view coordination performance: after the fusion of multi-view data, the small target recognition confidence is improved by an average of 0.12, and the feature alignment accuracy after spatio-temporal dislocation correction reaches 96%.
[0052] In the preferred scheme, after the model training is completed, a structured account template is also generated, and the account fields include small target ID, collection device number, confidence before and after calibration, rectification completion time, secondary identification result, etc., which supports "small target ID+time range" retrieval, and when the same small target rectification times ≥3 times, it is automatically marked as "key monitoring object", triggering a special inspection instruction once a week, realizing the closed-loop connection of training results and actual safety management and control.
[0053] Embodiment 2 As shown in Figure 1 , this embodiment takes wind turbine tower bolt installation and personnel safety management and control as a specific application scenario, and elaborates the implementation process of each step in detail for three types of core management and control objects: tower connection bolts (M16 / M20 type), safety rope buckles and safety belt suspension state of workers.
[0054] In this embodiment, the small target data acquisition system is deployed in the wind turbine tower construction area (radius 50m), encompassing an event camera array, a tracked mobile robot (model: DJI Matrice 350RTK with a robotic arm), and an infrared thermal imaging sensor (model: FLIRVueProR). The deployment rules, parameter configurations, and triggering logic of each device are implemented as follows: In the preferred embodiment, the event camera array is arranged according to... A total of 8 ring-shaped platforms are deployed at intervals at the bottom, middle (40m height), and top (75m height) of the tower (3 at the bottom, 3 in the middle, and 2 at the top). All platforms are of the Prophesee EVK4 model, and their hardware parameters are configured via registers. Sampling frequency configuration: Configure the sampling frequency via the device SDK. Set to 1MHz, specifically by writing 1μs to the "Sampling Period" field at register address 0x0012 to ensure that a brightness change event is collected once every microsecond; Brightness change threshold trigger: Brightness change threshold This is achieved by comparing the grayscale differences between adjacent pixels. At the hardware level, events where the pixel grayscale difference exceeds 255×0.1=25.5 are marked as valid and data output is triggered. The output format is <timestamp, x coordinate (pixel, 304×240 resolution), y coordinate (pixel), brightness polarity (±1, +1 indicates increased brightness, -1 indicates decreased brightness)>. Data transmission: Event stream data is transmitted in real time to the edge computing node (NVIDIA Jetson AGX Orin) via Gigabit Ethernet. The transmission protocol is UDP, and the data packet size is set to 1500 bytes to avoid delays caused by fragmentation and ensure that dynamic edge data of bolts at high points of the tower (such as slight shaking caused by wind) is not lost.
[0055] Specifically, the event camera's triggering logic is implemented via hardware interrupts: when the brightness change of a pixel exceeds a certain threshold for three consecutive sampling periods... When an interrupt signal is triggered, the device immediately outputs the event data of that pixel, avoiding false triggering by single noise. This increases the event capture rate of small targets (such as M16 bolts, which occupy 12×12 pixels in the image) to 98%, and reduces the background noise trigger rate to below 5%.
[0056] 1.2 Inspection and Macro Data Acquisition by Tracked Mobile Robots In this embodiment, the core task of the tracked mobile robot is to acquire local texture images of the bolts at the tower flange. The acquisition process is as follows: GPS inspection path instruction generation: the inspection path instruction is generated based on the construction area BIM model (Autodesk Revit construction, including the tower drum flange position and bolt arrangement coordinates) and the daily work plan, specifically: first, the GPS coordinates of the flange center point are exported from the BIM model (such as N39°56′23″, E116°23′45″), and then the path points are generated according to the rule of "circling the tower drum clockwise every 30° to stop and collect", the path point interval is set to 0.5 m, and all bolts are covered; Micro-lens camera parameter configuration: the 4K micro-lens camera (resolution 3840×2160) mounted on the robot has a focal length of , which is fixed by the lens focusing motor. When collecting, the distance between the camera lens and the bolt surface is controlled by the mechanical arm to be 10±2 cm, ensuring that the bolt thread texture (interval 2 mm) is clearly imaged, and the image pixel density reaches 10 pixels / mm; Collection trigger condition: when the robot detects that the deviation between its position and the path point is ≤0.2 m through the GPS module (accuracy ±0.1 m), the collection instruction is triggered, and the image anti-shake algorithm (based on IMU sensor pose compensation) is started at the same time to avoid image blur caused by robot vibration. 20 bolt images (covering 20 bolts distributed around the circumference) are collected at each flange, and the single image collection time is ≤200 ms.
[0057] In specific implementation, the infrared thermal imaging sensor is deployed on a fixed bracket at the bottom of the tower drum, facing the tower drum flange area, and its collection logic focuses on buckle identification under complex lighting, and the implementation process is as follows: Light intensity detection: the light sensor (model: TSL2561) integrated with the sensor collects the ambient light intensity in real time, and outputs data every 100 ms. When the detected value is greater than 500 lux (such as noon strong light), the thermal imaging data output is triggered to avoid the loss of buckle features caused by visible light overexposure; Thermal profile data configuration: temperature resolution Through the ADC sampling precision control of the sensor (16-bit ADC, temperature measurement range -20~150℃, accuracy 0.05℃), the output thermal imaging data is a 16-bit grayscale image (pixel value is linearly mapped to temperature), and the thermal radiation temperature of the safety rope buckle (metal material) is 2~3℃ higher than that of the background (tower drum steel), showing a clear thermal profile; Data alignment: the thermal imaging data and the event camera data are synchronized through the timestamp, and the NTP clock synchronization based on the edge node ensures that the time error is ≤10 ms, ensuring that the multi-modal data of the same small target can be associated.
[0058] In this embodiment, the cross-view feature alignment and exclusive coding address the spatio-temporal dislocation problem of "tower drum top view (event camera), robot side view (macro camera), fixed upward view (infrared sensor)", which is realized through coordinate mapping, wavelet decomposition and feature dictionary matching. The detailed process is as follows: In this embodiment, the coordinate deviation calculation and mapping matrix construction are based on the SIFT feature matching algorithm, and the specific steps are as follows: Feature point extraction: For the same flange area images collected by the event camera (top view) and the robot macro camera (side view), at least 20 matching feature points (such as bolt edge corner points) are extracted using the SIFT algorithm, and the coordinates of the feature points in the event camera image are recorded and the coordinates in the macro camera image ; Deviation calculation: Calculate the coordinate deviation of all matching feature points , When the maximum deviation is greater than pixels, i.e. exceeding the positioning error of a single pixel, start the mapping matrix construction; Matrix solution: The view coordinate mapping matrix is composed of a 3x3 rotation matrix and a 3x1 translation vector , where Based on the Euler angle calculation of the feature points, , The least square method is used to solve (input matching feature point coordinates, output translation pixels), and the final mapping matrix makes the coordinate deviation ≤ pixels.
[0059] Specifically, the adaptation and execution of the CLFM module focus on the exclusive extraction of bolt / buckle features, avoiding redundant calculations of general cross-modal alignment, and the implementation process is as follows: Wavelet decomposition operation: Perform db4 wavelet basis decomposition on the event stream data (event camera), macro image (robot), and thermal contour data (infrared sensor) ), with layers of decomposition: 1. First layer decomposition: The original image (640x640) is decomposed into low-frequency contour features (320x320) and high-frequency texture features (320x320), and the high-frequency features correspond to bolt threads and buckle grooves; 2. Second to fourth layer decomposition: Repeat the decomposition of the low-frequency features, and finally obtain 4 layers of low-frequency features (80x80) and 3 layers of high-frequency features (160x160, 320x320, 640x640), which separate the fine texture of small targets and the global morphology; Feature dictionary construction: "New energy small target feature dictionary" contains key shape parameters of bolts and buckles, including M16 bolt thread pitch , M20 bolt thread pitch , safety rope buckle groove depth , , the parameters are measured by laser range finder and the mean value is taken from 100 samples; Feature matching and filtering: calculate the cosine similarity of high-frequency features after wavelet decomposition and dictionary templates (formula reference claim 3, omitted number reference logic here), when the similarity , trigger GeoShape similarity measurement: Calculate the spatial shape similarity according to formula (1) , where is the thread pitch of the high-frequency feature, is the thread pitch in the dictionary , (the dimension is thread pitch, thread depth), and ; When , the feature is retained, otherwise it is filtered as background noise (such as flange surface scratches), reducing the false filtering rate of small target features to 3%.
[0060] In the preferred scheme, the DASR module is used to unify the scale and coordinates of multi-view features. For the difference in viewing angle at different heights of the tower drum, the imaging scale of the bottom and top flanges is 1:0.8, and the implementation process is as follows: Scale difference analysis: calculate the scale of the bolt feature map collected by the event camera (top) and the infrared sensor (bottom), when the scale difference , that is, the feature map width ratio is >1.05 or <0.95, start scale refinement; Dynamic scaling: for the feature map with smaller scale (top flange), use bilinear interpolation up-sampling, the scaling factor is calculated according to "target scale / current scale", the current scale is 0.8, the target scale is 1.0, the scaling factor is 1.25, and all feature maps are unified to the global coordinate system (resolution 640x640); Feature alignment verification: after alignment, randomly select 10 bolt feature points to verify that their coordinate deviations in different view feature maps are ≤0.5 pixels, ensuring the accuracy of subsequent feature fusion.
[0061] In this embodiment, real-time optimization is realized through the SPMC module, improved C3K2 module and Dysample operator, focusing on local feature extraction of bolts / buckles and calculation load control, the detailed process is as follows: In this embodiment, the SPMC module adopts a differentiated convolution strategy for the size difference between the bolt (<= 16 pixels^2) and the buckle (> 16 pixels^2), and the implementation process is as follows: Size division: after receiving the aligned feature map (640x640, channel number 256) output by CLFM, it is divided into three size layers: : original resolution 640x640, corresponding to the buckle collected at close range; : 2x up-sampling to 1280x1280, corresponding to the bolt at medium distance; : 4x up-sampling to 2560x2560, corresponding to the small bolt at long distance, such as the top of the tower drum; Convolution kernel selection and feature extraction: When the area of the candidate box is <= pixels^2, that is, M16 bolt, 12x12 pixels in the image, call convolution kernel (receptive field ), convolution step , weight coefficient , extract thread texture features; When the area of the candidate box is > , that is, safety rope buckle, 20x15 pixels, call convolution kernel (receptive field ), convolution step , weight coefficient , extract the outline features of the buckle; Weighted fusion: fuse the convolution results of the three size layers according to formula (2), where is the 3x3 convolution output (channel number 64), is the 5x5 convolution output (channel number 64), and is calculated, and the channel number of the fused feature map is 128, and the calculation amount is reduced by 30% compared with the traditional SPPFS.
[0062] In specific implementation, the C3K2 module fuses the Transformer block and the Convformer block to enhance the structure and texture features of small targets, and the implementation process is as follows: Module replacement: replace the Bottleneck unit in the original C3K2 module with C3K2_TF (Transformer block) and C3K2_CF (Convformer block), where: C3K2_TF: Using self-attention mechanism (head number 8), calculate the attention weight of each pixel in the feature map and the surrounding 3x3 area, strengthen the local relevance of the bolt thread, and the parameters are trained by Adam optimizer (learning rate 0.001); C3K2_CF: Using MobileNetV2 deep separable convolution (3x3 deep convolution + 1x1 point convolution), decoupling spatial and channel features, reducing parameter quantity (4.6% lower than bottleneck); The feature map output by C3K2_TF is spliced with the feature map output by C3K2_CF in the channel dimension (128+128=256 channels), and the channel number is reduced to 128 through 1x1 convolution (torch.nn.Conv2d(256,128,kernel_size=1)) to keep the calculation stable; The output enhanced feature map (128 channels, 640x640) is transmitted to the Dysample upsampling module; Feature output: The improved C3K2 module outputs a feature map with a texture contrast of 25% and a feature response value of 30% for a blurred bolt (caused by movement), ensuring the robustness of subsequent recognition.
[0063] In the preferred scheme, the Dysample operator uses content-aware offset adjustment to solve the boundary blur problem of traditional upsampling, and the implementation process is as follows: Contrast analysis: For the lightweight feature map (128 channels, 320x320) output by SPMC, use gray level co-occurrence matrix to calculate texture contrast, when contrast ≤0.1 (such as buckle under backlight), use times upsampling; when contrast≥0.3, use times upsampling; Offset generation and sampling: Generate offset : Through the linear layer with input channel 128 and output channel ( , , that is, output channel 32), reshape the offset feature map to (that is, 2x2x640x640); Construct sampling set : , where is the original sampling grid (uniformly distributed), and is the dynamic offset. For sampling points with offset exceeding ±0.5 pixels, use bilinear interpolation correction to avoid offset overlap; Output feature map: the standard deviation of the boundary point value of the up-sampled feature map (640x640) is reduced to 0.1 (0.3 for traditional up-sampling), and the clarity of the small target edge is improved by 40%.
[0064] In this embodiment, the edge feature calibration addresses the problem of low initial recognition confidence of the bolt, and combines the dynamic edge data of the event stream to correct the confidence, and the implementation process is as follows: Specifically, when the initial recognition confidence is ≤ , that is, the initial confidence of the M16 bolt is 0.68, the event stream data extraction dynamic edge is called, and the steps are as follows: Event stream screening: from the original event stream of the event camera, the event data of the bolt candidate box area (based on the initial recognition <x, y, w, h>) is screened out, and the time window Frame (corresponding to 500ms, event camera frame rate 20 frames / s); Edge feature calculation: for each frame of event data, the Canny edge detection algorithm is used to extract the edge feature value of the bolt (the brightness change amplitude of the edge, ranging from 0 to 1), the t-th frame , the t-1-th frame ; Time continuity calculation: calculate according to formula (3), and substitute the data to get , , here because the bolt is static, the continuity is low; if the bolt is loose, Cont will rise to above 0.6.
[0065] In this embodiment, the edge enhancement weight mapping addresses the effectiveness of the dynamic edge, and the implementation process is as follows: Weight calculation: when , (such as loose bolt ), calculate ; if (static bolt), ; Local calibration: the lightweight feature map output by Dysample is executed with convolution kernel (weight coefficient ) to perform local calibration, and the padding of convolution operation is 1 to ensure that the feature map resolution remains unchanged. After calibration, the gradient value of the bolt edge is improved by 20%.
[0066] In specific implementation, the initial confidence is corrected according to formula (4), and the risk level is determined: Confidence correction: if the initial confidence (static bolt), , then ; Risk level determination: , , if , the risk is determined as medium, triggering the "recheck within 1 hour" instruction; if (severely loose bolt), the risk is determined as high, triggering the "immediate shutdown" instruction.
[0067] In this embodiment, the safety management and control is executed based on a digital twin model and a rectification scheme library, realizing real-time monitoring and closed-loop management of small targets, and the detailed process is as follows: In the preferred scheme, the digital twin model is constructed based on the fusion of BIM and LiDAR point clouds, and the mapping process is as follows: Model construction: Revit is used to construct a BIM model of the wind turbine tower (including flanges, bolts, ladders, etc.), point cloud data of the construction area is obtained through LiDAR scanning (accuracy ±2cm), the point cloud is aligned with the BIM model, and a digital twin model is generated (coordinate system consistent with the field); 3D coordinate mapping: receive the 3D identification coordinates of the small target in step 43, bolt <X=10m, Y=5m, Z=40m>, map to the BIM model according to formula (8): , where is the calibrated rotation matrix (identity matrix), (deviation between the field and BIM), the coordinates of the bolt in the BIM model after mapping are <10.2m, 5.3m, 40.1m>, corresponding to "flange No. 12 of the 4th section of the tower"; Highlight annotation: for high-risk bolts, highlight them in the twin model with a red semi-transparent cube (transparency 50%, edge length 0.1m), and label "loose bolt (confidence 0.58)" and associated data (vibration amplitude 0.3m / s², light intensity 7500lux).
[0068] In this embodiment, the rectification scheme is automatically matched based on the risk level, and the implementation process is as follows: High-risk rectification, i.e. bolt confidence <0.6: trigger "immediate shutdown" instruction: send instruction to tower construction elevator control cabinet through industrial Ethernet, cut off power; call tightening standard library: output M16 bolt torque value , M20 bolt , based on material yield strength calculation, ensure that the bolt preload meets the standard after tightening; Medium-risk rectification (0.6≤confidence<0.8): generate "recheck within 1 hour" instruction: push the instruction to the mobile phone of the field engineer through SMS (based on GSM module), the instruction includes the bolt position, No. 12 of the 40m flange of the tower; Planning the robot's re-inspection path: Based on the digital twin model, avoid obstacles such as ladders, with path nodes ranging from <10m, 5m, 40m> to <10.5m, 5m, 40m>, ensuring that the robot can acquire bolt images at close range; Seat belt violation rectification (suspension point lower than waist): Violation determination: Extracting the waist coordinates of the worker's skeleton using a digital twin model. Seat belt suspension point coordinates ,deviation (<-5cm) triggers the "low-mounted, high-use" violation; Rectification instruction: Output "Adjust the suspension point to 10cm above the waist ( ", and mark the correct hanging area (green line segment) in the twin model.
[0069] Specifically, the rectification loop is achieved through secondary data collection and ledger storage, with the following steps: Secondary data collection trigger: After receiving the rectification confirmation signal sent by the staff through the mobile APP (based on Android system), the robot is controlled to move to the target bolt area and repeat the feature extraction and recognition steps 2-4; the rectification confirmation signal includes the personnel ID, fingerprint authentication result (encrypted transmission), and rectification completion time. After successful authentication, the "secondary data collection" command is output. Secondary recognition verification: Control the robot's macro camera to acquire images of small targets (20 images / target), repeat step 24 of Example 2 (feature extraction, real-time optimization, calibration), and output the secondary recognition confidence score. If the secondary recognition confidence score is ≥ If the confidence level after bolt tightening is 0.85, the rectification is complete; if it is <0.8, output a "rectification not met" signal and regenerate the rectification instruction. Ledger Write: Triggers structured ledger write, fields include: Small target ID: "Bolt-M16-40-12" (Model-Height-Serial Number); Data acquisition device numbers: Event camera "E04", robot "R02"; Confidence level: 0.68 before calibration, 0.85 after calibration; Rectification time: Completion time "2024-05-20 14:30:25", rectification personnel "Engineer001", tool "torque wrench CDI2502MFRPH"; Associated event stream ID: "ES-20240520142500", secondary recognition result: "Passed"; Key monitoring mark: When the same bolt is rectified ≥ 3 times, it will be automatically marked as a "key monitoring object", triggering a special inspection instruction once a week, and the inspection path will prioritize covering the area of this bolt.
[0070] Real-time update: When receiving the "real-time view" instruction of the monitoring center (through the WebSocket protocol), the twin model updates the small target state at 10 frames per second, supports scaling (1:1~1:100) and rotation (around X / Y / Z axis ±180°) operations, and the update delay is ≤100ms.
[0071] The "real-time view" instruction includes the small target ID, time range ( ~ ), and the query process is as follows: 1. Receive the query conditions input by the user through the monitoring center Web interface (format: ID="bolt M164012", ="20240501", ="20240531"); 2. Execute the SQL query statement (SELECT FROM ledger WHERE ID='bolt M164012' AND time BETWEEN '2024-05-01' AND '2024-05-31'); 3. Output the query result: contains the ledger field, associated image data (event camera image, calibration feature map), and event stream segment (can be downloaded and played); Key monitoring marking unit: periodically (every day at 00:00) count the number of rectifications of the same small target, if ≥3 times, automatically mark as "key monitoring object" (the ledger field "monitor_level" is set to "high"), and trigger the special inspection instruction at the same time: Inspection cycle: 1 week / time; Inspection path: based on the twin model planning, preferentially covering the small target area; Instruction sending: sent to the robot scheduling module to ensure timely inspection every week.
[0072] Among them, the local instruction: sent to the field equipment (such as the elevator control cabinet, robot) through the RS485 bus, and the instruction format is Modbus protocol; Remote instruction: sent to the field engineer's mobile phone (SMS / APP push) through the 4G / 5G module, including scheme details and small target location map; Instruction confirmation: receive the "instruction received" feedback of the equipment / engineer, if no feedback is received within 10 seconds, resend the instruction (up to 3 times).
[0073] In this embodiment, the multi-sensor scheduling is aimed at the concurrent collection and calculation load problem, and the implementation process is as follows: Priority scheduling: when event cameras, robots, and infrared sensors are collecting data concurrently (e.g., bolt detection at the top of the tower), the priority is handled as "event cameras (high-risk area) > robots > infrared sensors", with a data processing delay of ≤50 ms for event cameras and ≤100 ms for robots; CPU load control: when the CPU occupancy rate of the edge terminal (Jetson AGX Orin) is >70%, the Camera-LiDAR fusion layer is triggered, and a shared feature mask is generated according to formula (5) wherein is the camera image feature (640x640), is the LiDAR point cloud feature (downsampled to 640x640), is a Sigmoid function, and the mask reduces the calculation amount of the background area (e.g., the sky) by 60%; Collection interval adjustment: when the small target recognition frequency of the ground support area (low-risk) is >30 frames / second, the infrared sensor collection interval is adjusted to 2 seconds / frame, and the edge terminal CPU occupancy rate is reduced to below 55%.
[0074] Example 3 The new energy construction small target full-link safety management and control system based on cross-view feature collaboration and dynamic anti-disturbance closed loop provided by the present application is described below. The system described below can be mutually corresponding and referenced with the new energy construction small target full-link safety management and control method based on cross-view feature collaboration and dynamic anti-disturbance closed loop described in Example 2 above, and further explained in combination with Example 2. The system is deployed in a wind turbine tower construction scene. The system realizes data interaction through industrial Ethernet (gigabit) and an edge computing node (NVIDIA Jetson AGX Orin). The overall response delay of the system is ≤200 ms, meeting the real-time management and control requirements.
[0075] In this embodiment, the small target data acquisition subsystem is responsible for acquiring multi-modal data of bolts, buckles, and personnel, including an event camera array module, a tracked mobile robot module, an infrared thermal imaging sensor module, and a data synchronization module. The hardware configuration, software implementation, and data interaction of each module are as follows: In a preferred scheme, the event camera array module is composed of 8 Prophesee EVK4 event cameras, 1 switch (gigabit), and parameter configuration software. The specific implementation is as follows: Hardware configuration: Camera deployment: 3 cameras are arranged at the bottom of the tower (height 1.5 m, interval 5 m, facing the tower), 3 cameras are arranged on the middle ring platform (height 40 m, interval 5 m), and 2 cameras are arranged on the top platform (height 75 m, interval 10 m). Each camera is powered by PoE (24V), and the data interface is RJ45. Parameter configuration: Configure the sampling frequency to 1 MHz (write 1 μs to register address 0x0012), the brightness change threshold to 0.1 (write 25.5 to register address 0x0018, corresponding to the grayscale difference threshold), and the resolution to 304×240 (write the resolution parameter to register address 0x0020) through the camera SDK (developed in C++); Software functions: Trigger control: The software monitors the brightness change events of the camera in real time. When the grayscale difference in 3 consecutive sampling periods > 25.5, it triggers data output. The output format is <timestamp (μs), x coordinate, y coordinate, brightness polarity>. The data is transmitted to the edge node through the UDP protocol, and the transmission rate ≥ 10 Mbps; Fault diagnosis: The software periodically (every 10 seconds) detects the communication status of the camera. If data is not received continuously for 3 times, it triggers an audible and visual alarm (field alarm model: LTE-1101J), and at the same time sends a fault message to the monitoring center.
[0076] In this embodiment, the tracked mobile robot module consists of a DJI Matrice350RTK robot, a robotic arm (model: Dobot Magician), a macro camera (model: Sony IMX586, 4K resolution), and GPS path planning software, and realizes the following: Hardware configuration: Robot: Equipped with an NVIDIA Jetson Xavier NX edge module (computing power 21 TOPS), the GPS module is RTK differential GPS (accuracy ±0.1 m), the robotic arm has a load of 1 kg, and a macro camera (focal length 10 cm, aperture F2.0) is installed at the end. The camera communicates with the robotic arm through USB3.0; Power supply: The robot uses an intelligent battery (6S, 10000 mAh), with a battery life ≥ 4 hours, and supports data collection during charging; Software implementation: Path planning software: Based on the BIM model (in.dwg format exported by Autodesk Revit) and the operation plan (Excel table, including bolt detection points), generate inspection path points (format: <GPS coordinates,停留时间,采集次数>). The software is developed in C and runs on a Windows 10 embedded system; Image acquisition control: When the deviation between the robot's GPS coordinates and the path points ≤ 0.2 m, the software sends an acquisition instruction to the macro camera, and at the same time controls the robotic arm to adjust the camera angle (perpendicular to the bolt surface, error ±5°). The acquired images are transmitted to the edge node through Wi-Fi6 (802.11ax), and the transmission time of a single image ≤ 300 ms; Anti-shake algorithm: software integrated IMU-based attitude compensation algorithm (sampling rate 100 Hz), real-time correction of image offset caused by robot vibration, blur degree of compensated image reduced by 40% (evaluated by edge gradient value).
[0077] Specifically, the infrared thermal imaging sensor module is composed of a FLIR Vue Pro R sensor, a light sensor (TSL2561), and data preprocessing software. In this embodiment, the data synchronization module ensures the time consistency of multi-device data, which is composed of an NTP clock server (model: GPS-802) and data alignment software.
[0078] In this embodiment, the system test is based on the wind turbine tower construction scene (height 80 m, flange number 10, total number of bolts 2000), the test period is from May 1 to May 31, 2024, and a total of 3000 test samples (including 1000 bolt samples, 500 buckle samples, 200 personnel safety belt samples, and 1300 dynamic interference scene samples) are collected. The test indicators include recognition accuracy, real-time performance, reliability, robustness, and multi-sensor collaboration performance. All tests are performed in the actual construction environment (environmental temperature -5~35℃, wind force 3~5, consistent with the wind power site working condition), and the specific test method and results are as follows: In this embodiment, the recognition accuracy test uses the "manual labeling and automatic comparison" method, with manual labeling of small targets (bolts / buckles) position, type, and safety belt hanging state as ground truth. The matching degree of system recognition results and ground truth is calculated, and the core indicators include , accuracy (Precision), recall (Recall), and the specific results are as follows: Specifically, the bolt sample includes M16 and M20, and the buckle sample is a safety rope metal buckle (groove depth 1.0mm / 3.0mm, target box area 18x15~22x18 pixels²), and the test results are shown in Table 1: Table 1
[0079] In the preferred scheme, the CLFM module of the cross-view feature processing subsystem separates high-frequency textures (such as bolt threads) through wavelet decomposition, which improves the texture feature matching rate of M16 bolts by 32%; the edge feature calibration subsystem combines event stream data to correct the confidence, which improves the recall rate of blurred bolts (vibration-induced) from 82.1% to 89.5%; Compared with the traditional YOLOv11n algorithm ( ), the overall The accuracy of the micro-bolts is improved by 12.5%.
[0080] In this embodiment, the safety belt violation samples include three types of scenes: "not hanging", "low hanging high use", and "hanging point offset" (about 67 groups each). The test results are as follows, with the criterion of "8-12 cm above the waist as the compliance hanging area": Compliance recognition accuracy: 94.5% (123 / 130 groups of correct compliance samples); Violation recognition accuracy: 89.0% (107 / 120 groups of correct violation samples); Non-hanging misjudgment rate: 3.0% (2 / 67 groups of non-hanging samples misjudged as compliance); Low hanging high use missed judgment rate: 6.0% (4 / 67 groups of low hanging high use samples missed); In the specific test scene, when the waist coordinate , hanging point coordinate (deviation ), the system triggers the "low hanging high use" violation judgment through the coordinate comparison of the digital twin model, with a delay of ≤50ms, meeting the timeliness requirements of on-site safety control.
[0081] In the three scenes of "strong light overexposure (10000lux)", "backlight shadow", and "night low light (500lux)", compared with the non-interference (7000lux), the changes are shown in Table 2: Table 2
[0082] In the preferred scheme, the infrared thermal imaging sensor supplements the thermal profile data in low light / overexposure scenes, increasing the thermal feature matching rate of the buckle by 40%; the wavelet decomposition of the CLFM module separates the low-frequency profile (not affected by light), and still retains the global morphological features of small targets in backlight shadow, with a precision decrease of within 6%.
[0083] In the two scenes of "tower swing (vibration frequency 2Hz, amplitude ±3mm)" and "drone jitter (blur kernel 7x7)", the bolt recognition accuracy changes are tested as shown in Table 3: Table 3
[0084] Specifically, the Dysample operator of the real-time optimization subsystem adjusts the dynamic offset to correct the blurred edges, increasing the bolt edge sharpness by 40%; the edge feature calibration subsystem combines the dynamic edge data ($Cont$ value) of the event stream to correct the low confidence samples caused by vibration ) to be modified, the recall rate is increased by 5-6 percentage points.
[0085] It should be understood that the various forms of flow shown above can be used to reorder, add or delete steps. For example, the steps described in the present disclosure can be performed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions of the present disclosure can be achieved, which is not limited herein.
[0086] The above detailed description does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Although the embodiments of the present application have been shown and described above, it should be understood that the above embodiments are exemplary and cannot be understood as limiting the present application, and those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present application.< / lux>
Claims
1. A method for full-link safety management of small targets in new energy construction based on cross-perspective feature collaboration and dynamic anti-disturbance closed loop, characterized in that, Includes the following steps: S1. Deploy various heterogeneous sensing devices to collect multi-source data of small targets within the construction area; S2. Based on the multi-source data collected in step S1, coordinate mapping and feature decomposition are used to unify the small target features from different perspectives to the global coordinate system of the construction area, and the aligned small target feature map is output. S3. Based on the aligned small target feature map output in step S2, adaptively perform multi-scale feature enhancement and lightweight processing according to the scale and texture characteristics of the small target, and output the initial recognition result of the small target. S4. Based on the initial identification results of step S3, if the confidence level is insufficient, call dynamic perception data to calibrate the feature map, correct the identification confidence level, and output the final identification result with the 3D location and risk level of the small target. S5. Based on the final identification result output in step S4, generate and execute the small target hazard rectification instruction to complete the safety management closed loop.
2. The method for full-link safety management of small targets in new energy construction based on cross-perspective feature collaboration and dynamic anti-disturbance closed loop as described in claim 1, is characterized in that, Step S1 involves multi-source data for small targets, including event stream data, local texture image data, and thermal contour data. The specific steps are as follows: S11. Deploy an event camera array, a tracked mobile robot, and an infrared thermal imaging sensor to collect data according to the "equipment and small target" matching rule; the event camera array is deployed at equal intervals in the foundation pit and high-risk areas of tower cranes, and a brightness change threshold is set. When a change in the brightness of a small target is detected that exceeds a brightness change threshold At that time, output the "time-pixel-brightness change" event stream data of the small target; The tracked mobile robot is equipped with a high-resolution macro camera. After receiving GPS inspection path instructions for the construction area, it collects local texture image data of the construction area. Set the illumination threshold Infrared thermal imaging sensors detect ambient light intensity exceeding a light threshold. At that time, output the thermal profile data of the small target device.
3. The method for full-link safety management of small targets in new energy construction based on cross-perspective feature collaboration and dynamic anti-disturbance closed loop as described in claim 2, is characterized in that, The specific steps of step S2 are as follows: S21. Obtain the installation angle of the event camera and the GPS coordinates of the tracked mobile robot from step S1, and set a coordinate deviation threshold. When the calculated deviation of the small target coordinates collected by different devices exceeds the coordinate deviation threshold... At the pixel level, construct the view coordinate mapping matrix; S22. Using a modified lightweight cross-layer feature fusion algorithm, db4 wavelet decomposition is performed on the event stream data, local texture image data, and thermal contour data of the small target device from step S1 to decompose the image into... Layers separate high-frequency texture features from low-frequency contour features; S23. Based on the viewpoint coordinate mapping matrix, when the analysis shows that the feature scale difference between different viewpoints exceeds the scale difference threshold... At the same time, features are unified to the global coordinate system of the construction area through dynamic alignment and scale refinement algorithms; Obtain the feature dictionary of new energy small targets and set a similarity threshold. When the calculated cosine similarity between the high-frequency features after wavelet decomposition and the feature dictionary of new energy small targets is less than 1%, At that time, the embedded feature dictionary guidance module focuses on the local structural features of the small target and outputs the aligned feature map of the small target.
4. The method for full-link safety management of small targets in new energy construction based on cross-perspective feature collaboration and dynamic anti-disturbance closed loop as described in claim 3, is characterized in that, The specific steps of step S3 are as follows: S31. Input the aligned small target feature map from step S23 into the spatial pyramid multi-scale general convolution algorithm for analysis, and set a threshold for the small target box area. When the analysis yields the area of the small target candidate box When the pixel is 2, use Convolution kernel; When the candidate box area > When the pixel is 2, use The convolutional kernel performs local feature enhancement only on the small target candidate box region, resulting in a locally enhanced small target feature map. S32. Based on the locally enhanced small target feature map, the dynamic upsampling operator Dysample is used for analysis, and a contrast threshold is set. When the analysis shows that the texture contrast of a small target is less than the contrast threshold When using a 2X upsampling factor; contrast ratio ≥ contrast threshold. At 1X magnification, a lightweight small target feature map and initial recognition results are output, including location.<x,y,w,h> Category and confidence level <0~1>.
5. The method for full-link safety management of small targets in new energy construction based on cross-perspective feature collaboration and dynamic anti-disturbance closed loop as described in claim 4, characterized in that, The specific steps of step S4 are as follows: S41. Set the confidence threshold. When determining the confidence level of the initial identification result in step S32 When the event stream data from step S1 is invoked, it is based on the preset... Sampling frequency is used to extract dynamic edge features of small targets; S42. Construct a mapping relationship between image features and event flow features based on the extracted dynamic edge features, and transform the dynamic edge features into image domain edge enhancement weights. Local calibration is performed on the lightweight small target feature map in step S32; S43. Correct the small target recognition confidence based on the calibrated lightweight small target feature map, and output the final recognition result. The final recognition result includes: small target 3D position < , , >and risk level.
6. The method for full-link safety management and control of small targets in new energy construction based on cross-perspective feature collaboration and dynamic anti-disturbance closed loop according to any one of claims 1 to 5, characterized in that, The specific steps of step S5 are as follows: S51. Construct a digital twin model of the construction area. Based on the final identification results of S4, map the 3D location and risk level of the small target to the virtual model in real time and mark the specific construction location of the small target. S52. Automatically match rectification plans based on risk levels: If the risk level is high, then the confidence level is < This triggers the generation of an "immediate shutdown" command, calls the wind power construction torque standard library, and outputs the calibration parameters for small target fasteners. If it is medium risk, then Confidence level < ,trigger" The "re-inspect within 24 hours" instruction is associated with the re-inspection path of the tracked mobile robot; The system calls upon the skeletal feature data of the workers in the digital twin model to extract the 3D coordinates of key human body parts such as the shoulders, waist, and chest; obtains the spatial coordinates of the safety belt suspension point through a safety belt recognition algorithm; compares the 3D coordinates of the key human body parts with the spatial coordinates of the safety belt suspension point according to the preset human safety suspension area standard, generates comparison results, triggers a "safety belt suspension violation" judgment based on the comparison results, and outputs a rectification instruction to "adjust the suspension point to the human safety suspension area". The comparison results include: Not suspended: No valid seatbelt suspension point coordinates detected; Improper suspension point location: The coordinates of the suspension point fall outside the safe suspension area for the human body; S53. After completing the rectification according to the rectification instruction, generate a rectification record and a rectification confirmation signal. Based on the rectification confirmation signal, control the tracked mobile robot to perform secondary data collection on the target small object, generate secondary recognition results, and repeat steps S2 to S4. When the confidence level of the secondary recognition result is ≥ At that time, complete the closed-loop archiving of potential hazards.
7. The method for full-link safety management of small targets in new energy construction based on cross-perspective feature collaboration and dynamic anti-disturbance closed loop as described in claim 3, is characterized in that, The specific steps of adapting the "adapted cross-layer feature fusion module" in step S22 are as follows: Remove the general cross-modal alignment logic from the original COXNet CLFM module and add a small target feature matching layer; Set a similarity threshold When the calculated cosine similarity between the high-frequency features after wavelet decomposition and the feature dictionary of new energy small targets is less than the similarity threshold... When the GeoShape similarity measurement label assignment strategy is triggered, the spatial shape similarity between the feature and the template is calculated. When determined The feature is retained when it is active, otherwise it is filtered as background noise to ensure that small target features are not misjudged.
8. The method for full-link safety management of small targets in new energy construction based on cross-perspective feature collaboration and dynamic anti-disturbance closed loop as described in claim 4, is characterized in that, The specific steps of the spatial pyramid multi-scale universal convolution in step S31 are as follows: Based on the small target feature map aligned in step S23, the feature map is divided into three scale layers: 1X, 2X, and 4X, using the spatial pyramid multi-scale general convolution algorithm. When the analysis yields the area of the small target candidate box When pixel², call Convolution kernels extract local texture features, convolution stride Weighting coefficient ; When the analysis shows that the area of the small target candidate box is > When pixel², call Convolution kernels extract contour features, convolution stride Weighting coefficient ; The convolution results of the three scale layers are weighted and fused to output a locally enhanced small target feature map.
9. The method for full-link safety management of small targets in new energy construction based on cross-perspective feature collaboration and dynamic anti-disturbance closed loop as described in claim 5, is characterized in that, The specific steps for confidence correction in step S43 are as follows: When the confidence level of the initial identification result in step S32 is determined At that time, the event stream data is invoked to extract the dynamic edge features of the small target and calculate the temporal continuity of the edge features. ; Set a continuity threshold When judging At the same time, a mapping relationship between image features and event flow features is constructed, transforming dynamic edge features into image domain edge enhancement weights. ; Based on the confidence level of the initial recognition result output in step S32 Image domain edge enhancement weights The sum of these values yields the corrected confidence score for small target recognition. ; when When, it is determined to be low risk; when When, it was determined to be a medium risk; when When a high-risk situation is identified, a high-priority rectification order is triggered, which is only sent to the person in charge on-site via SMS.
10. The method for full-link safety management of small targets in new energy construction based on cross-perspective feature collaboration and dynamic anti-disturbance closed loop as described in claim 6, characterized in that, The specific steps of the digital twin model mapping process in step S51 are as follows: Based on the 3D location of the small target in step S43, the construction area BIM model is called, and the 3D coordinates of the small target are transformed using a coordinate transformation matrix. The coordinate transformation matrix formula for mapping to the corresponding component in the BIM model is shown in equation (8) below: (8); In equation (8), , , Represented as the 3D coordinates of the small target in the BIM model; It is represented as a 3×3 coordinate rotation matrix, obtained through BIM model and on-site calibration of the construction area; , , This is represented by the 3D recognition coordinates of the small target output in step S43; It is represented as a 3×1 coordinate translation vector, obtained through calibration between the BIM model and the actual construction area; When the risk level of a small target is determined to be high risk, the small target is highlighted in the twin model as a red semi-transparent cube, and the risk type and associated sensor data are marked. The sensor data includes vibration amplitude and light intensity. Obtain the "Real-time View" command sent by the monitoring center through the remote communication interface, and determine the twin model according to the "Real-time View" command. The small target status is updated every frame per second, and scaling and rotation operations can be performed optionally.
Citation Information
Patent Citations
Power operator behavior identification early warning system and method based on video analysis
CN120220241A
Cited By
Mountain foundation pit construction monitoring optimization method, system, equipment and medium
CN122020436A
A mountain foundation pit construction monitoring optimization method, system, device and medium
CN122020436B