Collaborative optimization method for intelligent monitoring of photovoltaic construction progress
By employing a collaborative optimization method involving multi-parameter acquisition, multi-modal adaptive identification, and dynamic dual-link transmission, the problems of insufficient identification accuracy, poor transmission stability, and low calibration efficiency in photovoltaic power plant construction progress monitoring have been solved, enabling automated and real-time monitoring of photovoltaic panel installation progress.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-19
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies for monitoring the construction progress of photovoltaic power plants suffer from several problems, including insufficient accuracy in automatically identifying the progress of photovoltaic panel installation due to complex environments such as shading, light fluctuations, and dynamic interference; poor stability of multi-level monitoring dual-link transmission; and low efficiency of dual-source data calibration.
A collaborative optimization method is adopted, which includes multi-parameter acquisition, multi-modal adaptive recognition, dynamic dual-link transmission and PLC noise calibration. This method involves UAV carrying multiple sensors for data acquisition, combining feature encoders and axial attention processing of photovoltaic panel features, real-time link quality monitoring and switching, and data calibration combined with terrain feature matching.
It improves the accuracy of photovoltaic panel installation progress identification, ensures the stability and real-time transmission of multi-level monitoring data, enhances the efficiency of data calibration, and adapts to the automated monitoring needs of complex construction scenarios.
Smart Images

Figure CN121787625A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart photovoltaic power plant technology, and in particular to a collaborative optimization method for intelligent monitoring of photovoltaic construction progress. Background Technology
[0002] In the construction of large-scale new energy power plants such as photovoltaic power stations, accurate monitoring of construction progress is crucial to ensuring on-time delivery and cost control. Currently, these projects generally rely on on-site staff to monitor progress by taking photos, filling out daily reports, or uploading construction information via handheld terminals. This manual method is not only inefficient but also susceptible to subjective judgment biases, missing construction data, and interference from factors such as sandstorms and inclement weather, making it difficult to create continuous and objective records of construction progress. To improve visibility, some companies have attempted to deploy fixed-point cameras for video recording. However, such equipment suffers from poor deployment flexibility, fixed monitoring angles, and low levels of intelligence, only meeting basic needs for post-event review or remote live streaming, and failing to achieve automatic identification of construction status and quantitative assessment of progress.
[0003] In recent years, drone technology has become an important means of data collection in construction scenarios due to its advantages of flexibly covering vast construction sites and quickly acquiring high-definition image / video data. Currently, CN116301055B discloses a drone inspection method and system based on building construction. The scheme involves the drone automatically inspecting along a preset route, using a single image recognition algorithm to detect construction targets, uploading the inspection data to the management terminal through a single-level or simple multi-level data transmission link, and quantifying the construction progress based on the recognition results without active dual-source data calibration. However, this technical solution has several drawbacks. At the identification level, it fails to consider the complex interference factors unique to photovoltaic construction scenarios, such as shading of photovoltaic panels by scaffolding or material accumulation, light fluctuations in dawn / dusk or rainy weather, and dynamic interference caused by the movement of construction personnel. This leads to multimodal weak misalignment and missed detection of small targets in photovoltaic panel installation progress identification, and the identification accuracy drops significantly due to environmental factors. At the transmission level, the link switching logic is not clearly defined. For remote photovoltaic sites with wide distribution and complex terrain, data packet loss is easily caused by wireless signal attenuation, failing to meet the real-time requirements of multi-level supervision. Ultimately, it cannot solve the problems of weak misalignment in multimodal identification of photovoltaic panel installation progress, missed detection of small targets, wasted bandwidth and packet loss in multi-level supervision dual-link transmission, and noise interference and low efficiency in dual-source data calibration caused by shading, light fluctuations, and dynamic interference during photovoltaic farm construction. Summary of the Invention
[0004] To address the shortcomings of the existing technologies, the technical problem to be solved by this invention is to provide a collaborative optimization method for intelligent monitoring of photovoltaic construction progress. This method can solve the problems of insufficient accuracy in automatic identification of photovoltaic panel laying progress, poor stability of multi-level monitoring dual-link transmission, low efficiency of dual-source data calibration, and coexistence of noise data interference caused by complex environments such as shading, light fluctuations, and dynamic interference during photovoltaic farm construction, without significantly increasing hardware costs and computational latency.
[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: The present invention provides a collaborative optimization method for intelligent monitoring of photovoltaic construction progress, comprising the following steps: S1. UAV data acquisition steps: When the output from S6 is received... When a periodic trigger signal is received, the "UAV data acquisition and control" action is triggered, in which... To preset the cycle period, the specific steps include sub-steps S11 to S12: S11, Multi-parameter Acquisition Sub-step: Control the drone equipped with a high-definition RGB camera, infrared thermal imager, light sensor, and shading sensor to perform periodic data acquisition on the photovoltaic construction area; the drone's flight parameters are based on the preset layout of the photovoltaic power station equipment, and the flight altitude is set to the distance from the photovoltaic panels. Horizontal flight speed ,Every RGB and IR images are captured at intervals; among them, , To preset the height threshold, , To preset the shooting distance interval, The preset maximum flight speed is for a windless environment; when the light sensor detects that the sampling interval has reached... At that time, among them, To preset the light intensity collection interval, trigger the "light intensity collection" action: collect the light intensity once; when the shading sensor detects a change in distance to the photovoltaic panel via laser ranging exceeding [a certain threshold], [the action continues]. At that time, among them, To preset the threshold for changes in occlusion distance, calculate the occlusion ratio; when the GPS module detects a change in the drone's position exceeding... At that time, among them, To preset the GPS location change threshold, the latitude and longitude of the collected location are recorded synchronously; S12, Data Storage and Transmission Sub-step: After the multimodal raw dataset is generated, it is stored on the UAV's local SD card; the multimodal raw dataset includes RGB images, IR images, and environmental parameters; the resolution of the RGB images is [resolution missing]. , , The default resolution is the RGB image resolution, and the default resolution is the IR image resolution. , , The system presets the IR image resolution and environmental parameters, including illumination intensity, occlusion ratio, and terrain type. When the SD card storage is complete and the wireless connection signal of the ground control station is detected, the multimodal raw dataset is transmitted to the ground control station's cache unit via the UAV's wireless module. S2, Multimodal Adaptive Recognition Step: When the ground control station's buffer unit outputs the multimodal raw dataset generated in S1, the "Multimodal Recognition Start" action is triggered, specifically including sub-steps S21 to S23: S21. Construct a dual-modal encoder, including a specific feature encoder and an invariant feature encoder. These two encoders process the multimodal raw dataset in parallel, achieving separation and extraction of specific and common features. When the specific feature encoder detects an RGB image input, it employs a convolutional network containing a C3 module, with convolutional kernel sizes of [sizes to be filled in]. , , ,in, , To preset the convolution kernel size, Number of output channels ,in, To preset the number of RGB feature channels, through Convolutional layers capture the texture of the photovoltaic panel's edge. Convolutional layers reduce dimensionality, bottleneck layers enhance feature representation, and output RGB-specific features. When a specific feature encoder detects an IR image input, it captures the temperature distribution differences of the photovoltaic panel using the same network structure, outputting IR-specific features. When an invariant feature encoder detects both RGB and IR images simultaneously, it employs an activation function containing SiLU. Convolutional layer, where the number of output channels , To preset the number of common feature channels, Extract cross-modal common features such as photovoltaic panel frame and wiring. To solve the feature mismatch problem in multimodal weak misalignment scenarios; when Once the extraction is complete, the "Spatial Migration Prediction Start" action is triggered, which is executed in two steps: S211, when When input is sent to the space enhancement module, the "space enhancement processing" action is triggered: [This action is applied to...] Perform max pooling and average pooling to obtain and ;Will and Concatenate, input Convolutional layers, activated by Sigmoid, yield spatial attention weights. ;when After generation, and Performing the Hadamard product operation yields spatially enhanced features. To enhance the characteristics of the target area; among them, Pooling core, Preset pooling kernel size; Pooling core; Number of output channels: 1. To preset the convolution kernel size,
[0006] S212, when When input is sent to the channel enhancement module, the "channel enhancement processing" action is triggered: [The text abruptly ends here, likely due to an incomplete sentence or a formatting error.] Perform global average pooling and global max pooling to obtain and The two are input into a shared multilayer perceptron, and the output is the channel attention weights. ;when After generation, the "channel weighting" action is triggered: and Element-wise multiplication yields the offset prediction features. Suppress channel noise; when offset prediction features After generation, the "spatial offset calculation" action is triggered: based on offset prediction features. The enhanced features after offset prediction are calculated using the formula. : Calculate spatial offset (1); where, This represents the enhanced features after the offset prediction; This represents the Sigmoid activation function; This represents a shared multilayer perceptron; AvgPool represents the global average pooling operation; MaxPool represents the global max pooling operation. This indicates the differences in shared features between RGB and IR. ; Represents the cross-modal shared features of RGB images; Represents the cross-modal shared features of IR images; Indicates spatial enhancement features, ; This represents the spatial offset between the RGB and IR images; S22, when S21 outputs When input is fed into the CSP layer of the backbone network, the "axial attention computation start" action is triggered, which is performed through bidirectional compression and positional encoding injection steps: S221. When the axial attention calculation start signal is triggered, for The "horizontal compression" action is triggered along the horizontal direction, through the formula. (2) Calculate and output the horizontal compression features. , where the dimension is ; Trigger the "vertical compression" action along the vertical direction, through the formula (3) Calculate and output the vertical compression feature. , where the dimension is Reduce computational complexity, from Down to ; in, Indicates the first horizontally compressed number. Line characteristics; express The width; express No. Line 1 The pixel values of the column; express Height; Represents the first vertically compressed [unit]. Column features; Represents the horizontally compressed feature matrix; Represents the vertically compressed feature matrix; Indicates the number of shared feature channels; S222, when and After generation, the positional encoding generation action is triggered to generate learnable positional codes. , and , Once the location code is generated, the "code overlay" action is triggered: ... and query vector Add, and key vector Add, and query vector Add, and key vector Add them together, preserving the global spatial relationships; where, Dimensions , The dimension is and The dimension is , The dimension is ; in, The positional encoding of the horizontal query vector; Represents the positional encoding of the horizontal key vector; This represents the positional encoding of the query vector in the vertical direction; This represents the positional encoding of the vertical key vector; express The query vector; express The key vector; express The query vector; express The key vector; S223. After the encoding is superimposed, trigger the "horizontal self-attention" action along the horizontal direction, through the formula. (4) Calculation; Trigger the "vertical self-attention" action along the vertical direction, using the formula (5) Calculation; After the horizontal and vertical self-attention calculations are completed, trigger the "attention fusion" action: and Adding them together yields the global attention enhancement feature. Improve the accuracy of small target recognition; in, This represents the result of the horizontal self-attention calculation; Indicates the first in the horizontal direction softmax normalization of each element; express The query vector's first One component; The first horizontal position code One component; express The first key vector One component; The first horizontal position code One component; express The value vector of the first One component; This represents the result of vertical self-attention calculation; Indicates the first vertical direction softmax normalization of each element; express The query vector's first One component; The first vertical position code One component; express The first key vector One component; The first vertical position code One component; express The value vector of the first One component; This indicates a global attention enhancement feature; S23, when S22 outputs When input is applied to the three branch networks for low light, high occlusion, and dynamic scenes, a multi-level parameter fusion startup action is triggered: S231, When the branch network initialization is complete and After input, for the identity layer of each branch, a virtual layer construction action is triggered to construct the virtual layer. Convolutional layers and virtual batch normalization layers; for each branch's batch normalization layer, the virtual convolutional layer construction action is triggered, and only the above virtual convolutional layers are constructed to achieve multi-branch parameter alignment; Among them, the kernel index The value is 1 at one location and 0 at the rest. For the center index of the convolution kernel, express Integer division results; moving weighted variance of virtual batch normalization layer Training parameters Moving weighted average Training parameters , Indicates the first Moving weighted variance of each batch of standardized layers; Indicates the first Scaling parameters for each batch of normalized layers; Indicates the first Moving weighted mean of each batch of standardized layers; Indicates the first Offset parameters for each batch of standardized layers; S232. After the alignment phase is complete, for each branch's convolutional layer and batch normalized layer, trigger the parameter fusion action using the formula. (6) and (7) Merge into a biased convolutional layer; in, Indicates the fused first Each convolutional kernel parameter; Indicates the original number Each convolutional kernel parameter; This indicates a preset minimum value, used to avoid a denominator of 0; Indicates the fused first Each convolutional layer bias parameter; S233, When single-level fusion is completed and the biased convolutional layer parameters of each branch are ( , , After generation, a parameter addition action is triggered, adding the parameters of the convolutional kernels from the fused three branches to obtain an equivalent single convolutional BN layer, achieving a significant reduction in the number of parameters, with a reduction ratio of [percentage missing]. , To reduce the threshold for the preset parameter quantity; where, , , These represent the convolutional kernel parameters after branch fusion in low light, high occlusion, and dynamic scene conditions, respectively. This represents the minimum threshold for reducing the number of parameters; when an equivalent single convolutional BN layer is generated and the output prediction box is... Then, the SSIoU loss function is used to optimize the photovoltaic panel bounding box, and the positioning accuracy is improved through multi-dimensional loss fusion. The formula is as follows: (8); where SSIoU represents the Softscaled intersection of union loss; IoU represents the intersection of the predicted box and the ground truth box; Indicates angle loss; Indicates distance loss; Indicates shape loss; Furthermore, the formula for calculating IoU is: (9); in, Represents the prediction box With real frame The area of their intersection; Represents the prediction box With real frame The area of the union of the sets; The bounding boxes represent ground truth boxes from a manually labeled dataset; The calculation formula is (10); in, Indicates the dynamic correction factor; Indicates the preset angle coefficient; This represents the angle between the line connecting the center points of the ground truth bounding box and the center point of the predicted bounding box, and the horizontal line of the ground truth bounding box. ; , These represent the prediction boxes. Realistic frame The y-coordinate of the center point; This represents the straight-line distance between the center point of the predicted bounding box and the center point of the ground truth bounding box. ; , These represent the prediction boxes. Realistic frame The x-coordinate of the center point; The calculation formula is (11); , They represent enclosed areas respectively. and The width and height of the smallest rectangle; The calculation formula is (12); in, This represents the horizontal distance loss component. (13); This represents the vertical distance loss component. (14); The calculation formula is (15); in, This represents the shape loss component in the width direction. (16); This represents the shape loss component in the height direction. (17); , These represent the prediction boxes. Width and height; , Representing the actual bounding boxes Width and height; Indicates the preset shape loss coefficient; Once the bounding box optimization is complete, the "Recognition Result Output" action is triggered: outputting the quantitative data of photovoltaic panel installation. and data priority labels Among them, quantitative data on photovoltaic panel installation. Data priority labels include the number of installations and the percentage of the area covered. Based on the quantitative data of photovoltaic panel installation Label the image as "high priority" and label it as "low priority" based on the auxiliary image; S3, Dynamic Dual-Link Transmission Steps: When the photovoltaic panel deployment quantization data output by S2... and data priority labels Upon arrival at the transmission node and detection of the transmission start signal, a dynamic dual-link transmission start action is triggered, specifically including sub-steps S31 to S33: S31, Link Quality Monitoring Sub-step: Deploy link quality monitoring units at the Level 5 transmission nodes. When the monitoring unit receives the start signal, each... Intermittent triggering of link parameter acquisition actions; among which, the five-level transmission nodes include the field acquisition terminal, the booster station, the regional video monitoring center, the company-level monitoring room, and the headquarters-level monitoring room; For the preset monitoring interval, the link parameters include link bandwidth. Real-time signal strength of the link and link real-time transmission delay After each link parameter collection is completed, the link score calculation is triggered using the formula. (18) Calculate the link stability score; where Score represents the link stability score; , , This represents the preset weight coefficients, and satisfies... ; S32. When link quality monitoring is started, trigger the threshold parameter initialization action: initialize parameters, including the initialization of link switching thresholds. Preset target update rate Preset threshold adjustment step size When the system clock detects that the time interval has reached At that time, through the formula (19) Statistical link update rate; among which, Indicates the first Link update rate for each statistical period; Indicates as of the date The number of link updates at any given time (taken from the monitoring unit log); Indicates as of the date The number of link updates at any given time; Indicates the update rate statistical period; when After the calculation is complete, the "threshold adaptive adjustment" action is triggered: if (If the update rate is lower than the target), then ;like (If the update rate is higher than the target), then (20) To achieve a stable link update rate; among which, Indicates the first Link switching threshold for each cycle; Indicates the first Link switching threshold for each cycle; Indicates the threshold adjustment step size; This indicates the preset link target update rate; S33. After the Score calculation output by S31 is completed, trigger the link adaptive selection action: If Choose a 5G link; if Switch to LoRa link; among which, The handover threshold indicating link quality; when the link is selected as a 5G link and Once loading is complete, trigger high-priority data transfer actions: [to...] Using UDP protocol, transmission rate ,in To preset a high-priority transmission rate, the breakpoint resume enable action is triggered after UDP protocol initialization: Add the following to the data frame header: Position number + Bit verification, where , For preset bit length, ;in, Indicates the port number for UDP protocol transmission; This indicates the transmission rate of high-priority data; The number of bits representing the data frame sequence number; Indicates the bit length of the data frame checksum; when high-priority transmission starts and the auxiliary image is loaded, low-priority data transmission is triggered: the auxiliary image uses the TCP protocol, and image compression optimization is triggered after the TCP connection is established: JPEG2000 compression is used, where the compression ratio is... , For the preset compression ratio, Irreversible mode; in, Indicates the port number used for TCP protocol transmission; Indicates the image compression ratio; when the link is selected as a LoRa link and the photovoltaic panel deployment quantization data is used. Once loading is complete, the core data priority transmission action is triggered: quantification data of photovoltaic panel installation. The LoRaWAN protocol is adopted; where SF represents the spreading factor of the LoRa protocol. This indicates the transmission bandwidth of the LoRa protocol; once LoRa link transmission starts, each Check the score once at intervals: if This triggers the auxiliary image retransmission action; among which, Indicates the judgment interval for auxiliary image retransmission; When the transmission is complete and the photovoltaic panel installation data is quantified When the auxiliary image arrives at the calibration end, complete data is obtained. According to complete data Trigger the data output action: output the complete data Send to the calibration end; among which, This represents the complete data transmitted to the calibration end, including quantization data and auxiliary images; S4, when S3 outputs Arrived at the calibration end and detected Includes photovoltaic panel identification confidence level When this occurs, the PLC noise calibration start action is triggered, specifically including sub-steps S41 to S43: S41, to The photovoltaic panel identification results are categorized according to the definition of PMD noise, where PMD noise is feature-related noise, constrained to the middle confidence region: if ,but (21); if ,but (22); among which, This represents the probability that a true label of 1 is mislabeled as 0. c1 and c2 represent the probability that a true label of 0 is mislabeled as 1; c1 and c2 represent preset noise coefficients, taken from historical noise statistics. Indicates the confidence level of photovoltaic panel identification; The confidence interval threshold represents the PMD noise. S42. After the multimodal recognition model output by S2 is loaded onto the calibration end, the Warmup training start action is triggered, with batch size... Learning rate SGD optimizer training Rounds, when the number of training rounds reaches And the validation set At this point, training is stopped, and a preliminary model is obtained. ;in, Indicates the batch size during training; Indicates the initial learning rate; This represents the momentum coefficient of the SGD optimizer; This represents the weight decay coefficient of the SGD optimizer; This indicates the number of training epochs in Warmup; mAP represents the mean precision. The minimum threshold representing the verification accuracy; This represents the initial model after Warmup training; when After generation, set the initial threshold. Calculate the calibration threshold ;when Generate and Once the image samples are loaded, the confidence screening and calibration action is triggered: for each image sample, calculate... softmax output ,like and If the predicted label differs from the manually labeled label, a dynamic label update is triggered: through the formula... (23) Update the tags; among which, Indicates the initial calibration threshold; This indicates the calibration threshold actually used for screening; express For the sample The softmax output (recognition confidence); This indicates the updated tag; Indicates the indicator function; 3) When the label update is complete and the cumulative training rounds reach At that time, Updated to Synchronous updates Repeat label calibration until continuous. The cycle of labelless updates yields a calibration label set. ;in, Indicates the interval between iterations of the threshold; Indicates the first The threshold of the wheel; Indicates the step size of the threshold iteration; This indicates the number of consecutive iterations without updates that have stopped. This represents the calibrated tag set; To preset the stopping iteration round, when continuous iterations are detected... The stop check is triggered when the wheel label update count reaches 0. S43. When the terrain type parameters output by S1 reach the calibration end, the terrain adaptive matching action is triggered: If Matching mountain landmark segment threshold ;like Matching plain section threshold ;like Matching the threshold of the gentle slope section ;in, Indicates the elevation difference of the photovoltaic construction area; The threshold representing the elevation difference in mountainous terrain; The threshold for elevation difference in plain terrain; This indicates the deviation threshold for mountain sections; This indicates the deviation threshold for the plain section; This represents the deviation threshold for the gentle slope section; once the terrain assessment is completed and the corresponding threshold is matched, the deviation calculation and calibration action is triggered: Calculation Compared with manually reported data deviation (24); among which, This indicates the percentage deviation between the quantitative data and the manually reported data; This represents the quantitative data on photovoltaic panel installation output by S2; This represents manually reported photovoltaic panel installation data; when Calculation completed and When the deviation exceeds the matching threshold, a secondary precision calibration is triggered, repeating step S42. When the deviation calculation is complete and no secondary calibration is required, or after the secondary calibration is completed, the calibration result output is triggered: outputting the calibrated photovoltaic construction progress data. And model optimization parameters, including the dual-modal encoder convolution kernel parameters. Axial attention weight , Used to generate inspection reports and As input to the S5 model iteration steps; where, This indicates the calibrated photovoltaic construction progress data; This represents the optimized convolution kernel parameters of the dual-modal encoder; This represents the optimized weight parameters of the axial attention module; S5, when S4 outputs and Upon reaching the model iteration module, the model parameter iteration update action is triggered, specifically including sub-steps S51 to S52: S51, when After loading is complete, Inputting the bimodal encoder S21, updates the convolution kernel weights of the feature-specific encoder C3 module and the SiLU activation layer parameters of the invariant feature encoder; where, The momentum coefficient representing the momentum gradient descent; S52, when Once loading is complete, the "Attention Weight Optimization" action is triggered: [The action will be adjusted accordingly.] Input the SeaAttention module of S22, update the weights of horizontal / vertical self-attention, and use an adaptive learning rate with an initial learning rate. When the absolute value of the gradient of the loss function Trigger learning rate multiplied by ( When the absolute value of the gradient Trigger learning rate multiplied by ( ), , To preset the gradient threshold, ;in, This represents the initial learning rate for updating attention weights; , Indicates the gradient threshold; This represents the learning rate decay coefficient; This represents the learning rate increment factor; after the parameters are updated, the updated multimodal recognition model will be... The model is stored at the ground control station as the identification model for the next S2 step, enabling iterative optimization of the model. S6. After the system starts and an initialization completion signal is detected, the "loop control start" action is triggered, which includes sub-steps S61 to S63: S61. Deploy a timer at the ground control station. When the timer detects that the system clock has reached its maximum speed... At intervals, a periodic trigger signal generation action is initiated; among which, Indicates the period of the loop execution; S62. After the trigger signal is generated, the trigger signal is sent to the UAV and each transmission node to start the S1 UAV data acquisition step. S63. After all steps S1 to S5 are completed, a loop log is recorded. The loop log includes the loop start time, the execution time of each step, and the amount of output data. The log is stored in the database. Storage is triggered when the database connection is normal, supporting subsequent querying and analysis. The format of the loop log is "timestamp, step name, execution time, and amount of output data". The database is MySQL or other relational database. In the preferred embodiment, in S21, when the invariant feature encoder... After the convolutional layer outputs features, the SiLU activation calculation is triggered: the formula for the SiLU activation function is as follows. (25), of which (26); among which, This represents the output value of the SiLU activation function; This represents the input value of the activation function; This represents the Sigmoid activation function; Represents the natural exponential function; It is implemented by computer through exponential operations (exponential calculation is triggered when x is input) and division operations (division calculation is triggered when the exponential result is generated), which enhances the nonlinear expression capability of features.
[0007] In the preferred embodiment, in S22, after feature compression is completed, the following steps are taken: and Then, position encoding The generation logic is as follows: based on The dimension is used to generate a randomly initialized parameter matrix. After the matrix initialization is complete, backpropagation is triggered, and the model is updated iteratively through backpropagation. After the model loss is calculated, parameter adjustment is triggered to make the encoded matrix more stable. It can capture the spatial relationship of photovoltaic panels; among them, This represents the preset initialization range threshold for the positional encoding parameters; after the query and key vectors of the self-attention function are added together, the softmax calculation is triggered: the softmax function in the self-attention calculation is determined by the computer using the formula... Among them, horizontal self-attention, when (27) Triggered after calculation is completed; Among them, vertical self-attention, when (28) Triggered after calculation is completed; in, Indicates the horizontal direction. The softmax output of each element; Indicates the horizontal direction. Attention score for each element; express The natural index; This represents the sum of the indices of all attention scores in the horizontal direction; Indicates the vertical direction of the first The softmax output of each element; Indicates the vertical direction of the first Attention score for each element; express The natural index; This represents the exponential sum of all attention scores in the vertical direction.
[0008] In the preferred scheme, in S23, when multi-level parameter fusion is initiated and the number of channels of the branch input tensor is... Once confirmed, the "Parameter Reduction Rate Calculation" action is triggered. The calculation logic for the parameter reduction rate of multi-level parameter fusion is as follows: convolution kernel size Branch, parameter reduction rate (29); kernel size Branch, parameter reduction rate (30); among which, express The rate of reduction in the number of parameters in the convolution kernel branch; express The rate of reduction in the number of parameters in the convolution kernel branch; This represents the number of channels in the branch input tensor; n1 and n2 represent preset constants; Represent the square of K1; The computer performs calculations using integer and division operations to ensure that the number of parameters is reduced to meet the requirements. Threshold requirements.
[0009] In the preferred embodiment, in S31, after link quality monitoring is initiated, the monitoring unit is triggered to operate. The operating logic of the link quality monitoring unit is as follows: when the bandwidth sensor detects... At time intervals, bandwidth calculation is triggered: through statistics Calculation of the number of bytes in the internal data packet When the signal strength detector receives a wireless signal, it triggers a signal strength conversion process: converting the RSSI value of the received signal into a signal strength value. When the delay timer detects a data packet transmission signal, it triggers a delay recording action: recording the time difference between the data packet transmission time from the sender to the receiver. ;in, The time interval for bandwidth statistics is indicated; RSSI indicates received signal strength. when , , Once all data collection is complete, a score normalization process is triggered: the score is calculated by the computer first... , , Normalization is performed, where, Normalize to [0,1], Normalize to [0,1], Normalize to [0,1]. Once normalization is complete, trigger weight summation, then press... , , Weighted summation; Normalization means mapping parameter values to the interval [0,1], and the formula is as follows: (31), among which, The normalized value. The original value, For the minimum value of the parameter, Set to the maximum value of the parameter; ensure that the score objectively reflects the link quality.
[0010] In the preferred embodiment, in S32, when the system clock... When an interruption is triggered, an update rate statistics action is initiated, and the link update rate is recorded. The statistical logic is as follows: each time the monitoring unit records a change in link parameters, the following data is recorded: fluctuation or fluctuation or fluctuation Triggered after parameter fluctuation detection is completed Add 1, then Add 1; in, The threshold representing bandwidth fluctuation; The threshold representing signal strength fluctuations; The threshold representing transmission delay fluctuation; Indicates the end time The number of link updates; Computer every Read The current value, and Before Subtraction, when After reading is complete, subtraction is triggered, and the result is obtained by dividing by t3. ;when and After the size relationship is determined, a threshold adjustment operation is triggered: the threshold adjustment operation is performed by the computer to update the threshold by performing addition or subtraction operations based on the size relationship. Ensure the link update rate remains stable at [value missing]. nearby.
[0011] In the preferred embodiment, in step S42, after the multimodal recognition model has been loaded and the training / validation set has been partitioned, according to... Division, , To pre-determine the division ratio, Once the dataset is partitioned, the "Warmup Training" action is triggered. The Warmup Training logic is as follows: the computer loads the multimodal raw dataset output by S1, and... It is divided into a training set and a validation set; among which, Indicates the proportion of the training set; Indicates the proportion of the validation set; During each iteration, when After all training samples are loaded, trigger the sample input action: retrieve Input of training samples Calculate the SSIoU loss, and once the loss is calculated, trigger backpropagation to update the model parameters; training. After the validation set is loaded, the validation set evaluation action is triggered: the computer calculates on the validation set. ,like If so, stop Warmup training; otherwise, increase... Round training; among them, Indicates the round of supplementary training; When the confidence level is satisfied and the label comparison is inconsistent, the label overwriting action is triggered: the label update action, which is performed by the computer for comparison. If the predicted labels and manually labeled labels are inconsistent, This will overwrite the original label with the predicted label, thus improving label accuracy.
[0012] In the preferred embodiment, in step S43, after the GPS data output from step S1 is loaded, the elevation difference calculation is triggered. The calculation logic for the elevation difference is as follows: based on the GPS data, the computer selects... The grid is used to calculate the difference between the maximum and minimum elevations within the grid. ;in, This represents the grid edge length used to calculate the elevation difference; when and Once the units are unified, the "deviation calculation" action is triggered: Deviation The calculation is first performed by the computer. and Standardize the units, then calculate the percentage deviation using the formula. The computer automatically sends calibration instructions to the manual annotation end to obtain the corrected threshold. Ensure the deviation is within the allowable range.
[0013] In the preferred scheme, in S5, when After all current parameters of the dual-modal encoder have been loaded, the parameter fusion action is triggered. The parameter update logic is as follows: the computer will... The parameters are then weighted and fused with the current parameters of the dual-modal encoder, where the weights are: occupy The current parameter accounts for , , Once the weights are set, fusion is triggered; among them, express The fusion weights; This indicates the fusion weights of the current parameters; it avoids model instability caused by sudden parameter changes; after the gradient of the loss function is calculated, it triggers the learning rate adjustment: the adaptive learning rate is updated by the computer according to the gradient change of the loss function, if the absolute value of the gradient is... Learning rate multiplied by If the absolute value of the gradient Learning rate multiplied by ; in, This represents the learning rate decay coefficient; This represents the learning rate amplification factor; In the preferred embodiment, in step S6, after system initialization is complete and the timer start signal is triggered, the timer operation is initiated. The timer's operation logic is as follows: after the computer starts the timer, every... An interrupt signal is triggered. Once the interrupt signal is generated, the UAV's flight control module executes the S1 data acquisition action. When the flight control module receives the interrupt signal, it triggers the S3 monitoring action, which is triggered simultaneously by the monitoring units of each transmission node upon receiving the interrupt signal. When all S1 and S5 steps of each cycle are completed and the log data is collected, the log storage action is triggered: the cycle log is stored in the database by the computer in the format of "timestamp, step name, execution duration, output data volume", supporting subsequent query and analysis; thus achieving full-link process traceability.
[0014] This invention provides a collaborative optimization method for intelligent monitoring of photovoltaic construction progress. Compared with existing methods, this method offers the following advantages through the coordination of the aforementioned structures: First, by integrating modal separation coding, axial attention and multi-level parameter fusion technology at the recognition end, it specifically addresses common issues in photovoltaic construction such as scaffolding, material shading, morning and evening, light fluctuations due to rain and cloudy weather, and dynamic interference from construction personnel. This effectively solves the problems of weak misalignment in multimodal images and missed detection of small targets on photovoltaic panels with incomplete edges, thereby improving the accuracy and environmental robustness of photovoltaic panel laying progress recognition and avoiding the impact of environmental factors on progress judgment due to recognition deviations. Secondly, it can adjust the link switching threshold in real time based on bandwidth, signal strength, and transmission delay, and allocate transmission resources according to data priority. This effectively addresses the signal attenuation problem in remote photovoltaic power stations, reduces data packet loss, avoids bandwidth waste caused by a single link, ensures the stability and real-time transmission of multi-level regulatory data from the field to the headquarters, and breaks down the barriers between regulatory levels. Third, it can introduce PMD noise classification and PLC progressive label calibration algorithm through the calibration end, and combine the photovoltaic site terrain features to match the segmented dynamic deviation threshold. It can accurately filter out noise interference caused by shading and insufficient light in the identification data, reduce the workload of manual review, improve the efficiency of dual-source data calibration, and at the same time feed back the correct calibrated labels to the identification model to continuously optimize the model parameters and enhance the model's adaptability to complex construction scenarios. Attached Figure Description
[0015] The present invention will be further described below with reference to the accompanying drawings and embodiments: Figure 1 This is a main view diagram of the process structure of this invention; Figure 2 This is a flowchart of Embodiment 1 of the present invention. Detailed Implementation
[0016] To better understand the purpose, system architecture, and functional implementation of this embodiment, the embodiments and features in the embodiments of this application can be combined with each other without conflict. The exemplary embodiments disclosed in this application will be described below with reference to the accompanying drawings, which include specific technical details disclosed in this embodiment to aid understanding; however, these details should be considered exemplary rather than restrictive. Therefore, those skilled in the art should understand that various improvements and adjustments can be made to the embodiments described herein without departing from the scope and core ideas of the invention. Similarly, for clarity, detailed descriptions of well-known technologies, functions, and structures (hardware configuration of conventional obstacle avoidance modules for UAVs, basic communication logic of general wireless transmission protocols, conventional storage architecture of standard databases, etc.) are omitted in the following description.
[0017] In the field of new energy power plant construction technology, especially in the area of construction progress monitoring for large-scale photovoltaic power plant projects, achieving automated, precise, and real-time monitoring of construction progress in complex environments is a core technological requirement for ensuring on-time project delivery and cost control. Currently, photovoltaic power plants are generally characterized by wide construction areas, remote site distribution, and complex terrain. Traditional manual inspections are inefficient and easily affected by subjective factors. Fixed-point camera monitoring lacks flexibility and intelligence, only meeting the needs of post-event review. Although drone technology is used for data collection, existing solutions mostly remain at the "shooting and recording" level, lacking deep integration with multimodal image recognition, dynamic link transmission, and noise data calibration, making it difficult to form a complete monitoring system from data collection to progress quantification. With the development of artificial intelligence and communication technologies, multimodal adaptive recognition, dynamic dual-link transmission, and progressive noise calibration (PLC algorithm) have become key technological directions for overcoming these bottlenecks. Through multi-technology collaboration, automatic identification, stable transmission, and precise calibration of photovoltaic panel installation progress can be achieved, providing reliable support for multi-level monitoring.
[0018] However, the relevant technologies have significant limitations when applied to photovoltaic construction progress monitoring. For example, the solution disclosed in CN116301055B, which involves UAV pre-set route inspection, single image recognition, simple multi-level transmission, and no dual-source calibration, has obvious defects in its core technologies: In terms of recognition, it lacks an adaptive feature processing mechanism designed for shading and light fluctuations in photovoltaic construction scenarios, relying solely on a single recognition algorithm. This leads to weak misalignment in multimodal images, missed detection of small targets on incomplete photovoltaic panels, and a significant drop in recognition accuracy due to environmental influences. In terms of transmission, it lacks clear link quality monitoring and switching logic, making it prone to data packet loss in remote sites with signal attenuation. It cannot balance real-time transmission with bandwidth utilization, failing to meet the needs of multi-level monitoring. In terms of calibration, it lacks an active calibration mechanism for feature-related noise, failing to effectively filter noise interference and resulting in low calibration efficiency. In summary, the relevant technologies only improve a single aspect in isolation, failing to form a collaborative optimization mechanism for recognition, transmission, and calibration, making it difficult to adapt to the automated monitoring needs of complex photovoltaic power plant construction scenarios.
[0019] Example 1 This embodiment will describe in detail the training process of the multimodal recognition model and the PLC calibration module.
[0020] Specifically, the model training data comes from historical construction data of the photovoltaic base in Xilingol League, Inner Mongolia, including 20,000 sets of RGB-IR dual-modal image pairs and corresponding photovoltaic panel laying progress labels. The image resolution has been uniformly adjusted to [resolution value missing]. (RGB) (IR), environmental parameters include light intensity, occlusion ratio, and terrain type. The data is divided into a training set (16,000 sets) and a validation set (4,000 sets) in an 8:2 ratio.
[0021] In the preferred embodiment, the training of the dual-modal encoder is performed according to step S21. Based on the OAFA method: the feature-specific encoder uses a convolutional network containing a C3 module, and the convolutional kernel size... , Number of output channels The C3 module contains 3 convolutional layers (stride 1, padding 1) and 2 bottleneck layers (1×1 convolutions reduced to 32 channels, then restored by 3×3 convolutions); the invariant feature encoder uses 1×1 convolutional layers ( Number of output channels The activation function is the SiLU function, obtained through the formula... (25) and formula (26) Implement nonlinear mapping; where, This represents the output value of the SiLU activation function; This represents the input value of the activation function; Represents the Sigmoid activation function; where, This represents the output value of the Sigmoid activation function; Represented by natural constant as the base It is the natural exponential function of the exponent; This represents the input value for the Sigmoid activation function; The exponential operations mentioned above are accelerated by the GPU's CUDA cores, and the division operations are optimized using floating-point precision to ensure improved feature extraction efficiency.
[0022] In this embodiment, the training of the axial attention module is performed according to step S22. The Sea-Attention mechanism: the feature compression stage is performed according to the formula... (2) and formula =(3)to (Dimensions 32×256×512) undergoes horizontal / vertical compression to obtain (32×256×1) and (32×1×512); in, Indicates the first horizontally compressed number. row eigenvalues; Representation of spatial augmentation features The width; express No. Line 1 The pixel values of the column; Indicates to No. Row from column 1 to column 2 Sum the pixel values of the column; Represents the first vertically compressed [unit]. Column feature values; Representation of spatial augmentation features Height; express No. Line 1 The pixel values of the column; Indicates to No. Columns from row 1 to row 2 Sum the pixel values of the row; Location coding , , , Initialize to The random matrix is updated iteratively through backpropagation, and the formula is used during training. (27) and formula (28) Calculate the self-attention weights; where, Indicates the horizontal direction. The softmax normalized output of each element; Indicates the horizontal direction. Attention score for each element; Represented by natural constant as the base It is the natural exponential function of the exponent; This indicates the horizontal direction from the 1st to the 1st. Sum the attention score index values of each element; Representation of spatial augmentation features Height; Indicates the vertical direction of the first The softmax normalized output of each element; Indicates the vertical direction of the first Attention score for each element; Represented by natural constant as the base It is the natural exponential function of the exponent; This indicates the vertical direction from the 1st to the th. Sum the attention score index values of each element; Representation of spatial augmentation features The width; , After attention fusion The recall rate for identifying photovoltaic panels with incomplete edges is improved, which is an improvement over traditional two-dimensional self-attention.
[0023] In practice, the training of the multi-level parameter fusion module is performed according to step S23, with branch fusion technology: targeting low light conditions ( ), high shading ( ), dynamic scenes ( The three branches first construct a virtual 3×3 convolutional layer and a virtual batch normalization layer according to the "alignment stage"; Then follow the formula (6) and formula (7) ) are merged into a biased convolutional layer; where, Indicates the fused first Each convolutional kernel parameter; Indicates the first Scaling parameters for each batch of normalized layers; Indicates the original number Each convolutional kernel parameter; Indicates the first Moving weighted variance of each batch of standardized layers; This indicates a preset minimum value, used to avoid a denominator of 0; Indicates the fused first Each convolutional layer bias parameter; Indicates the first Scaling parameters for each batch of normalized layers; Indicates the first Moving weighted mean of each batch of standardized layers; Indicates the first Moving weighted variance of each batch of standardized layers; Indicates the preset minimum value; Indicates the first Offset parameters for each batch of standardized layers; Finally, the three branches were merged according to the "multi-level fusion" rule. , , Adding them together, we get an equivalent single convolutional-BN layer.
[0024] In one feasible approach, the training of the PLC calibration module is performed according to step S42, using a progressive label calibration algorithm: the warm-up phase is performed in batches. Initial learning rate SGD optimizer training In the first round, when the mAP on the validation set reaches 85.2%, the warm-up process stops, and a preliminary model is obtained. Initial threshold during calibration phase ,according to Screening high-confidence samples, for Furthermore, for samples whose predicted labels differ from manually labeled labels, the formula is used. (23) Update tags; in, This indicates the updated tag; This indicates an indicator function that outputs 1 if the condition within the parentheses is true, and 0 otherwise. Representing the preliminary model For the sample The softmax output (recognition confidence); Each cumulative After round of training, according to Update the threshold until it is continuous No tag update in the cycle, at this time the calibrated tag set The noise rate decreased from 12% to 2.1%, PMD noise constraint, formula: (21), formula (twenty two); in, c1 represents the probability that a true label of 1 is mislabeled as 0; c1 represents the preset noise coefficient. c2 represents the confidence level for photovoltaic panel identification; c2 represents the preset noise figure. express of Power of; c1 represents the probability that a true label of 0 is mistakenly labeled as 1; c1 represents the preset noise coefficient, with a value of 0.8. c1 represents the confidence level for photovoltaic panel identification; c2 represents the preset noise figure, with a value of 0.5. express of Power of; This represents the confidence interval threshold for PMD noise, with a value of 0.1.
[0025] In this embodiment, the model iteration is performed according to step S5: the optimized parameters output by the PLC calibration are used to generate the dual-modal encoder convolution kernel. With axial attention weight The input model uses momentum gradient descent to update the weights of the feature-specific encoder and an adaptive learning rate to update the attention weights, where the initial... absolute value of gradient Multiply by time , Multiply by time ;in, The momentum coefficient representing the momentum gradient descent; This represents the initial learning rate for updating attention weights; This represents the learning rate decay coefficient; This represents the learning rate amplification factor; Example 2 This embodiment focuses on a sunny day scenario for a photovoltaic power plant in the Ordos Plain of Inner Mongolia.
[0026] In this embodiment, after the system initialization is completed, S6 deploys the timer according to the timer initialization steps and sets the cycle period. Every 30 minutes, when the system clock reaches a 30-minute interval, a periodic trigger signal is generated. The generated trigger signal is distributed throughout the signal chain and sent to the DJI M300 drone and the level 5 transmission node, starting the S1 drone data acquisition step. At this time, the trigger signal of S6 is the only condition for starting S1.
[0027] Specifically, S1 is executed according to steps S11 to S12: the UAV flight parameters are set. , (Height above photovoltaic panel), horizontal flight speed ,Every Capture an RGB-IR image once; the light sensor captures an RGB-IR image every time. A light intensity was collected once, and the shading sensor detected a change in distance from the photovoltaic panel. (Unobstructed trigger) The GPS module detected a change in location. Latitude and longitude are recorded in real time; after the data collection is completed, the multimodal raw dataset (RGB: 1920×1080, IR: 640×512, environmental parameters: illumination, occlusion, plain terrain) is stored to an SD card and transmitted to the ground control station cache unit via a 5G wireless module.
[0028] In the preferred scheme, S2 is executed according to steps S21-S23. The OAFA method is as follows: In the S21 dual-modal encoder, the specific feature encoder extracts the RGB texture features (clear photovoltaic panel border) and the IR temperature features (photovoltaic panel temperature 25℃, background 30℃), while the invariant feature encoder extracts cross-modal common features. (Rectangular outline of photovoltaic panel); Spatial offset prediction stage, for Performing 2×2 pooling yields and After 7×7 convolution and Sigmoid, we get According to the Hadama product, Then, through global pooling and MLP, we obtain... Finally, according to the formula (1) Calculate spatial offset Pixel-level, minimal multimodal misalignment in sunny scenes.
[0029] In this embodiment, the S22 axial attention module is executed according to formulas (2), (3), (4), and (5); In the case of formulas (2) and (3), the variables are explained in the same way as in Example 1; in, This represents the result of the horizontal self-attention calculation; Indicates the first in the horizontal direction softmax normalization of each element; Indicates horizontal compression features The query vector of the first One component; The horizontal position code is represented by the first one. One component; express transpose; Indicates horizontal compression features The key vector of the first One component; The horizontal position code is represented by the first one. One component; Indicates horizontal compression features The value vector of the first One component; Representation of spatial augmentation features Height; This represents the result of vertical self-attention calculation; Indicates the first vertical direction softmax normalization of each element; Indicates vertical compression features The query vector of the first One component; Indicates the vertical position code number One component; express transpose; Indicates vertical compression features The key vector of the first One component; Indicates the vertical position code number One component; Indicates vertical compression features The value vector of the first One component; Representation of spatial augmentation features Width; for (32×256×512) horizontal compression obtained (32×256×1), obtained by vertical compression (32×1×512), after injecting the positional encoding, calculate the self-attention according to formulas (4)-(5), and fuse to obtain The positioning error for densely arranged photovoltaic panels (2m spacing) in a plain scene was reduced to 1.2 pixels.
[0030] In specific implementation, the S23 multi-level parameter fusion is performed according to formulas (6) to (7). Since there is no low light / high occlusion / dynamic interference in the sunny scene, only the dynamic scene branch is enabled. The bounding box optimization of the equivalent single convolution-BN layer after fusion is performed according to formulas (8) to (17). Wherein, IoU represents the intersection-union ratio between the predicted box and the real box. in, This represents the horizontal distance loss component; Represents the prediction box With real frame The absolute value of the difference in the x-coordinates of the center point; Indicates the width of the enclosing rectangle; express The square of; in, This represents the vertical distance loss component; Represents the prediction box With real frame The absolute value of the difference in the y-coordinates of the center points; Indicates the height of the enclosing rectangle; express The square of; in, Represents shape loss; exp represents the natural exponential function; This indicates the shape loss component in the width direction; This indicates the shape loss component in the height direction; This represents the preset shape loss coefficient (value 4). This indicates that the result within the parentheses is taken. Power of; in, This indicates the shape loss component in the width direction; Represents the prediction box With real frame The absolute value of the width difference; This represents the maximum width between the predicted bounding box and the ground truth bounding box; in, This indicates the shape loss component in the height direction; Represents the prediction box With real frame The absolute value of the height difference; This represents the maximum height between the predicted bounding box and the ground truth bounding box; Actual measurement , ( ), , ( ), , ( ),final Output (1200 tiles to be laid) and ( (High priority, auxiliary image low priority), this output serves as the input to S3, triggering S3's dynamic dual-link transmission start action.
[0031] In one feasible implementation, S3 is executed according to steps S31-S33, with link adaptive updates: S31 Link Quality Monitoring Unit every Collect parameters and perform actual measurements. , , According to the formula (18) Calculate the score; Wherein, Score represents the link stability score; , , Indicates the preset weighting coefficient ( , , ),and ; Indicates the real-time bandwidth of the link (unit: Mbps); Indicates the real-time signal strength of the link (unit: dBm); Indicates the real-time transmission delay of the link (unit: ms); Represents the reciprocal of the transmission delay; The calculation yields: (After normalization), that is ; S32, Threshold Initialization , times / second ,Every statistics times / second According to the formula (20) Calculate ; in, Indicates the first Link switching threshold for each cycle; Indicates the first Link switching threshold for each cycle; Indicates the threshold adjustment step size; S33, Select 5G link. According to the UDP protocol (port) )by Transmission (breakpoint resume: 16-bit sequence number + 32-bit checksum), auxiliary images are transmitted via TCP protocol (port). )by Compressed transmission, transmission completed (including) (with auxiliary images) arrive at the calibration end.
[0032] In this embodiment, S4 is executed according to steps S41-S43: S41 PMD noise classification, According to the formula (21) Calculate S42PLC calibration right of , No need to update tags; S43 terrain matching plain segment threshold. According to the formula (24) Calculate ;in, This indicates the percentage deviation between the quantitative data and the manually reported data; This represents the photovoltaic panel deployment quantization data output by multimodal recognition; This represents manually reported photovoltaic panel installation data; This represents the absolute value of the difference between the two; actual measurement. piece, No secondary calibration, output (1200 tiles to be laid) and , .
[0033] Specifically, S5 is executed according to steps S51-S52: Update the weights (momentum) of the C3 module of the feature encoder. ), Update the axial attention weights, and then update the model. Stored at the ground control station; S6 records a loop log (timestamp 2024-06-01 10:00, S1 execution 120s, S2 execution 8s, S3 execution 5s, S4 execution 3s, S5 execution 2s, output data volume 200MB), stores it in the MySQL database, and completes one loop.
[0034] Example 3 This embodiment focuses on a cloudy day scenario at the Ulanqab Mountain Photovoltaic Power Plant in Inner Mongolia.
[0035] In this embodiment, the timer of S6 is set to... Triggered every minute, the generated signal is sent to the drone and transmission node, initiating S1: the drone's flight parameters are adjusted to... , (Altitude gained from mountainous terrain), horizontal flight speed (Reduce speed to ensure image clarity), each Capture; Light sensor per Light intensity collected (measured at 380 lux) (Triggered low-light branch), occlusion sensor detects distance change. (Mountain shadows cause obstruction) (Triggering high occlusion branch), GPS records latitude and longitude; after collection, the multimodal raw dataset (RGB is dark due to cloudy weather, IR is clear, environmental parameters: illumination, occlusion, mountainous terrain) is stored to SD card and transmitted to ground control station.
[0036] Specifically, in the S21 dual-mode encoder of S2, due to the weak misalignment between RGB and IR, and the sensor viewing angle deviation caused by the mountainous terrain, the invariant feature encoder extracts... (Characteristics of photovoltaic panel terminals), according to the formula (1) Calculate spatial offset Pixels (increased in sunny scenes); where the variables in formula (1) are explained in the same way as in Example 2; at this time, the modal separation coding effectively avoids the modal difference between the dark part of RGB and the bright part of IR. Improved feature retention rate of photovoltaic panels; The S22 axial attention module is executed according to formulas (2) to (5); wherein, the variables in formulas (2)-(5) are interpreted the same as in Example 2; for (32×256×512) horizontal compression obtained (32×256×1), obtained by vertical compression (32×1×512), after injecting the positional encoding, calculate the self-attention according to formulas (4)-(5), and fuse to obtain .
[0037] In the preferred scheme, S23 multi-level parameter fusion utilizes two branches: low illumination and high occlusion. A virtual 3×3 convolutional layer is constructed during the alignment stage, and single-level fusion is performed according to the formula. (6)-Formula (7) Calculation; where the variables in formulas (6)-(7) are interpreted the same as in Example 1; Here (Compensate for low illumination feature intensity), bounding box optimization is performed according to formulas (8) to (17); The variables in formulas (8)-(17) are interpreted the same as in Example 2; Due to obstruction , ( ), , ( ), , ( ), Output (850 tiles to be laid) and .
[0038] In this embodiment, the S31 link quality monitoring of S3 shows signal attenuation in cloudy mountainous terrain, as measured in actual measurements. , , According to the formula (18) Calculate the score; The variables in formula (18) are interpreted the same as in Example 2; The calculation yields: (After normalization) S32 Statistics times / second> According to the formula (20) Calculate ; The variables in formula (20) are interpreted the same as in Example 2; S33 switches to the LoRa link, according to the LoRaWAN protocol (spreading factor). ,bandwidth )transmission ,Every Detect the score, 3 minutes later No additional auxiliary images will be uploaded at this time.
[0039] In practical implementation, the S41PMD noise classification of S4 is as follows: According to the formula (21) Calculate The measured noise rate was 8.2% (cloudy weather caused a decrease in recognition confidence). The variables in formula (21) are interpreted the same as in Example 1; S42PLC Calibration: After Warm-up of , For 210 noise samples, according to the formula (23) Update labels; where the variables in formula (23) are explained in the same way as in Example 1; after 5 iterations, the noise rate drops to 2.3%; S43 terrain matching mountain section threshold Manual reporting Blocks, according to the formula (24) Calculate ; In formula (24), the variables are interpreted the same as in Example 2; thus, we obtain... Output (850 blocks) and optimization parameters are used as inputs to S5.
[0040] In one feasible approach, S5 updates the model: Increase the IR mode feature weight to 0.6 (to compensate for RGB dark areas). Increase the vertical attention weight to 0.55 (to adapt to mountain slopes), and the updated model improves mAP in mountainous cloudy scenes; S6 records logs (timestamp 2024-06-01 10:30, S1 execution 150s, S2 execution 10s, S3 execution 8s, S4 execution 12s, S5 execution 3s, output data volume 120MB), and completes the loop.
[0041] Example 4 This embodiment focuses on a dynamic scenario of strong winds at the Bayannur photovoltaic power plant in Inner Mongolia, where the terrain is a gentle slope.
[0042] In this embodiment, after the timer of S6 is triggered, the UAV S1 adjusts its flight parameters due to strong winds: , Horizontal flight speed (Reduce speed to resist wind), each Capture image (increase the interval to avoid blur); the light sensor collects 5500 lux of light, the occlusion sensor detects construction workers passing by causing occlusion, and GPS records location changes. After the data collection is completed, the multimodal raw dataset (RGB is slightly blurred due to wind, while IR is stable) is transmitted to the ground control station.
[0043] Specifically, in the S21 dual-mode encoder of S2, dynamic interference causes RGB-IR misalignment. Pixel-invariant feature encoder extraction (Photovoltaic panel metal frame), according to the formula (1) When calculating the offset, channel attention is used. Suppressing wind fuzz noise, The signal-to-noise ratio has been improved to 30dB; In this case, the variables in formula (1) are interpreted in the same way as in Example 2; The S22 axial attention module is executed according to formulas (2) to (5); The variables in formulas (2)-(5) are interpreted the same as in Example 2; Horizontal compression captures the movement trajectory of construction workers (eliminating dynamic interference), while vertical compression focuses on the fixed arrangement of photovoltaic panels.
[0044] In the preferred scheme, S23 multi-level parameter fusion enables dynamic scene branching ( Construction workers' movement speed 1.5m / s > 1m / s: During the alignment phase, a virtual 3×3 convolutional layer is constructed (convolutional kernel index). The value is 1.2, enhancing dynamic feature response), and single-level fusion is performed according to the formula. (6)-Formula (7) Execution; The variables in formulas (6)-(7) are interpreted the same as in Example 1; Here , After multi-level fusion, the number of parameters is reduced by 52.1%, the inference speed remains at 34 FPS, and the bounding box optimization follows the formula. (8)-Formula (17) Execute; The variables in formulas (8)-(17) are interpreted the same as in Example 2; Dynamic interference caused , ( ), , ( ), , ( ), Output (Number of tiles to be laid) and .
[0045] In this embodiment, the S31 link quality monitoring of S3: strong winds caused signal fluctuations, as measured... , , According to the formula (18) Calculate the score; In formula (18), the variables are interpreted the same as in Example 2; the calculation yields: (After normalization) S32 Statistics times / second> According to the formula (20) Calculate ; The variables in formula (20) are interpreted the same as in Example 2; S33 switches to LoRa link transmission ,Every Detect the score, 10 minutes later Supplementing auxiliary images (compression ratio) (to speed up transmission).
[0046] In practical implementation, the S41PMD noise of S4 is classified as: caused by dynamic interference. According to the formula (21) Calculate The measured noise level was 6.8% (caused by interference from construction workers). The variables in formula (21) are interpreted the same as in Example 1; S42PLC Calibration: After Warm-up of , The noise samples caused by 180 dynamic interferences were analyzed according to the formula. (23) Update tags; The variable interpretation in formula (23) is the same as in Example 1; S43, Threshold for Terrain Matching Slope Section Manual reporting Blocks, according to the formula (24) Calculate The variables in formula (24) are interpreted the same as in Example 2. have to Output (980 blocks) and optimization parameters.
[0047] In one feasible approach, S5 updates the model: Enhance the convolutional kernel response of dynamic scene branches (increase the suppression weights for moving targets). Increasing the horizontal attention update frequency (from 1 time / frame to 2 times / frame) improves the model's mAP in dynamic scenes after the update; S6 records logs (timestamp 2024-06-01 11:00, S1 executes for 180s, S2 executes for 11s, S3 executes for 10s, S4 executes for 15s, S5 executes for 4s, output data volume 180MB), and completes the loop.
[0048] Example 5 This embodiment is designed for nighttime and dusty weather scenarios at the Baotou photovoltaic power plant in Inner Mongolia. In this embodiment, S6 follows a preset cycle. The timer is triggered every minute, and the generated trigger signal is sent to the DJI M350RTK drone (equipped with an enhanced IR sensor) and the Level 5 transmission node via "signal end-to-end distribution," initiating the S1 drone data acquisition step. At this time, the trigger signal of S6 is the only condition for S1 to start, and it needs to be adapted to the special environmental parameters of nighttime sandstorms.
[0049] Specifically, S1 is executed according to steps S11-S12: the drone flight parameters are optimized for nighttime sandstorms, and the flight altitude is... , Horizontal flight speed ,Every Capture an RGB-IR image once (increase the interval to reduce sand and dust trails); nighttime light sensor captures an RGB-IR image every time. Light intensity collected (measured at 85 lux) (Triggered low-light branch), the shading sensor detects changes in the distance between dust particles and the photovoltaic panel via laser ranging. (Dust caused dynamic shading) (Triggered high obstruction branch), GPS module activates anti-interference mode (dust affects signal, location changes) The latitude and longitude were recorded, ranging from E109° to 110° and N40° to 41°. After acquisition, the multimodal raw dataset (RGB images appear grayish-brown and blurred due to nighttime dust, while IR images clearly capture the temperature differences of the photovoltaic panels) was stored on a 128GB industrial-grade SD card (read and write speeds are fast). The signal is transmitted to the ground control station buffer unit via an enhanced 5G module (anti-dust signal attenuation).
[0050] In the preferred scheme, S2 is executed according to steps S21-S23. S21 uses the dual-modal encoder OAFA method, which adjusts modal weights for low-light nighttime scenes. The specific feature encoder uses a "low-light enhancement convolutional layer" (3×3 convolutional kernel, stride 1, padding 2) to extract texture features (weakening of the photovoltaic panel border under sand and dust blur) for RGB images, and a "high-sensitivity temperature convolutional layer" for IR images. The invariant feature encoder extracts cross-modal common features. (Common characteristics of temperature and shape of photovoltaic panel metal brackets), due to dust, there is a significant misalignment between RGB and IR (approximately 5.6 pixels), according to the formula (1) Calculate spatial offset Pixel; In the spatial migration prediction stage, for Performing 3×3 pooling yields and Spatial attention weights are obtained by 7×7 convolution and Sigmoid. (Weight reduced by 0.3 for dusty areas), obtained according to Hadamard volume. Then through the channel attention (IR channel weight 0.7) Suppresses RGB blur noise. The signal-to-noise ratio has been improved to 28dB.
[0051] In this embodiment, the S22 axial attention module's Sea-Attention mechanism optimizes feature focusing for sand and dust blurring, according to the formula... (2) To (40×256×512) Perform horizontal compression, focusing on capturing the temperature consistency of the horizontal arrangement of photovoltaic panels (excluding horizontal ambiguity interference from sand and dust); according to the formula (3) Perform vertical compression to focus on the shape characteristics of the longitudinal support of the photovoltaic panel (the longitudinal shading of sand and dust has a relatively small impact). Injection location encoding , (Horizontal encoding weight 0.4) , After applying a vertical coding weight of 0.6, follow the formula. (4) Calculate the level of self-attention according to the formula. (5) Calculate vertical self-attention; Fusion Subsequently, the recall rate for identifying photovoltaic panels that were half-obscured by sand and dust was effectively improved, which is an improvement over traditional CNNs and effectively solves the problem of missed detection of small targets caused by sand and dust blurring.
[0052] In practice, a dual-branch approach of "low light and high occlusion" is adopted to address nighttime sandstorms: a virtual 3×3 convolutional layer (convolutional kernel index) is constructed during the alignment stage. The value is 1.3, enhancing the IR characteristic response) and the virtual batch normalization layer ( , (IR branch) / 1.0 (RGB branch) , (Compensation for low-light feature shift); Single-level fusion according to formula (6) and formula (7) Execution; in, Indicates the fused first Each convolutional kernel parameter; Indicates the first Scaling parameters for each batch of normalized layers (IR branch) (enhancing temperature characteristics) Indicates the original number One convolutional kernel parameter (IR branch parameter accounts for 0.7); Indicates the first The moving weighted variance of each batch of standardized layers ( (Adapting to characteristic fluctuations under sandstorms). To avoid the denominator being 0; in, Indicates the fused first Each convolutional layer bias parameter; Indicates the first Scaling parameters for each batch of normalized layers; Indicates the first Moving weighted mean of each batch of standardized layers; Indicates the first Moving weighted variance of each batch of standardized layers; Indicates the first Offset parameters of each batch of normalized layers ( (Enhance low-light feature intensity) After multi-level fusion, the number of parameters is reduced, while the inference speed remains at 32 FPS (meeting the real-time monitoring requirements under nighttime sandstorm conditions); bounding box optimization is performed according to the formula. (8)-Formula (17) Execution: Dust caused an increase in the misalignment between the predicted bounding box and the actual bounding box. According to the formula (9) Calculate ; Where IoU represents the intersection-union ratio between the predicted bounding box and the ground truth bounding box; according to the formula (10) Calculate ,in (Dust caused an increase in the angle deviation of the line connecting the center points). ( Pixels Pixels (pixels) According to the formula (12) Calculate ,in ( (pixels) ( (pixels) According to the formula (15) Calculate ,in ( (pixels) ( (pixels) ; final Output and ,in, The number of tiles laid is 720. High priority, auxiliary images are marked as "low priority - delayed transmission" due to sand and dust blurring.
[0053] In one feasible implementation, S3 is executed according to steps S31-S33, with a link state adaptive update mechanism: S31 link quality monitoring unit every Acquisition parameters (dust caused increased signal fluctuations, extending the monitoring interval), measured link bandwidth Signal strength (32dBm lower than on a sunny day), transmission delay (Dust causes increased signal reflection delay), according to the formula (18) Calculate the score; Wherein, Score represents the link stability score; , , (Weights remain unchanged); (After normalization) ); (After normalization) ); (After normalization) ); Calculated (After normalization) S32 Threshold Initialization , times / second ,Every Statistical link update rate According to the formula (20) Calculate (Lowering the threshold to adapt to frequent link fluctuations); among which, Indicates the first Link switching threshold for each cycle; Indicates the first The threshold for each cycle (0.5); Indicates the threshold adjustment step size; This indicates the update rate caused by frequent changes in link parameters under dust conditions; S33, Switch to LoRa Link Transmission: Transmission Only (720 units), using the LoRaWAN protocol (spreading factor) ,bandwidth (Increase bandwidth to improve anti-interference)), each The dust storm subsided after 25 minutes by adjusting the score (extending the detection interval to reduce energy consumption). We will not upload supplementary images for now, but will wait until the sandstorm dissipates the next day. Supplementary transmission will be provided later.
[0054] In this embodiment, S4 is executed according to steps S41-S43. The PLC algorithm addresses the high noise caused by nighttime sandstorms: S41 PMD noise classification, nighttime sandstorms improve the confidence level of photovoltaic panel recognition. According to the formula (21) Calculate The measured noise level was 4.2% (mainly due to false identification caused by sand and dust); among which, This represents the probability that a true label of 1 (photovoltaic panels have been installed) is mistakenly labeled as 0. , The preset noise figure; Confidence level for identification under nighttime sandstorm conditions; This represents 0.32 raised to the power of 1.5. S42, PLC calibration: Warm-up stage, batch size Learning rate (Reduce learning rate to avoid overfitting noise), SGD optimizer (momentum) Weight decay )train Warm-up is stopped when the mAP on the validation set reaches 79.5% (extending the number of training rounds to adapt to complex noise), resulting in a preliminary model. Initial threshold during calibration phase (Increase the threshold to filter more reliable samples), calculate ,right Furthermore, for the 198 samples where the predicted labels did not match the manually labeled labels, the formula was used... (23) Update the labels, prioritizing the retention of IR feature prediction results; among which, This indicates the updated tag; For indicator functions ( Output 1 if the condition is met, otherwise output 0. for The softmax output (IR feature contribution ratio 0.7); each cumulative After round of training, according to Update the threshold, iterate 6 times, and then continuously. With unlabeled updates, the noise rate decreased from 4.2% to 1.8%. S43, Terrain Matching (This area has a gentle slope) ), matching the threshold of the gentle slope section Manual night patrol reporting Blocks, according to the formula (24) Calculate ;in, Indicates percentage deviation; Block (model output); Block (manual inspection value); block; get No secondary calibration, output (720 blocks, 58%) and model optimization parameters ( IR modality convolution kernel weights, Vertical attention weights).
[0055] Specifically, S5 performs model iterations according to steps S51-S52: Update the dual-modal encoder, increase the weight of the C3 module of the IR specific feature encoder to 0.75 (to enhance the nighttime temperature feature), and adjust the SiLU activation layer parameters of the invariant feature encoder; The Sea-Attention module was updated, with the vertical attention weight increased from 0.6 to 0.7 and the horizontal attention update frequency increased from 1 time / frame to 1.5 times / frame (to exclude dynamic interference from sand and dust in real time). After the update, the model's mAP in nighttime sand and dust scenes improved to 82.3%, an improvement of 2.8% compared to before optimization (79.5%). S6 records a loop log (timestamp 2024-06-02 02:30, S1 execution 210s, S2 execution 15s, S3 execution 12s, S4 execution 20s, S5 execution 5s, output data volume 95MB), stores it in a MySQL database, and completes one monitoring loop of a nighttime sandstorm scene.
[0056] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0057] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.
Claims
1. A collaborative optimization method for intelligent monitoring of photovoltaic construction progress, characterized in that, Includes the following steps: S1. Periodically collect multimodal data of the photovoltaic construction area using drones and transmit the multimodal data to the ground control station; S2. Trigger multimodal recognition based on multimodal data, extract features through a dual-modal encoder, calculate spatial offset, enhance features through an axial attention module, perform multi-level parameter fusion and bounding box optimization, and output photovoltaic panel laying quantification data and data priority labels; S3. When the photovoltaic panel's quantification data and data priority tags reach the transmission node, dynamic dual-link transmission is triggered. The transmission link is dynamically selected through link quality monitoring and link switching thresholds to transmit the data to the calibration end. S4. When the transmission data output in step S3 arrives at the calibration end, PLC noise calibration is triggered. Through PMD noise classification, progressive label calibration and standard deviation threshold matching, the calibrated photovoltaic construction progress data and model optimization parameters are output. S5. Trigger iterative updates of model parameters based on model optimization parameters, optimize the parameters of the dual-modal encoder and axial attention module, and store the updated model; S6. Repeat steps S1 to S5, generate periodic trigger signals through a timer, distribute the signals to periodically trigger step S1, and record a cycle log to periodically monitor the progress of photovoltaic construction.
2. The collaborative optimization method for intelligent monitoring of photovoltaic construction progress according to claim 1, characterized in that, The UAV data acquisition step in step S1 includes the following steps: S11. Control the drone to perform periodic data collection tasks over the photovoltaic construction area; the drone's flight parameters are preset according to the photovoltaic power station equipment layout, and the flight altitude is set to the distance from the photovoltaic panels. - Horizontal flight speed ,Every RGB and IR images are captured at intervals, where... , To preset the height threshold, , To preset the shooting distance interval, The preset maximum flight speed in a windless environment; When the drone detects that the data collection interval has reached Light intensity is collected when the time is set, where Preset the light collection interval; When the drone detects a change in distance to the photovoltaic panel exceeding [a certain value] using laser ranging... When this occurs, the occlusion ratio calculation is triggered, where, The preset threshold for changes in occlusion distance; When the detected change in the drone's position exceeds The location latitude and longitude are recorded at the time of triggering, among which, Preset GPS location change threshold; S12. Data storage is triggered after the multimodal raw dataset is generated; data transmission is triggered when data storage is completed and a wireless connection signal from the ground control station is detected.
3. The collaborative optimization method for intelligent monitoring of photovoltaic construction progress according to claim 1, characterized in that, The multimodal adaptive recognition step in step S2 includes the following steps: S21. Construct a dual-modal encoder, including a specific feature encoder and an invariant feature encoder. When the specific feature encoder detects an IR image input, it triggers IR-specific feature extraction: capturing the temperature distribution differences of the photovoltaic panel through the same network structure and outputting IR-specific features; when the invariant feature encoder detects simultaneous input of RGB and IR images, it triggers cross-modal common feature extraction, extracting cross-modal common features of the photovoltaic panel frame and wiring. When cross-modal shared features When input is sent to the spatial augmentation module, spatial augmentation processing is triggered to obtain spatial augmentation features. When spatial enhancement features When the input is sent to the channel enhancement module, channel enhancement processing is triggered to obtain the offset prediction features. Based on offset prediction features The enhanced features after offset prediction are calculated using the formula. ; S22. Embed a Sea-Attention axial attention module in the CSP layer of the backbone network; when the Sea-Attention axial attention module calculates the activation signal, it triggers feature compression, which enhances the features after the offset prediction. Horizontal compression is triggered along the horizontal direction, and the horizontal compression feature is output. Enhanced features after offset prediction Trigger vertical compression along the vertical direction and output vertical compression features. ; Based on horizontal compression characteristics and vertical compression features Trigger positional encoding injection to generate learnable positional codes: , and , Once the position code is generated, the code overlay is triggered: and query vector Add, and key vector Add, and query vector Add, and key vector Add them together, preserving the global spatial relationships; where, This represents the positional encoding of the horizontal query vector. The positional encoding of the horizontal key vector; This represents the positional encoding of the query vector in the vertical direction; This represents the positional encoding of the vertical key vector; After the encoding is superimposed, a horizontal self-attention action is triggered along the horizontal direction, and the horizontal self-attention calculation result is obtained. Trigger vertical self-attention actions along the vertical direction, and obtain the vertical self-attention calculation result. ;Calculate the results of horizontal self-attention Comparison with vertical self-attention calculation results Adding them together yields the global attention enhancement feature. ; S23, Enhance global attention features Input is fed into three branch networks: low light intensity, high shading, and dynamic scene. This triggers a multi-level parameter fusion initiation action. Based on the three branch networks, multi-level parameter fusion is performed, and the resulting photovoltaic panel deployment quantification data is output. and data priority labels .
4. The collaborative optimization method for intelligent monitoring of photovoltaic construction progress according to claim 3, characterized in that, The specific steps of step S23 are as follows: S231, When the branch network initialization is complete and After input, for the identity layer of each branch, a virtual layer construction action is triggered to construct the virtual layer. Convolutional layers and virtual batch normalized layers; For each branch's batch normalized layer, the virtual convolutional layer construction action is triggered; among them, Preset convolution kernel size; S232. Trigger single-level parameter fusion action based on the convolutional layer and batch normalized layer of each branch, and calculate the biased convolutional layer parameters of the three branches after single-level fusion is completed. S233. Trigger a multi-level fusion action based on the parameters of the biased convolutional layer of the three branches after single-level fusion is completed. Add the parameters of the convolutional kernels after the three branches are fused to obtain an equivalent single convolution-BN layer. Use the SSIoU loss function to optimize the photovoltaic panel bounding box positioning. Once the bounding box optimization is complete, the recognition result output action is triggered, outputting the quantitative data of photovoltaic panel laying. and data priority labels .
5. The collaborative optimization method for intelligent monitoring of photovoltaic construction progress according to claim 1, characterized in that, Step S3, the dynamic dual-link transmission step, includes the following steps: S31. Deploy link quality monitoring units at the Level 5 transmission nodes to trigger periodic monitoring actions. Intermittently trigger link parameter collection actions, where, For the preset monitoring interval, the link parameters include link bandwidth. Signal strength and transmission delay ; After each link parameter collection is completed, the link stability score is calculated based on the collected link parameters. S32. When the link quality monitoring unit is started, it triggers the threshold parameter initialization action to obtain the initialization parameters, including the initialization link switching threshold. Preset target update rate Preset threshold adjustment step size When the system clock detects that the time interval has reached The update rate statistics action is triggered on time to count the link update rate. ;in, Set the preset update rate statistical interval; Based on link update rate Update rate with preset target Trigger threshold adaptive adjustment action: If ,but ;like ,but ; in, Indicates the first Link switching threshold for each cycle; Indicates the first Link switching threshold for each cycle; Indicates the threshold adjustment step size; S33, Preset link quality switching threshold Based on the link stability score, the link adaptive selection action is triggered: if Choose a 5G link; if Switch to LoRa link; Among them, when the link is selected as a 5G link and the photovoltaic panel deployment quantification data Once loading is complete, a high-priority data transmission action is triggered. When high-priority transmission is initiated and the auxiliary image is loaded, a low-priority data transmission action is triggered. When the link is selected as a LoRa link and the photovoltaic panel deployment quantization data is available Once loading is complete, the core data priority transmission action is triggered; after the LoRa link transmission starts, each Link stability score is calculated once at intervals: If This triggers the auxiliary image retransmission action, retransmitting the auxiliary image; among which, A preset retransmission judgment interval is used to assist in image retransmission; when the transmission is completed and the photovoltaic panel laying quantization data is available... When the auxiliary image arrives at the calibration end, complete data is obtained. , will complete data Send to the calibration end.
6. The collaborative optimization method for intelligent monitoring of photovoltaic construction progress according to claim 1, characterized in that, Step S4, the PLC noise calibration step, includes the following steps: S41. Classify the noise from photovoltaic construction sites as PMD noise, using the formula... and Constrain the noise range; where, This represents the probability that a true label of 1 is mislabeled as 0. c1 and c2 represent the probability that a true label of 0 is mislabeled as 1; c1 and c2 represent preset noise coefficients, taken from historical noise statistics. Indicates the confidence level of photovoltaic panel identification; S42. Trigger the Warm-up training start action based on the multimodal recognition in step S2 to obtain the preliminary model. Set the initial calibration threshold. Calculate the calibration threshold When the calibration threshold is reached Generate and complete data After the image samples are loaded, the confidence screening and calibration action is triggered: for each image sample of the transmitted photovoltaic panel with quantized data, based on the preliminary model... Calculate the softmax and output the recognition confidence score. ,like And the confidence level of identification If the predicted label differs from the manually labeled label, a dynamic label update action is triggered. When the label update is complete and the cumulative number of training rounds reaches [a certain threshold], When triggered Updated to Synchronous updates Repeat label calibration until continuous. The cycle of labelless updates yields a calibration label set. ; in, This indicates the interval between iterations for the preset threshold. Indicates the first The threshold of the wheel; Indicates the step size; Indicates the step size of the preset threshold iteration; This indicates the number of consecutive iterations without updates that are preset to stop the iteration; when consecutive iterations are detected... A stop check is triggered when the wheel label update count reaches 0. S43. Trigger terrain adaptive matching action based on terrain type parameters collected in S1: If Matching mountain landmark threshold ;like Matching plain section threshold ;like Matching the threshold of the gentle slope section ; in, Indicates the elevation difference of the photovoltaic construction area; This represents the preset mountain elevation difference threshold for mountainous terrain. This represents the elevation difference threshold for plain terrain, and ; This indicates the deviation threshold for mountain sections; This represents the deviation threshold for the plain section, and ; This indicates the deviation threshold for the gentle slope section, and ; Once the terrain is determined and the corresponding threshold is matched, the manually reported photovoltaic panel installation data is retrieved. Calculate quantitative data for photovoltaic panel installation. Compared with manually reported photovoltaic panel installation data percentage deviation ; When percentage deviation Calculation completed and percentage deviation When the matching threshold is exceeded, a second precise calibration action is triggered, and step S42 is repeated. Once the deviation calculation is complete and no secondary calibration is required, or after secondary calibration is completed, output the calibrated photovoltaic construction progress data. And model optimization parameters, including the dual-modal encoder convolution kernel parameters. and axial attention weight .
7. The collaborative optimization method for intelligent monitoring of photovoltaic construction progress according to claim 6, characterized in that, Step S5, the model iteration step, includes the following steps: S51, Convert the convolution kernel parameters of the dual-modal encoder. The input to the dual-modal encoder S21 updates the convolution kernel weights of the feature-specific encoder and the SiLU activation layer parameters of the invariant feature encoder. S52, Adjust axial attention weight Input the Sea-Attention module of S22 to update the weight coefficients of horizontal and vertical self-attention. After the parameters are updated, use the multimodal recognition model with updated parameters. Stored at the ground control station.
8. The collaborative optimization method for intelligent monitoring of photovoltaic construction progress according to claim 1, characterized in that, Step S6, which is executed cyclically, includes the following steps: S61. Deploy a timer at the ground control station. When the timer detects that the system clock has reached its maximum speed... At intervals, a periodic trigger signal generation action is initiated; among which, Indicates the period of the loop execution; S62. After the trigger signal is generated, the trigger signal is sent to the UAV and each transmission node, and the UAV data acquisition step S1 is started. Steps S1 to S5 are executed in sequence. After each loop is completed, the loop log is recorded and stored in the database.
9. The collaborative optimization method for intelligent monitoring of photovoltaic construction progress according to claim 3, characterized in that, In step S21, after the convolutional layer of the invariant feature encoder outputs features, the SiLU activation calculation is triggered: the formula for the SiLU activation function is as follows. ,in ; in, This represents the output value of the SiLU activation function; This represents the input value of the activation function; This represents the Sigmoid activation function; Represents the natural exponential function; when Inputting the data triggers exponent calculation, yielding the exponent result. Once the exponent result is generated, division calculation is triggered. When cross-modal shared features Once extraction is complete, spatial offset prediction is initiated, which is executed in two steps: S211, Common features across modalities Perform max pooling and average pooling to obtain the max pooling result. and average pooling results ; Max pooling results and average pooling results Spatial attention weights are obtained by splicing and activating with a Sigmoid function. Spatial attention weights Common features across modalities Performing the Hadamard product operation yields spatially enhanced features. ; S212, Spatial Enhancement Features Perform global average pooling and global max pooling to obtain the global average pooling result. and global max pooling results ; The result of global average pooling and global max pooling results Input shared multilayer perceptron, output channel attention weights Channel attention weights Spatial Enhancement Features Element-wise multiplication yields the offset prediction features. Suppress channel noise.
10. The collaborative optimization method for intelligent monitoring of photovoltaic construction progress according to claim 3, characterized in that, Based on the horizontal compression characteristics in S22 and vertical compression features The specific steps for triggering the location encoding generation action are as follows: based on the horizontal compression features The dimension generates a randomly initialized parameter matrix. After the matrix initialization is complete, backpropagation is triggered, and the model is updated iteratively through backpropagation. After the model loss is calculated, parameter adjustment is triggered to adjust the encoded query vector. It can capture the spatial relationship of photovoltaic panels; among them, This represents the preset initialization range threshold for the position encoding parameters; Once the self-attention query and key vector are summed, the softmax calculation is triggered.
Citation Information
Patent Citations
A UAV inspection method and system based on building construction
CN116301055B