An Unmanned Aerial Vehicle-Based Environmental Monitoring Method and Monitoring System
The environmental monitoring images are obtained through drones and the image generation and optimization are generated and optimized by using environmental monitoring and labeling neural networks. The problems of low environmental monitoring efficiency and insufficient accuracy in the prior art are solved, and efficient and accurate environmental monitoring and labeling are achieved.
Patent Information
- Application Number
- CN202510294533.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-03-13
AI Technical Summary
Existing environmental monitoring methods are inefficient and insufficient in processing high-spatial-spatial-resolution data. Traditional remote sensing data analysis relies on manual annotation to easily lead to hysteresis and subjective deviations. Automatic annotation systems have registration errors in multi-sensor data fusion, making it difficult to adapt to spectral variations of complex surface coverage types.
The drone-based environmental monitoring method is adopted, by obtaining the initial environmental monitoring image and target monitoring labeling information, the environmental monitoring labeling neural network is used for image generation and optimization, and combining the feature weight focusing component and the context cache matrix, the error index of image chunking is iteratively optimized to generate a monitoring labeling image that meets the labeling requirements.
It improves the efficiency and accuracy of environmental monitoring image annotation, reduces the dependence on pre-debug, and the generated monitoring and annotation images more meet the preset annotation requirements, reducing the need for manual intervention.
Smart Images

Figure CN119832458B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of machine learning, and more particularly, to an unmanned aerial vehicle (UAV)-based environmental monitoring method and monitoring system. Background Art
[0002] In the field of environmental monitoring, traditional remote sensing data analysis methods mainly rely on manual visual interpretation and semi-automated processing tools to identify specific ecological indicators and perform manual annotation through satellite or aerial images. However, such methods face significant efficiency bottlenecks when processing high spatio-temporal resolution data. Especially when dealing with sudden environmental pollution incidents or large-scale ecological monitoring tasks, the lag and subjective bias of manual annotation often lead to missed detection of key information. In the prior art, although the automated annotation system based on a rule engine improves the processing speed, its rigid feature extraction algorithm is difficult to adapt to the spectral variation of complex surface cover types, and registration errors are easily introduced during the multi-sensor data fusion process, limiting the application accuracy in scenarios such as vegetation dynamic monitoring and water pollution tracking. In recent years, deep learning technology has shown potential in the field of remote sensing image classification. Models such as convolutional neural networks (CNNs) are used to automatically extract ground object features and perform automatic annotation of images according to set rules based on user requirements. However, a large number of templates are required, and these templates are designed manually or a large number of training samples with annotation constraints are collected for model training. The effect of the model is greatly affected by the sample size. In other words, the existing methods are not efficient and accurate enough in monitoring and annotating images. Summary of the Invention
[0003] In view of this, embodiments of the present application provide an unmanned aerial vehicle (UAV)-based environmental monitoring method and monitoring system.
[0004] According to one aspect of the embodiments of the present application, an unmanned aerial vehicle (UAV)-based environmental monitoring method is provided, including:
[0005] Obtaining an initial environmental monitoring image in a target monitoring environment area captured by an unmanned aerial vehicle (UAV), and obtaining target monitoring annotation information, where the target monitoring annotation information is used to indicate the monitoring annotation requirements for the target monitoring environment area;
[0006] Loading the initial environmental monitoring image into a target environmental monitoring annotation neural network to generate an unoptimized image, and outputting an unoptimized monitoring annotation image;
[0007] Load the image block features corresponding to each image block in the unoptimized monitoring annotation image into the target generation result evaluation component respectively, determine the error metrics of the target generation result evaluation component respectively, and iterate the context cache matrix based on the error metrics, where the error metrics are used to represent the consistency of the image block corresponding to the image block feature with the target monitoring annotation information;
[0008] Load the initial environmental monitoring image into the target environmental monitoring annotation neural network for image generation with optimization, and output an optimized monitoring annotation image, where the optimized monitoring annotation image represents the environmental monitoring image output by the target environmental monitoring annotation neural network after iterating the context cache matrix based on the error metrics, and each image block in the optimized monitoring annotation image satisfies the target monitoring annotation information.
[0009] According to another aspect of the embodiments of the present application, a monitoring system is provided, including: one or more processors; and one or more memories, where computer-readable code is stored in the memories, and when the computer-readable code is run by the one or more processors, the one or more processors are caused to execute the method as described above.
[0010] The technical effects of this application at least include: This application obtains an initial environmental monitoring image and target monitoring annotation information in a target monitored environmental area, where the target monitoring annotation information is used to indicate the monitoring annotation requirements of the target monitored environmental area; loads the initial environmental monitoring image into a target environmental monitoring annotation neural network for unoptimized image generation, and outputs an unoptimized monitoring annotation image. Among them, the unoptimized monitoring annotation image includes image blocks that do not meet the target monitoring annotation information, and the unoptimized monitoring annotation image is constructed from the image blocks output by the target environmental monitoring annotation neural network. The target environmental monitoring annotation neural network is used to determine the output image blocks one by one according to the loaded image blocks and the context cache matrix based on the image block scanning order, and determine the corresponding monitoring annotation image based on the output image blocks. The target environmental monitoring annotation neural network includes multiple feature weight focusing components, and the context cache matrix includes feature mapping binary tuples corresponding to each feature weight focusing component in the multiple feature weight focusing components; loads the image block features corresponding to each image block in the unoptimized monitoring annotation image into the target generation result evaluation component respectively, determines the error metrics of the target generation result evaluation component respectively, and optimizes the context cache matrix based on the error metrics, where the error metrics are used to represent the consistency of the image blocks corresponding to the image block features meeting the target monitoring annotation information; loads the initial environmental monitoring image into the target environmental monitoring annotation neural network for optimized image generation, and outputs an optimized monitoring annotation image. The optimized monitoring annotation image represents the environmental monitoring image output by the target environmental monitoring annotation neural network after iterating the context cache matrix based on the error metrics. The way that each image block in the optimized monitoring annotation image meets the target monitoring annotation information does not require modifying and repeatedly debugging the pre-debugged environmental monitoring annotation neural network compared with directly debugging an environmental monitoring annotation neural network for an additional annotation method. In the content generation process of the environmental monitoring annotation neural network of this application, a generation result evaluation component is used to evaluate whether the generated image meets the preset annotation requirements, and the generation error of the generation result evaluation component is backpropagated to optimize the state of the environmental monitoring annotation neural network, so that the probability of meeting the annotation requirements is greater, thus guiding the pre-debugged environmental monitoring annotation neural network to generate a monitoring annotation image that meets the annotation requirements.
[0011] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the technical solution of this application. Brief Description of the Drawings
[0012] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.
[0013] Figure 1 It is a schematic structural diagram of an application scenario provided by the present application;
[0014] Figure 2 It is a schematic flowchart of an environment monitoring method based on an unmanned aerial vehicle provided by the present application;
[0015] Figure 3 It is a schematic structural diagram of a monitoring system provided by an embodiment of the present application. Detailed implementation manners
[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.
[0017] To facilitate a clearer understanding of the present application, first, the application scenario for implementing the media data processing method of the present application will be introduced. As Figure 1 shown, the application scenario includes a monitoring system 10 and a drone cluster. The drone cluster can include one or more drones, and the number of drones will not be limited here. As Figure 1 shown, the drone cluster can specifically include Drone 1, Drone 2,..., Drone n; it can be understood that Drone 1, Drone 2, Drone 3,..., Drone n can all be network-connected to the monitoring system 10, so that each drone can perform data interaction with the monitoring system 10 through the network connection.
[0018] It is understandable that the monitoring system 10 may refer to a device that executes the drone-based environmental monitoring method provided in the embodiments of the present application. The monitoring system 10 may also be used to store images or videos captured by the drone. Among them, the monitoring system 10 may be a server. For example, it may be an independent physical server, or a server cluster or distributed system composed of at least two physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The drone may specifically refer to an industrial drone, an agricultural drone, a special drone, etc., but is not limited thereto. Each drone and the monitoring system 10 may be directly or indirectly connected through wired or wireless communication methods. At the same time, the number of drones and the monitoring system 10 may be one or at least two, and the present application does not make any restrictions here.
[0019] Further, please refer to Figure 2 , which is a schematic flowchart of a drone-based environmental monitoring method provided in the embodiments of the present application. As Figure 2 shown, this method can be executed by the monitoring system 10 in Figure 1 . Among them, the drone-based environmental monitoring method may include the following steps:
[0020] Step S100: Obtain the initial environmental monitoring image in the target monitoring environment area captured by the drone, and obtain the target monitoring annotation information, where the target monitoring annotation information is used to indicate the monitoring annotation requirements for the target monitoring environment area.
[0021] In step S100, an initial environmental monitoring image of the target monitoring environment area is collected by a drone, and the target monitoring annotation information associated with this area is synchronously obtained. During this process, the parameters of the sensor module carried by the drone can be configured first, including adjusting the resolution, focal length, exposure time of the optical sensor, and the band range of the multispectral sensor, to ensure that the collected image data can cover the key monitoring elements of the target environment area. For example, if the target monitoring environment is an industrial pollution area, the drone may be configured to carry a high-resolution visible light sensor (resolution ≥ 4096×2160 pixels) and a short-wave infrared sensor (band range 1400 - 3000 nm) to capture surface visible features and thermal radiation features simultaneously. The acquisition of the initial environmental monitoring image needs to follow the principle of spatio-temporal consistency: the flight path of the drone is planned through the Geographic Information System (GIS) module to ensure that the spatial coverage of the aerial images meets the monitoring requirements, and a timestamp alignment mechanism is used to ensure the temporal coherence of the continuously captured frame sequence. During the image acquisition process, data preprocessing operations are performed in real time, including removing the influence of atmospheric scattering by the dark channel prior-based dehazing algorithm, enhancing the image contrast by adaptive histogram equalization, and image stitching based on feature matching to generate a regional panoramic view.
[0022] The acquisition of the target monitoring annotation information involves interaction with the annotation rule database, which stores the preset annotation specifications for different environmental monitoring tasks, such as the types of annotated regions, the styles used for annotation (such as colors, lines, patterns, annotation content), and so on. For example, in the forest fire monitoring scenario, the annotation information may include the infrared feature threshold of the fire area (such as the HSV color space range of the area with temperature ≥ 300℃ is defined as H ∈ [0,15], S ∈ [80,100], V ∈ [90,100]), the vector annotation rule for the smoke diffusion direction (using the Histogram of Oriented Gradients HOG descriptor), and the unit standard for the affected area. By parsing the structured description file of the annotation information (such as JSON or XML format), it is converted into an operable tensor form, for example, encoding the annotation style constraints as the conditional vector of the style transfer network , mapping the content detection requirements to the semantic segmentation template matrix . For dynamic monitoring scenarios, an annotation information update mechanism can also be established. For example, when an abnormal water body pH value is detected, the detection item of the heavy metal pollutant ion concentration in the annotation rule is automatically triggered and appended. At the data association level, the initial environmental monitoring image is bound to the geographical coordinates through a spatial registration algorithm, and an affine transformation matrix is used to implement the conversion from pixel coordinates to the UTM coordinate system, and an R-tree index structure is used to establish the correspondence between the image blocks and the annotation information. To address the challenges of data acquisition in complex environments, a multi-modal data fusion strategy can be implemented: fusing visible light images , infrared images and synthetic aperture radar images generate multi-channel input data through a weighted fusion algorithm ( , where ), and set the confidence weights of each modality data in the annotation information. In the quality control link, perform abnormal data detection, use a texture analysis algorithm based on local binary pattern (LBP) to identify the defective areas of the image, and trigger the UAV re-flight mechanism for data re-acquisition. The finally generated initial environmental monitoring image set and the target monitoring annotation information set will meet technical indicators such as spatial resolution δ ≤ 0.1m, number of spectral channels k ≥ 5, and time synchronization error Δt ≤ 50ms, and provide input data that meets the format requirements for subsequent neural network processing.
[0023] Step S200: Load the initial environmental monitoring image into the target environmental monitoring annotation neural network for unoptimized image generation, and output the unoptimized monitoring annotation image. Among them, the unoptimized monitoring annotation image includes image blocks that do not meet the target monitoring annotation information. The unoptimized monitoring annotation image is constructed from the image blocks output by the target environmental monitoring annotation neural network. The target environmental monitoring annotation neural network is used to determine the output image blocks one by one according to the loaded image blocks and the context cache matrix according to the image block scanning order, and determine the corresponding monitoring annotation image based on the output image blocks. The target environmental monitoring annotation neural network includes multiple feature weight focusing components, and the context cache matrix includes the feature mapping binary groups corresponding to each feature weight focusing component in the multiple feature weight focusing components.
[0024] In step S200, the obtained initial environmental monitoring image is input into the target environmental monitoring annotation neural network for unoptimized image generation, that is, the forward propagation of the complete network, to generate the unoptimized monitoring annotation image. For example, based on the Transformer architecture, the serialization processing of image blocks and context correlation modeling can be realized through feature weight focusing components. First, perform image block preprocessing: divide the input image into a block sequence with a size of p×p, where H and W are the height and width of the image respectively, C is the number of channels (for example, for an RGB image, C = 3), and the number of blocks . For a visible light image with a resolution of 2048×2048, if p = 16 is set, then 128×128 = 16384 image blocks are generated. Each block is converted into an embedding vector with d = 512 dimensions through a linear projection layer , and position encoding is added to retain spatial information, and its calculation follows the formula 。In the feature weight focusing component, a global dependency relationship between blocks is established through the multi-head attention mechanism. The calculation process of each attention head is , where the query matrix , the key matrix , and the value matrix are obtained by linear transformation from the input embedding. , where h is the number of attention heads. For example, when h = 8 is set, the dimension d of each head k = 64, and the attention score matrix represents the association strength between all image patches. The context cache matrix (l is the sequence length) remains fixed at this stage, and it stores the set of key-value pairs generated by each feature weight focusing component when processing the previous t - 1 image patches. For example, when processing the t-th image patch, the of the current patch is concatenated with the existing in the cache to form an extended key-value matrix for calculating the attention output of the current patch. The generation of the unoptimized monitored annotation image follows an autoregressive pattern: predicting image patches one by one in raster scan order (from left to right, top to bottom). The generation of the s-th patch B s depends on the previous s - 1 patches and their corresponding cache matrix . Specifically, when processing the s-th patch, the following operations are performed: (1) Input the embedding vector of the current patch B s into the Transformer encoder layer. After layer normalization ( , where is the mean and standard deviation of the features, and are learnable parameters), it enters the multi-head attention module; (2) Each attention head calculates the attention weights of the current patch and the historical patches in the cache in parallel to generate context-aware features , where represents the attention degree of the s-th patch to the j-th patch in the i-th attention head, and is the value vector of the j-th patch in the i-th attention head; (3) Concatenate the outputs of all heads and generate the predicted patch through a feed-forward neural network (FFN). The FFN usually contains two fully connected layers and the GELU activation function: , where is the weight matrix. During this process, the immutability of the context cache matrix is strictly maintained - that is, after each image patch is processed, its corresponding key-value pair is only added to the cache for subsequent patches to use, but its value is not updated through backpropagation of gradients. For example, in the scenario of river pollution monitoring, when processing the 5th image patch containing the pollution source, the attention weights are calculated using the cache information of the previous 4 patches (such as the color characteristics of the upstream water body, flow velocity vector, etc.). However, due to the unoptimized cache, it may lead to prediction deviations in the pollution diffusion range, such as mislabeling the actual polluted area R p as the normal area R n . The construction of the unoptimized monitoring annotation image I unop t is achieved through reverse chunking operation: the predicted patches are recombined into a complete image in the original scanning order. If the chunk size is 16×16, bilinear interpolation is used to eliminate the gaps between chunks during recombination, and the interpolation weight is calculated as , where ( )) are the coordinates of the nearest neighbor chunk center. Quality verification is performed synchronously at this stage. For example, the global consistency between and I is compared through the structural similarity index (SSIM): , where and represent the pixel means of the original image I and the unoptimized annotation image I unopt respectively, and represent the pixel value variances of the original image I and the unoptimized annotation image I unopt respectively, and c1 and c2 are constants to avoid a zero denominator. When the SSIM value is lower than the threshold, the chunk regeneration mechanism is triggered. To improve processing efficiency, a memory optimization strategy is adopted: for large-scale images exceeding 4096×4096 pixels, tiling is implemented, and the image is divided into overlapping sub-regions (the number of overlapping pixels = 32), which are respectively input into the neural network and then synthesized into a complete annotation image through weighted fusion (the fusion function , decays linearly with the distance from the edge). x and y represent the pixel values or feature vectors of adjacent sub-regions in the overlapping part, is the weight coefficient, controlling the fusion ratio of the two sub-regions in the overlapping area.
[0025] Step S300: Load the image patch features corresponding to each image patch in the unoptimized monitoring annotation image into the target generation result evaluation component respectively, determine the error metrics of the target generation result evaluation component respectively, and iterate the context cache matrix based on the error metrics, where the error metrics are used to represent the consistency of the image patch corresponding to the image patch feature with the target monitoring annotation information.
[0026] In step S300, the target generation result evaluation component performs error analysis on the unoptimized monitoring annotation image, and optimizes the context cache matrix based on the error index to improve the annotation quality. First, each image patch in the unoptimized monitoring annotation image (where p is the patch size and c is the number of feature channels) is input into the target generation result evaluation component, which consists of a fully connected layer , a non-linear activation layer (such as the GELU function , where is the cumulative distribution function of the standard normal distribution), and an output layer . Its forward propagation formula is , where represents the operation of converting the image patch B s from a multi-dimensional tensor to a one-dimensional vector. and are the biases of the fully connected layer, and the output dimension k corresponds to the number of annotation information categories (for example, in soil pollution monitoring, k = 5 can represent the concentration levels of heavy metals cadmium, lead, arsenic, mercury, and chromium). The error index is calculated using the contrastive loss function , where is the mean squared error index of the s-th patch, k is the total number of annotation categories, is the true annotation vector of the s-th block in the target monitoring annotation information. For example, when = 0.9, it means that the lead pollution concentration reaches the threshold of 90%. After calculating the loss value for each patch independently, the error gradient is calculated through the backpropagation algorithm, focusing on updating the parameters of the context cache matrix . The specific optimization process uses the projected gradient descent method: Let the current cache matrix be , the learning rate be , and the gradient of the error function with respect to the cache matrix be . Then the update formula is , where the projection operation constrains the matrix elements within the range [-1, 1] to prevent numerical overflow.
[0027] In a case of predicting the diffusion of river pollution, if the error index = 0.76 corresponding to the 15th image patch is significantly higher than the average error = 0.23, calculate the gradients of the key-value pairs of the 3rd attention head of the Transformer for this patch, for example, they are and , and propagate them layer by layer to the corresponding positions of the cache matrix for correction through the chain rule. For example, when there is a deviation in the pollution boundary annotation, the corrected key matrix It will increase the weight of the water flow direction feature, enabling more accurate prediction of the pollution diffusion range during subsequent block generation. To improve the optimization efficiency, a hierarchical optimization strategy is implemented: a smaller learning rate is adopted for the bottom-layer feature weight focusing component (the layer closer to the input) to retain the basic features, while a larger learning rate is used for the high-layer components to quickly adapt to the semantic-level annotation requirements. The dynamic learning rate adjustment follows the cosine annealing rule , where T is the total number of iterations and t is the current iteration index.
[0028] In an industrial waste gas emission monitoring scenario, if the annotation rule requires simultaneous detection of the SO2 concentration (continuous value regression task) and the plume shape (discrete classification task), a multi-task loss function will be constructed , where the regression loss , the classification loss , and the weight coefficient is determined through grid search on the validation set. During the optimization process, an asynchronous parameter update mechanism is adopted: when processing the s-th block, only the cache matrix entries before the position x corresponding to its scanning order are updated , while the subsequent entries are kept unchanged, which is achieved through a mask matrix , where if and only if and , ensuring that the temporal dependency relationship is not damaged. In the case of crop pest and disease monitoring, if the annotation of the rust area in the 23rd block is missing (the error index exceeds the threshold = 0.5), the 2nd, 4th, and 7th attention heads used when generating this block will be located, the gradient direction of their key-value matrices will be calculated, and the parameters will be updated through the momentum acceleration method: , where the momentum coefficient = 0.9 can accelerate convergence. The quality monitoring module runs synchronously, and the average error change rate is calculated every 10 iterations. When , the early stopping mechanism is triggered.
[0029] Step S400: Load the initial environmental monitoring image into the target environmental monitoring annotation neural network for image generation with optimization, and output the optimized monitoring annotation image. Among them, the optimized monitoring annotation image represents the environmental monitoring image output by the target environmental monitoring annotation neural network after iterating the context cache matrix based on the error index, and each image block in the optimized monitoring annotation image meets the target monitoring annotation information.
[0030] Step S400 re - inputs the initial environmental monitoring image into the target environmental monitoring annotation neural network with the optimized context cache matrix to generate an optimized monitoring annotation image that meets the annotation requirements. This process re - uses the block scanning mechanism in step S200, but uses the context cache matrix optimized by gradient in the critical image block processing stage , and its element values are updated through the error back - propagation in step S300. First, perform cache matrix injection: write the optimized key - value pairs into the feature weight focusing components of each layer of the neural network according to the attention head number. For example, in the 3rd attention head, the updated key matrix , where is the parameter increment calculated in step S300
[0031] In an offshore oil spill monitoring case, the optimized key matrix will increase the attention weight of the oil film edge texture features, so that the originally un - labeled thin oil film area (thickness < 0.1mm) can be accurately labeled in the optimized image. In the image generation stage, when processing each block B s in raster order, perform improved attention calculation: for the s - th block, its query vector is calculated for similarity with the optimized key matrix using the formula , where is the attention weight of the s - th block to the j - th historical block, reflecting the feature correlation is the key vector of the j - th historical block of the i - th attention head in the optimized cache matrix, d k is the dimension of the key vector . For example, when processing the 25th block containing the oil spill front, the optimized cache matrix increases the attention weights of this block to the 7th, 12th, and 18th upstream blocks from 0.08 to 0.23, effectively capturing the trajectory features of the oil spill moving with ocean currents. The feed - forward neural network (FFN) adopts a dynamic width adjustment strategy at this stage: for high - complexity regions (such as the building - green space intersection zone in urban heat island effect monitoring), automatically expand the FFN hidden layer dimension to , where , d is the original dimension of the FFN hidden layer is the block complexity factor, calculated jointly through the block texture energy (G is the amplitude of the block grayscale gradient) and the spectral heterogeneity index ( is the variance of the k - th band, c is the total number of bands) (such as , where , is the maximum value of texture energy and spectral heterogeneity index of all blocks in the training set, which is used for normalization. , is the weight coefficient), u and v are spatial frequency domain coordinates, representing the position index of the block in the two-dimensional discrete Fourier transform (DFT) or gradient amplitude space, Indicates the gradient magnitude of the block at position (u,v). Optimize monitoring and annotation images The generation of introduces the residual connection mechanism: each block prediction result With initial block Perform weighted fusion, the formula is , where the weight coefficient By block confidence score Dynamic decision, when >0.9 =0.8 to maintain the stability of the annotation. <0.5 = 0.3 to enhance the original feature retention. After generating the complete image, the Laplacian pyramid decomposition is used to Decompose into Three frequency bands, where L0 is the original resolution layer, L1 is the mid-frequency layer downsampled by 2 times, and L2 is the low-frequency layer downsampled by 4 times. Each layer independently calculates the consistency index with the target annotation information. For example, the high-frequency layer L0 focuses on edge sharpness verification (using Canny edge matching ,in, To optimize monitoring annotation images, E A The low-frequency layer L2 verifies the consistency of regional distribution (through Wasserstein distance measure, To optimize the probability distribution of labeled images, is the probability distribution of the true annotation, is the set of all possible joint distributions, is the transfer plan, which represents the cost of moving a unit mass from position x (labeled result) to position y (true label), is the Euclidean distance between the source point x and the target point y, which measures the spatial offset cost). When any layer indicator exceeds the threshold, the local regeneration process is triggered: locate the abnormal block position , before keeping If the blocks remain unchanged, Start re-executing a limited number of generation iterations (usually 3-5 times), using the Partial Cache Reset mechanism to Replace the key-value pairs with the initial values, and then combine incremental optimization to generate the corrected region. In the final image output stage, perform cross-modal alignment: Align the optimized annotated image with the multi-spectral data cube (where b is the number of bands) for spatial registration. Eliminate the sub-pixel offset caused by the attitude change of the drone through Thin-Plate Spline interpolation. Its energy function is minimized to achieve non-rigid alignment. Among them, is the Thin-Plate Spline interpolation function, representing the spatial deformation field of the image, is the coordinate of the i-th control point, which is the feature matching point between the multi-spectral data and the visible light image, is the target displacement vector, representing the offset that the control point needs to correct, is the regularization coefficient, which balances the data fitting term and the deformation smoothing term. The quality assessment module runs synchronously to calculate the comprehensive score of the optimized image , where the intersection over union measures the accuracy of region annotation. Among them, is the overlapping area, is the union area. The F1 score evaluates the classification accuracy, and the Hausdorff distance quantifies the boundary matching degree. Among them, X is the set of edge points of the predicted annotation region, and Y is the set of edge points of the true annotation region, is the Euclidean distance between points x and y, and sup and inf represent the supremum (greatest lower bound) and infimum (least upper bound) respectively. The weight coefficient . When S > 0.9, it is determined that the optimization is successful and I opt is stored in the database; otherwise, trigger the full process for re-optimization. At this time, the backup cache matrix will be enabled for rollback iteration to avoid error accumulation.
[0032] As an implementation, the method further includes:
[0033] Step 10: Load the initial environmental monitoring image into the target environmental monitoring annotation neural network to obtain a first image patch features, where the target environmental monitoring annotation neural network includes r feature weight focusing components, the first image patch feature is the output result of the r-th feature weight focusing component, the target environmental monitoring annotation neural network outputs a image sub-blocks one by one through repeated iteration in the scanning order of a image patches, the (s + 1)-th image sub-block is determined by loading the s-th image sub-block output by the target environmental monitoring annotation neural network in the s-th image patch scanning order and the s-th context cache matrix into the target environmental monitoring annotation neural network in the (s + 1)-th image patch scanning order, the s-th context cache matrix is used to represent the r feature mapping binary tuples corresponding to the s-th image patch scanning order, the r feature mapping binary tuples are in one-to-one correspondence with the r feature weight focusing components, the a first image patch features are used to determine the unoptimized monitoring annotation image constructed by the a image sub-blocks, where r≥2, a≥s + 1, s≥1;
[0034] Step 20: Load the a first image patch features into the target generation result evaluation component respectively to obtain a evaluation results, and optimize the target context cache matrix, where the target context cache matrix represents the context cache matrix corresponding to the target image sub-block output by the target environmental monitoring annotation neural network after being loaded into the target environmental monitoring annotation neural network, the target image sub-block includes the image sub-blocks that do not meet the target monitoring annotation information determined by the a evaluation results;
[0035] Step 30: Load the initial environmental monitoring image into the target environmental monitoring annotation neural network, determine a second image patch features based on the optimized target context cache matrix, and generate an optimized monitoring annotation image, where the acquisition process of the second image patch feature is the same as the acquisition process of the first image patch feature.
[0036] In step S10, the initial environmental monitoring image is loaded into the target environmental monitoring annotation neural network to obtain a first image patch features. The target environmental monitoring annotation neural network includes r feature weight focusing components. The first image patch feature is the output result of the r-th feature weight focusing component. The target environmental monitoring annotation neural network will output a image sub-blocks one by one through repeated iteration in the scanning order of a image patches. The (s + 1)-th image sub-block is determined by loading the s-th image sub-block output by the target environmental monitoring annotation neural network in the s-th image patch scanning order and the s-th context cache matrix into the target environmental monitoring annotation neural network in the (s + 1)-th image patch scanning order. The s-th context cache matrix is used to represent r feature mapping binary tuples corresponding to the s-th image patch scanning order. The r feature mapping binary tuples are in one-to-one correspondence with the r feature weight focusing components. The a first image patch features are used to determine the unoptimized monitoring annotation image constructed from the a image sub-blocks, where r ≥ 2, a ≥ s + 1, and s ≥ 1.
[0037] The initial environmental monitoring image is the original image data obtained by the drone shooting the target monitoring environment area, which covers various environmental information of the target area. For example, when monitoring a forest environment, the image will contain information about elements such as trees, grasslands, and streams. The target environmental monitoring annotation neural network is a pre-trained model, such as a transformer network, which has the ability to process and analyze the input image to generate the corresponding monitoring annotation image. When loading the initial environmental monitoring image into this neural network, it is necessary to ensure that the format and size of the image meet the input requirements of the neural network. Image preprocessing techniques can be used, such as using the OpenCV library to resize and normalize the image, and convert the image into a format suitable for neural network processing.
[0038] The r feature weight focusing components in the target environmental monitoring annotation neural network are very crucial parts. For example, they can be attention mechanism components. These components can enable the neural network to focus on the important features in the image, improving the performance and accuracy of the model. In the process of processing image sub-blocks, each feature weight focusing component will weight the features of the current image sub-block according to the feature mapping binary tuples in the context cache matrix, thereby highlighting the important feature information. Taking the monitoring of the urban environment as an example, the attention mechanism component can focus more attention on important features such as buildings and roads, while ignoring some irrelevant background information.
[0039] In terms of the image block scanning order, the initial environmental monitoring image is usually divided into blocks and processed sequentially in the order from left to right and from top to bottom. The determination of the (s + 1)-th image block depends on the s-th image block and the s-th context cache matrix. The context cache matrix contains the feature mapping binary tuples corresponding to each feature weight focusing component in multiple feature weight focusing components. These feature mapping binary tuples record the feature information of the previously processed image blocks, which helps the target environmental monitoring annotation neural network better understand the context information of the current image block. Through repeated iteration, the target environmental monitoring annotation neural network outputs a image blocks in sequence in the a-image block scanning order, and these a first image block features are the output results of the r-th feature weight focusing component, which are jointly used to construct the unoptimized monitoring annotation image.
[0040] In step S20, the a first image block features are respectively loaded into the target generation result evaluation component to obtain a evaluation results, and the target context cache matrix is optimized. The target context cache matrix represents the context cache matrix corresponding to the target image block output by the target environmental monitoring annotation neural network after being loaded into the target environmental monitoring annotation neural network. The target image block includes the image blocks that do not meet the target monitoring annotation information determined by the a evaluation results.
[0041] The structure of the target generation result evaluation component includes a fully connected layer, a non-linear activation function layer, and a fully connected layer connected in sequence. Its function is to evaluate the image blocks output by the target environmental monitoring annotation neural network. After loading the a first image block features into this component, the component will evaluate the corresponding image blocks according to these features and obtain a evaluation results. These evaluation results reflect the situation of each image block meeting the target monitoring annotation information. The target monitoring annotation information stipulates the monitoring annotation requirements for the target monitoring environment area, such as the annotation style, the image content to be annotated, etc. If the evaluation result of an image block shows that it does not meet the target monitoring annotation information, then the image block belongs to the target image block.
[0042] The target context cache matrix is optimized according to the a evaluation results. The purpose of the optimization is to enable the target environmental monitoring annotation neural network to refer to more accurate context information when processing subsequent image blocks, thereby improving the accuracy of the output results. For example, when the annotation color of an image block does not meet the requirements of the target monitoring annotation information, the feature mapping binary tuple related to the image block in the target context cache matrix will be adjusted, changing the decision of the target environmental monitoring annotation neural network when processing subsequent image blocks.
[0043] In step S30, the initial environmental monitoring image is loaded into the target environmental monitoring annotation neural network, and a second image block features are determined based on the optimized target context cache matrix to generate an optimized monitoring annotation image. The acquisition process of the second image block features is the same as that of the first image block features.
[0044] The initial environmental monitoring image is loaded into the target environmental monitoring annotation neural network again. At this time, the target environmental monitoring annotation neural network uses the optimized target context cache matrix. Since the target context cache matrix has been optimized, the target environmental monitoring annotation neural network can refer to more accurate context information when processing image blocks, so that the generated a second image block features better meet the requirements of the target monitoring annotation information. Similar to the acquisition process of the first image block features, the target environmental monitoring annotation neural network outputs a image blocks in sequence according to the image block scanning order, combined with the optimized context cache matrix, and then obtains a second image block features.
[0045] An optimized monitoring annotation image is generated according to the a second image block features. During the generation process, the image blocks corresponding to these image block features are stitched and combined according to certain rules to construct a complete optimized monitoring annotation image. Since the optimized context cache matrix is used, each image block in the optimized monitoring annotation image meets the requirements of the target monitoring annotation information and can more accurately reflect the actual situation of the target monitoring environment area.
[0046] In the implementation of the feature weight focusing component, relevant algorithms of the attention mechanism can be adopted, such as the multi-head attention mechanism, and its calculation formula is: , where , , here, Q, K, and V are the query matrix, key matrix, and value matrix respectively, and are learnable weight matrices, d k is the dimension of the key vector.
[0047] In the implementation of the target generation result evaluation component, a model is constructed using a fully connected layer and a non-linear activation function layer, and the evaluation result is obtained through forward calculation. For the optimization of the context cache matrix, optimization algorithms such as gradient descent can be adopted, and its update formula is: , where is the element value in the context cache matrix at the t-th iteration, is the learning rate, is the gradient of the error index with respect to .
[0048] As an implementation manner, in step S20, loading a first image patch features into a target generation result evaluation component respectively to obtain a evaluation results, and optimizing a target context cache matrix, including:
[0049] Step S21: Loading a first image patch features into a target generation result evaluation component respectively to determine a error metrics, and the a evaluation results are in one-to-one correspondence with the a error metrics;
[0050] Step S22: Determining a target error metric corresponding to the x-th image patch scanning order among the a error metrics, where the target error metric is an error metric that does not meet the requirements of the set error metric, and x ≤ a;
[0051] Step S23: Optimizing the target context cache matrix based on the target error metric.
[0052] In step S21, loading a first image patch features into a target generation result evaluation component respectively to determine a error metrics, and the a evaluation results are in one-to-one correspondence with the a error metrics. The first image patch feature is the result output by the r-th feature weight focusing component in the target environmental monitoring annotation neural network after loading the initial environmental monitoring image into the target environmental monitoring annotation neural network, and these features can reflect the key information of each image block. The target generation result evaluation component is a component for evaluating the output result of the target environmental monitoring annotation neural network, and its structure includes a fully connected layer, a non-linear activation function layer, and a fully connected layer connected in sequence. The fully connected layer is used to perform a linear transformation on the input feature vector, and the non-linear activation function layer introduces non-linear factors to increase the expression ability of the model.
[0053] Inputting a first image patch features into the target generation result evaluation component respectively, and the component will process each feature according to the internal calculation rule to obtain the corresponding error metric. The error metric is used to represent the consistency of the image block corresponding to the image patch feature with the target monitoring annotation information. The target monitoring annotation information stipulates the monitoring annotation requirements for the target monitoring environment area, such as the annotation style, the image content to be annotated, etc. Feasible error metric calculation methods include mean square error (MSE), cross entropy loss, etc.
[0054] Taking the monitoring of the urban environment as an example, assuming that the target monitoring annotation information requires buildings to be annotated with a specific color, and the annotation color of the building in a certain image block does not meet the requirement, then the error metric corresponding to the image block will be larger. By calculating a error metrics, a evaluation results can be obtained, and each evaluation result corresponds to an error metric, intuitively reflecting the degree to which each image block meets the target monitoring annotation information.
[0055] In step S22, determine the target error metric corresponding to the scanning order of the x-th image block among the a error metrics, where the target error metric is the error metric that does not meet the requirements of the set error metric, and x ≤ a. The requirement of the set error metric is a preset standard used to determine whether the error metric is within an acceptable range. The a error metrics will be checked and compared one by one to find the error metrics that do not meet the requirements of the set error metric.
[0056] For example, when setting the requirements of the error metric, it is stipulated that the threshold of the error metric is 0.1. When an error metric is greater than 0.1, it is considered that the error metric does not meet the requirements. Among the a error metrics, those greater than 0.1 will be found. Suppose the error metric corresponding to the scanning order of the x-th image block is greater than 0.1, then this error metric is the target error metric. The image block corresponding to the target error metric is the image block that does not meet the target monitoring annotation information and needs to be further processed and optimized.
[0057] In step S23, optimize the target context cache matrix based on the target error metric. The target context cache matrix is the context cache matrix corresponding to the target image block output by the target environmental monitoring annotation neural network after being loaded into the target environmental monitoring annotation neural network. The context cache matrix contains the feature mapping binary tuples corresponding to each feature weight focusing component in multiple feature weight focusing components. These feature mapping binary tuples record the feature information of the previously processed image blocks and are used to help the target environmental monitoring annotation neural network better understand the context information of the current image block.
[0058] By optimizing the target context cache matrix, the context information referred to by the target environmental monitoring annotation neural network when processing image blocks can be adjusted, thereby improving the accuracy of the output result. The optimization process is actually to adjust the element values in the context cache matrix according to the target error metric, so that the target environmental monitoring annotation neural network can make more accurate decisions when processing subsequent image blocks.
[0059] To achieve the optimization of the target context cache matrix, optimization algorithms such as gradient descent can be used. By calculating the gradient of the error metric with respect to the elements in the context cache matrix, and then updating the element values in the context cache matrix along the opposite direction of the gradient, the error metric can be gradually reduced.
[0060] In practical applications, first obtain the target error metric, and then calculate the parameter change rate of the target generation result evaluation component based on this error metric. The parameter change rate reflects the rate at which the error metric changes as the element values in the context cache matrix change. Debug the original parameter matrix according to the parameter change rate and a preset debugging learning rate. The magnitude of the debugging learning rate is positively correlated with the magnitude of the parameter update amount for the iterative debugging of the original parameter matrix, that is, the larger the learning rate, the larger the parameter update amount for each iteration.
[0061] Through continuous iterative debugging, optimize the target context cache matrix with the parameter matrix after debugging for each generation. At the same time, after each optimization of the target context cache matrix, the error metric and parameter change rate corresponding to the x-th image block scanning order of the corresponding generation are obtained again until the error metric corresponding to the x-th image block scanning order meets the set error metric requirements. At this time, the parameter matrix after debugging for the corresponding generation is determined as the target parameter matrix, and the element values in the target parameter matrix are the values for optimizing the target context cache matrix.
[0062] Add the target context cache matrix and the target parameter matrix, and determine the added matrix as the optimization result of the target context cache matrix. In this way, the optimization of the target context cache matrix is completed, enabling the target environmental monitoring annotation neural network to refer to more accurate context information when processing image blocks subsequently, thereby improving the quality of the optimized monitoring annotation image output.
[0063] By executing steps S21 - S23, it is possible to evaluate the image blocks output by the target environmental monitoring annotation neural network, find the image blocks that do not meet the target monitoring annotation information, and optimize the target context cache matrix. In practical applications, the effective implementation of these steps can improve the accuracy and reliability of environmental monitoring annotation.
[0064] As an implementation method, step S23, optimizing the target context cache matrix based on the target error metric, includes:
[0065] Step S231: Obtain an adjustable original parameter matrix;
[0066] Step S232: Determine the parameter change rate of the target generation result evaluation component according to the target error metric, and debug the original parameter matrix based on the parameter change rate until the target parameter matrix is obtained, and the element values in the target parameter matrix are the values for optimizing the target context cache matrix;
[0067] Step S233: Add the target context cache matrix and the target parameter matrix, and determine the added matrix as the optimization result of the target context cache matrix.
[0068] In step S231, an adjustable original parameter matrix is obtained. The original parameter matrix is a basic matrix for debugging and optimizing the target context cache matrix, and it contains a series of adjustable parameter values. These parameter values will be adjusted according to the target error metric in the subsequent optimization process to achieve the optimization of the target context cache matrix. The original parameter matrix can be obtained in various ways. One way is random initialization. For example, when using a deep learning framework for model training, the element values in the original parameter matrix can be randomly generated according to a certain distribution (such as Gaussian distribution). Another way is to initialize the original parameter matrix based on experience or prior knowledge. Suppose in the field of environmental monitoring, according to past monitoring experience, certain features have specific importance weights in the context cache matrix, and the element values in the original parameter matrix can be set according to these experiences so that the initial parameter matrix is more targeted. The dimension and structure of the original parameter matrix usually match the target context cache matrix for effective subsequent calculation and adjustment.
[0069] In step S232, according to the target error metric, the parameter change rate of the target generation result evaluation component is determined, and the original parameter matrix is debugged based on the parameter change rate until the target parameter matrix is obtained. The element values in the target parameter matrix are the values for optimizing the target context cache matrix. The target error metric is determined in step S22, which represents the error metric corresponding to the x-th image block scanning order and does not meet the requirements of the set error metric. The target generation result evaluation component is a component for evaluating the output result of the target environmental monitoring annotation neural network, and the error metric it outputs reflects the consistency of the image blocks meeting the target monitoring annotation information.
[0070] Calculate the parameter change rate of the target generation result evaluation component according to the target error metric. The parameter change rate can be obtained by calculating the gradient of the error metric with respect to the elements in the original parameter matrix. In deep learning, the gradient represents the change rate of a function at a certain point. By calculating the gradient of the error metric with respect to the parameters, it can be known how to adjust the parameters to reduce the error metric. For example, using the mean squared error (MSE), the chain rule can be used to calculate the partial derivative of the MSE with respect to the elements in the original parameter matrix to obtain the parameter change rate.
[0071] Based on the obtained parameter change rate, the original parameter matrix is debugged. During the debugging process, an important parameter is required, namely the debugging learning rate. The size of the debugging learning rate is positively correlated with the size of the parameter update amount for the iterative debugging of the original parameter matrix. If the debugging learning rate is set too large, the parameter update amount will be very large, which may cause the algorithm not to converge and even lead to the divergence of parameter values. If the debugging learning rate is set too small, the parameter update amount will be very small, and the algorithm convergence speed will be very slow, requiring more iteration times to achieve the ideal optimization effect.
[0072] Iteratively debug the original parameter matrix according to the debugging learning rate. In each generation, optimize the target context cache matrix through the debugged parameter matrix. At the same time, after each optimization of the target context cache matrix, obtain the error index and parameter change rate corresponding to the x-th image block scanning order of the corresponding generation again. This process is repeated until the error index corresponding to the x-th image block scanning order meets the set error index requirements. At this time, the debugged parameter matrix of the corresponding generation is determined as the target parameter matrix.
[0073] In step S233, add the target context cache matrix and the target parameter matrix, and determine the added matrix as the optimization result of the target context cache matrix. After the iterative debugging in step S232, the target parameter matrix is obtained. The element values in this matrix are obtained by optimizing the original parameter matrix according to the target error index. Add the corresponding elements of the target context cache matrix and the target parameter matrix to obtain a new matrix. This new matrix combines the information of the original target context cache matrix and the optimized and adjusted target parameter matrix, enabling the target environmental monitoring annotation neural network to refer to more accurate context information when processing image blocks. For example, assume the target context cache matrix is C and the target parameter matrix is P, then the optimized target context cache matrix C' is C' = C + P.
[0074] By executing steps S231 - S233, the target context cache matrix can be effectively optimized based on the target error index. This optimization can significantly improve the quality of the monitoring annotation images generated by the target environmental monitoring annotation neural network.
[0075] As an implementation manner, step S232, determining the parameter change rate of the target generation result evaluation component according to the target error index, and debugging the original parameter matrix based on the parameter change rate to obtain the target parameter matrix, includes:
[0076] Step S2321: Obtain a pre-determined debugging learning rate, where the magnitude of the debugging learning rate is positively correlated with the magnitude of the parameter update amount for the iterative debugging of the original parameter matrix;
[0077] Step S2322: Iteratively debug the original parameter matrix according to the debugging learning rate. In each generation, optimize the target context cache matrix through the debugged parameter matrix. At the same time, after each optimization of the target context cache matrix, obtain the error index and parameter change rate corresponding to the x-th image block scanning order of the corresponding generation again until the error index corresponding to the x-th image block scanning order meets the set error index requirements, and determine the debugged parameter matrix of the corresponding generation as the target parameter matrix, where the debugged parameter matrix of each generation is determined by the parameter matrix before each generation of debugging, the corresponding parameter change rate of each generation, and the debugging learning rate together.
[0078] In step S232, through these two steps, the debugging of the original parameter matrix is gradually realized to obtain a target parameter matrix that can optimize the target context cache matrix, thereby improving the performance of the target environmental monitoring annotation neural network and making the generated monitoring annotation image more conform to the target monitoring annotation information.
[0079] In step S2321, a pre-determined debugging learning rate is obtained. The magnitude of the debugging learning rate is positively correlated with the magnitude of the parameter update amount of the iterative debugging of the original parameter matrix. The debugging learning rate is an important hyperparameter that controls the update amplitude of the original parameter matrix during the iterative debugging process. In optimization algorithms such as the gradient descent algorithm, the update of parameters is based on the current parameter change rate and the debugging learning rate. If the debugging learning rate is set too large, the update amount of parameters in each iteration will be very large, which may cause the algorithm to fail to converge, and even the parameter values may diverge, making the optimization process unable to achieve the expected effect. On the contrary, if the debugging learning rate is set too small, the update amount of parameters in each iteration will be very small, and the convergence speed of the algorithm will be very slow, requiring more iteration times to achieve the ideal optimization effect, which will increase the computing time and resource consumption.
[0080] The debugging learning rate can be pre-determined in various ways. One way is to use a fixed learning rate, that is, according to experience or experimental results, a suitable constant is selected as the debugging learning rate. For example, in many deep learning tasks, common fixed learning rate values are 0.01, 0.001, etc. Another way is to adopt a learning rate decay strategy, that is, as the number of iterations increases, the debugging learning rate is gradually decreased. This strategy can use a larger learning rate at the initial stage of the optimization process to enable the parameters to quickly approach the optimal value, and then use a smaller learning rate in the later stage for more refined adjustment to avoid missing the optimal solution. Feasible learning rate decay methods include step decay, exponential decay, etc. Taking exponential decay as an example, its calculation formula is: , where is the debugging learning rate at the t-th iteration, is the initial debugging learning rate, is the decay coefficient, and t is the number of iterations.
[0081] In step S2322, the original parameter matrix is iteratively debugged according to the debugging learning rate. In each generation, the target context cache matrix is optimized through the debugged parameter matrix. Meanwhile, after each optimization of the target context cache matrix, the error index and the parameter change rate corresponding to the x-th image block scanning order of the corresponding generation are obtained again until the error index corresponding to the x-th image block scanning order meets the set error index requirement. The debugged parameter matrix of the corresponding generation is determined as the target parameter matrix. The debugged parameter matrix of each generation is determined together by the parameter matrix before debugging in each generation, the corresponding parameter change rate in each generation, and the debugging learning rate.
[0082] After obtaining the debugging learning rate, iterative debugging of the original parameter matrix begins. In each generation of iteration, first, the parameter change rate of the target generation result evaluation component is determined according to the current original parameter matrix and the target error index. The parameter change rate can be obtained by calculating the gradient of the error index with respect to the elements in the original parameter matrix. In deep learning, the gradient represents the change rate of a function at a certain point. By calculating the gradient of the error index with respect to the parameters, it can be known how to adjust the parameters to reduce the error index. For example, using the mean squared error (MSE), the partial derivative of MSE with respect to the elements in the original parameter matrix is calculated using the chain rule to obtain the parameter change rate. Then, the original parameter matrix is updated according to the debugging learning rate and the parameter change rate. Suppose the current original parameter matrix is , the debugging learning rate is , and the parameter change rate of the t-th generation is , then the update formula for the parameter matrix of the (t + 1)-th generation is: , this formula means that in each generation of iteration, the original parameter matrix is updated along the opposite direction of the parameter change rate, and the update amplitude is controlled by the debugging learning rate.
[0083] Next, the target context cache matrix is optimized using the debugged parameter matrix. Specifically, the debugged parameter matrix is added to the target context cache matrix to obtain the optimized target context cache matrix. Then, the initial environmental monitoring image is loaded into the target environmental monitoring annotation neural network, new image blocks are generated based on the optimized target context cache matrix, and the error index and the parameter change rate corresponding to the x-th image block scanning order of the corresponding generation are obtained again.
[0084] This iterative process is continuously repeated until the error metric corresponding to the scanning order of the x-th image patch meets the set error metric requirements. The set error metric requirements are a pre-set standard used to determine whether the error metric is within an acceptable range. For example, if the threshold of the set error metric is 0.1, when the error metric corresponding to the scanning order of the x-th image patch is less than 0.1, it is considered that the error metric meets the requirements. At this time, the parameter matrix after corresponding iteration and debugging is determined as the target parameter matrix, and the element values in the target parameter matrix are the values for optimizing the target context cache matrix. During the iterative debugging process, a loop structure and an optimization algorithm (such as the gradient descent algorithm) are used to update the parameters. At the same time, the entire iterative process can be monitored and recorded to analyze the changes in the error metric, so as to adjust the debugging learning rate or optimization algorithm in a timely manner to ensure that the iterative process can converge efficiently.
[0085] By executing steps S2321 - S2322, the original parameter matrix can be effectively iteratively debugged according to the target error metric and the debugging learning rate, and finally the target parameter matrix can be obtained. It can significantly improve the quality of the monitoring annotation images generated by the target environmental monitoring annotation neural network.
[0086] As an implementation, in step S400, loading the initial environmental monitoring image into the target environmental monitoring annotation neural network for image generation with optimization and outputting the optimized monitoring annotation image includes:
[0087] Step S410: Loading the initial environmental monitoring image into the target environmental monitoring annotation neural network to generate a second image patch features, and the context cache matrix corresponding to the scanning order of the x-th image patch has been optimized;
[0088] Step S420: Loading the a second image patch features into the target generation result evaluation component respectively to obtain a evaluation results again, where the a evaluation results obtained again all represent that the corresponding image patches meet the target monitoring annotation information;
[0089] Step S430: Sampling the a second image patch features to obtain the optimized monitoring annotation image.
[0090] In step S400, through these steps, the previously optimized context cache matrix is fully utilized to generate high-quality optimized monitoring annotation images that meet the target monitoring annotation information.
[0091] In step S410, the initial environmental monitoring image is loaded into the target environmental monitoring annotation neural network to generate a second image patch features. The context cache matrix corresponding to the scanning order of the x-th image patch has been optimized. The initial environmental monitoring image is the original image data obtained by the drone shooting the target monitoring environment area, which contains rich environmental information of the target area. For example, when monitoring an urban area, the image may contain information about elements such as buildings, roads, and green areas. The target environmental monitoring annotation neural network is a pre-trained and optimized model, such as a transformer network, which has the ability to process and analyze the input image and then generate the corresponding monitoring annotation image. When loading the initial environmental monitoring image, it is necessary to ensure that the format and size of the image meet the input requirements of the target environmental monitoring annotation neural network. Image preprocessing techniques can be used, such as using the OpenCV library to resize and normalize the image, and convert the image into a format suitable for neural network processing.
[0092] Since in the previous step, the context cache matrix corresponding to the scanning order of the x-th image patch has been optimized based on the target error metric. The context cache matrix contains the feature mapping binary groups corresponding to each feature weight focusing component in multiple feature weight focusing components. These feature mapping binary groups record the feature information of the previously processed image patches, which helps the target environmental monitoring annotation neural network better understand the context information of the current image patch. The optimized context cache matrix can provide more accurate reference information for the neural network, enabling the neural network to make more reasonable decisions when processing image patches.
[0093] The target environmental monitoring annotation neural network processes the image patches one by one according to the image patch scanning order. The image patch scanning order usually divides and processes the initial environmental monitoring image in the order from left to right and from top to bottom. When processing each image patch, the neural network combines the optimized context cache matrix and generates the corresponding image patch features through internal calculations and transformations. After a series of processes, a second image patch features are finally obtained. These features can more accurately reflect the content and features of the image patches, laying a foundation for generating high-quality optimized monitoring annotation images in the future.
[0094] In step S420, the a second image patch features are respectively loaded into the target generation result evaluation component to obtain a evaluation results again. Among them, the a evaluation results obtained again all represent that the corresponding image patches meet the target monitoring annotation information. The target generation result evaluation component is a component used to evaluate the output results of the target environmental monitoring annotation neural network. Its structure includes a fully connected layer, a non-linear activation function layer, and a fully connected layer connected in sequence. The fully connected layer is used to perform a linear transformation on the input feature vector, and the non-linear activation function layer introduces non-linear factors to increase the expression ability of the model.
[0095] Input a second image patch features into the target generation result evaluation component respectively. The component will process each feature according to the internal calculation rules to obtain the corresponding evaluation result. The evaluation result reflects the degree to which each image patch meets the target monitoring annotation information. The target monitoring annotation information stipulates the monitoring annotation requirements for the target monitoring environment area, such as the annotation style, the image content to be annotated, etc. Since the optimized context cache matrix is used in step S410, the generated second image patch features better meet the requirements of the target monitoring annotation information. Therefore, the a evaluation results obtained again all indicate that the corresponding image patches meet the target monitoring annotation information.
[0096] For example, when monitoring a forest environment, the target monitoring annotation information requires that the tree area be marked with a specific color and the grass area be marked with another color. After the optimized context cache matrix, the image patches corresponding to the second image patch features generated by the target environmental monitoring annotation neural network can be accurately marked according to the requirements, and the evaluation results given by the target generation result evaluation component show that these image patches all meet the marking requirements.
[0097] In step S430, sample the a second image patch features to obtain the optimized monitoring annotation image. Sampling is to select appropriate features from the a second image patch features to construct the final optimized monitoring annotation image. Multiple sampling methods can be used, such as random sampling, stratified sampling, etc. Random sampling is to randomly select a certain number of features from the a second image patch features; stratified sampling is to divide the image patches into different layers according to some features of the image patches (such as position, category, etc.), and then extract a certain number of features from each layer.
[0098] During the sampling process, it is necessary to consider the rationality and representativeness of the sampling to ensure that the selected features can accurately reflect the situation of the entire target monitoring environment area. For example, when monitoring a large urban area, the environmental characteristics of different areas may vary. Stratified sampling can be used to divide the city into different functional areas (commercial areas, residential areas, industrial areas, etc.) and then extract an appropriate amount of second image patch features from each layer.
[0099] After sampling, splice and combine the image patches corresponding to the selected second image patch features to construct a complete optimized monitoring annotation image. During the splicing process, it is necessary to ensure that the positions and orders of the image patches are correct to form a coherent and accurate image. The finally obtained optimized monitoring annotation image is the result generated by the optimized target environmental monitoring annotation neural network, in which each image patch meets the requirements of the target monitoring annotation information and can accurately reflect the actual situation of the target monitoring environment area.
[0100] By executing steps S410 - S430, the optimized context cache matrix can be fully utilized to generate a high-quality optimized monitoring annotation image, improving the accuracy and reliability of environmental monitoring and providing strong support for environmental monitoring work.
[0101] As an implementation, the method further includes:
[0102] Step S24: Normalize a first set of a image patch features and a second set of a image patch features respectively to obtain a first probability density and a second probability density;
[0103] Step S25: Determine a target distribution deviation based on the first probability density and the second probability density, where the target distribution deviation is used to evaluate the deviation between the first probability density and the second probability density;
[0104] Step S26: Optimize the target context cache matrix along the direction that minimizes the magnitude of the target distribution deviation.
[0105] Steps S24 - S26 are further supplements to the image generation and optimization process in the UAV-based environmental monitoring method. Through these steps, in-depth analysis of the first set of image patch features and the second set of image patch features is carried out to evaluate the optimization effect and further optimize the target context cache matrix, thereby improving the accuracy and reliability of the entire environmental monitoring annotation.
[0106] In step S24, the a first image patch features and the a second image patch features are normalized respectively to obtain a first probability density and a second probability density. The first image patch features are the results output by the r-th feature weight focusing component in the target environmental monitoring annotation neural network when the initial environmental monitoring image is loaded into the target environmental monitoring annotation neural network without optimizing the context cache matrix; the second image patch features are the corresponding features obtained after the context cache matrix is iteratively optimized based on the error metric and then the initial environmental monitoring image is loaded into the target environmental monitoring annotation neural network again. Normalization is a feasible method for data preprocessing, whose purpose is to convert data with different ranges and scales into data with the same scale and distribution for subsequent analysis and comparison. Feasible normalization methods include Z-score normalization and Min - Max normalization.
[0107] After normalizing a first set of a image patch features and a second set of a image patch features respectively, the normalized feature data is obtained. To further analyze the distribution of these feature data, their probability density is calculated. The probability density function describes the probability distribution of a random variable at each value point. For a continuous random variable, the probability density function can be estimated by methods such as kernel density estimation. The basic idea of kernel density estimation is to place a kernel function around each data point and then perform a weighted sum of these kernel functions to obtain the probability density function of the entire data set. For example, using a Gaussian kernel function, through kernel density estimation, the first probability density of the first image patch features and the second probability density of the second image patch features are obtained. For example, when monitoring a forest environment, the first image patch features may contain the feature information of areas such as trees and grasslands before optimization, and the second image patch features are the feature information of the corresponding areas after optimization. Through normalization and probability density estimation, the distribution changes of these features before and after optimization can be clearly seen.
[0108] In step S25, based on the first probability density and the second probability density, a target distribution deviation is determined. The target distribution deviation is used to evaluate the deviation between the first probability density and the second probability density. The target distribution deviation can measure how much the distribution of the image patch features has changed after optimizing the target context cache matrix, and thus reflect the impact degree of the optimization operation on the output result of the target environmental monitoring annotation neural network. Feasible metrics for measuring distribution deviation include KL divergence (Kullback - Leibler divergence) and JS divergence (Jensen - Shannon divergence). The calculation formula of KL divergence is: , where P(x) is the first probability density and Q(x) is the second probability density. KL divergence represents the amount of information loss from probability distribution P to probability distribution Q. The larger its value, the greater the difference between the two distributions. It should be noted that KL divergence is asymmetric, that is .
[0109] JS divergence is an improvement of KL divergence. It is symmetric, and its calculation formula is: , where . The value range of JS divergence is between [0, 1]. A value of 0 indicates that the two distributions are exactly the same, and a value of 1 indicates that the two distributions are completely different.
[0110] By calculating the target distribution deviation, the differences in the distributions of the first image patch features and the second image patch features can be intuitively understood. For example, if the target distribution deviation is small, it indicates that the optimization operation on the target context cache matrix has not caused much change in the distribution of the image patch features, which may mean that the optimization effect is not obvious; conversely, if the target distribution deviation is large, it indicates that the optimization operation has had a significant impact on the distribution of the image patch features, which may indicate that the optimization has achieved good results.
[0111] In step S26, the target context cache matrix is optimized along the direction that minimizes the magnitude of the target distribution deviation. The target context cache matrix is a caching mechanism used by the target environment monitoring annotation neural network when processing image patches. It contains the feature mapping binary tuples corresponding to each feature weight focusing component in multiple feature weight focusing components. These feature mapping binary tuples record the feature information of the previously processed image patches and are used to help the target environment monitoring annotation neural network better understand the context information of the current image patch. The goal is to adjust the element values in the target context cache matrix to minimize the target distribution deviation between the first probability density and the second probability density. This can be achieved through an optimization algorithm, such as the gradient descent algorithm. The basic idea of the gradient descent algorithm is to update the parameters along the negative gradient direction of the objective function to gradually decrease the value of the objective function. In this problem, the objective function is the target distribution deviation, and the parameters are the element values in the target context cache matrix.
[0112] Suppose the target distribution deviation is , where represents the element value in the target context cache matrix. The update formula of the gradient descent algorithm is: , where is the element value in the target context cache matrix at the t-th iteration, is the learning rate, is the gradient of the target distribution deviation with respect to .
[0113] By continuously iteratively updating the element values in the target context cache matrix, the target distribution deviation is gradually reduced. When the target distribution deviation reaches a small value or meets the preset convergence condition, the iterative process stops, and the target context cache matrix obtained at this time is the optimized result.
[0114] Through the execution of steps S24-S26, the distribution difference between the first image block feature and the second image block feature can be deeply analyzed, the effect of the target context cache matrix optimization can be evaluated, and the target context cache matrix can be further optimized to improve the quality of the monitoring and annotation image generated by the target environment monitoring and annotation neural network. In practical applications, these steps can help to monitor and annotate the target environment area more accurately, providing a more reliable basis for environmental monitoring work.
[0115] As an implementation mode, the method further includes a pre-training process of an environmental monitoring annotation neural network, specifically including:
[0116] Step S1: Obtain image training samples without labels;
[0117] Step S2: pre-debugging the initial environment monitoring annotation neural network based on the image training sample to obtain the target environment monitoring annotation neural network, wherein the initial environment monitoring annotation neural network and the target environment monitoring annotation neural network execute a global feature weight focusing strategy on the loaded image and execute a local feature weight focusing strategy on the output image.
[0118] Steps S1-S2 in the pre-training process of the environmental monitoring annotation neural network are an important basis for building an effective neural network model. Through these two steps, from data collection to model training, the pre-debugging of the initial environmental monitoring annotation neural network is gradually completed, and finally the target environmental monitoring annotation neural network is obtained.
[0119] In step S1, image training samples without labels are obtained. Image training samples are data used to train neural networks. Without labels, these images do not have pre-labeled monitoring information, such as the specific locations and categories of various environmental elements in the image (such as trees in the forest, buildings in the city, etc.) are not clearly marked. These image training samples can be obtained in a variety of ways, such as from public image data sets, satellite image databases, or through large-scale photography of drones in different environmental areas. Taking the monitoring of urban environment as an example, a large number of images of urban areas can be downloaded from the satellite image database. These images contain various elements such as streets, buildings, and greenery in the city, but these elements are not labeled in detail.
[0120] Obtaining unlabeled image training samples. On the one hand, a large amount of unlabeled image data is relatively easy to obtain, which can provide rich information for neural networks to learn the common features of images; on the other hand, unsupervised learning methods can allow neural networks to automatically discover potential patterns and structures in images without relying on a large amount of manual annotation.
[0121] In step S2, the initial environmental monitoring annotation neural network is pre - debugged based on image training samples to obtain the target environmental monitoring annotation neural network. The initial environmental monitoring annotation neural network and the target environmental monitoring annotation neural network execute the global feature weight focusing strategy on the loaded image and the local feature weight focusing strategy on the output image. The initial environmental monitoring annotation neural network is a pre - constructed neural network model with certain structures and parameters, but it has not been fully trained for the environmental monitoring annotation task.
[0122] The global feature weight focusing strategy, that is, global attention, means that when the neural network processes the input image, it can pay attention to the overall features and global information of the image. For example, when processing an image of a forest environment, the global attention mechanism allows the neural network to consider information such as the overall layout of the forest and the distribution range of vegetation. The global attention mechanism is usually implemented by calculating the correlation degree between different positions in the input image. A feasible calculation method is to use attention scores. Suppose the feature representation of the input image is , where N is the number of features in the image and D is the dimension of the features. The global attention mechanism calculates the attention score to measure the correlation degree between different features. The calculation of the attention score can use the dot - product attention formula: , where Q and K are the query matrix and the key matrix respectively, and d k is the dimension of the key vector.
[0123] The local feature weight focusing strategy, that is, local attention, means that when the neural network outputs an image, it can focus on the local regions and detailed information of the image. For example, when annotating trees in a forest, the local attention mechanism allows the neural network to more accurately identify details such as the boundaries and species of the trees. The local attention mechanism is usually implemented by restricting the attention range and only focusing on the local region around the current position.
[0124] When pre - debugging the initial environmental monitoring annotation neural network, unsupervised learning algorithms such as autoencoders and generative adversarial networks are used. Taking the autoencoder as an example, the goal of the autoencoder is to encode the input image into a low - dimensional representation and then decode an output image as similar as possible to the input image from this low - dimensional representation. The loss function of the autoencoder can use the mean squared error (MSE), and its calculation formula is , where is the input image, is the decoded output image, and n is the number of samples.
[0125] During the training process, the parameters of the initial environmental monitoring annotation neural network are continuously adjusted to gradually reduce the value of the loss function. As the training progresses, the neural network will learn the feature representations and potential structures of the images. After a sufficient number of training iterations, the target environmental monitoring annotation neural network is obtained, which can better process environmental monitoring images and perform better in subsequent environmental monitoring annotation tasks.
[0126] When implementing these steps, various technical means can be adopted. In terms of obtaining image training samples, web crawler technology is used to download images from public image datasets and databases under the permission of laws and regulations, or historical images are used as samples.
[0127] Please refer to Figure 3 , which is a schematic structural diagram of a monitoring system 10 provided by an embodiment of the present application. As Figure 3 shown, the monitoring system 10 may include: a processor 1001, a network interface 1004, and a memory 1005. In addition, the above-mentioned monitoring system 10 may further include: a user interface 1003 and at least one communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. Among them, the user interface 1003 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. Optionally, the memory 1005 may also be at least one storage device far from the aforementioned processor 1001. As Figure 3 shown, the memory 1005, as a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.
[0128] In Figure 3 the monitoring system 10 shown, the network interface 1004 can provide network communication functions; while the user interface 1003 is mainly used to provide an input interface; and the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement the methods provided in the above embodiments.
[0129] It should be understood that the monitoring system 10 described in the embodiments of the present application can execute the descriptions of the drone-based environmental monitoring method in the corresponding embodiments mentioned above. In addition, the descriptions of the beneficial effects of adopting the same method will not be repeated. Figure 2
[0130] As an example, the above program instructions can be deployed to execute on a monitoring system, or on at least two monitoring systems at a location, or, on at least two monitoring systems distributed at at least two locations and interconnected by a communication network. The at least two monitoring systems distributed at at least two locations and interconnected by a communication network can form a blockchain network.
[0131] The above computer-readable storage medium can be the central storage unit of the monitoring system provided in any of the foregoing embodiments, such as the hard disk or the medium memory of the monitoring system. The computer-readable storage medium can also be an external storage device of the monitoring system, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the monitoring system. Further, the computer-readable storage medium can also include both the central storage unit and the external storage device of the monitoring system. The computer-readable storage medium is used to store the computer program and other programs and data required by the monitoring system. The computer-readable storage medium can also be used to temporarily store the data that has been output or will be output.
[0132] In the description of the specification, claims and drawings of the embodiments of the present application, the terms "first", "second", etc. are used to distinguish the content in different media, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units is not limited to the listed steps or modules, but optionally further includes steps or modules not listed, or optionally further includes other step units inherent to these processes, methods, devices, products or equipment.
[0133] The embodiments of the present application also provide a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, they implement the description of the above drone-based environmental monitoring method in the corresponding embodiments Figure 2 Therefore, the description will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated. For the technical details not disclosed in the embodiments of the computer program product involved in the present application, please refer to the description of the method embodiments of the present application.
[0134] Those of ordinary skill in the art will realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0135] The methods and related devices provided by the embodiments of this application are described with reference to the method flowcharts and / or structural schematic diagrams provided by the embodiments of this application. Specifically, each process and / or block of the method flowchart and / or structural schematic diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable network-connected devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable network-connected devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable network-connected device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable network-connected device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or structural schematic one block or multiple blocks. The above disclosure is only for the preferred embodiments of this application, and of course, it cannot be used to limit the scope of rights of this application. Therefore, equivalent changes made according to the claims of this application still fall within the scope covered by this application.
Claims
1. An unmanned aerial vehicle-based environmental monitoring method, characterized in that, Including: Obtaining an initial environmental monitoring image in a target monitoring environment area captured by a drone, and obtaining target monitoring annotation information, where the target monitoring annotation information is used to indicate the monitoring annotation requirements for the target monitoring environment area; Loading the initial environmental monitoring image into a target environmental monitoring annotation neural network for unoptimized image generation, and outputting an unoptimized monitoring annotation image; Loading the image block features corresponding to each image block in the unoptimized monitoring annotation image into a target generation result evaluation component respectively, determining the error metrics of the target generation result evaluation component respectively, and iterating a context cache matrix based on the error metrics, where the error metrics are used to represent the consistency of the image block corresponding to the image block feature with the target monitoring annotation information; Loading the initial environmental monitoring image into the target environmental monitoring annotation neural network for image generation to be optimized, and outputting an optimized monitoring annotation image, where the optimized monitoring annotation image represents the environmental monitoring image output by the target environmental monitoring annotation neural network after iterating the context cache matrix based on the error metrics, and each image block in the optimized monitoring annotation image satisfies the target monitoring annotation information.
2. The method according to claim 1, characterized in that, Wherein, The unoptimized monitoring annotation image includes image blocks that do not satisfy the target monitoring annotation information, and the unoptimized monitoring annotation image is constructed from the image blocks output by the target environmental monitoring annotation neural network. The target environmental monitoring annotation neural network is used to determine the output image blocks one by one according to the loaded image blocks and the context cache matrix in the order of image block scanning, and determine the corresponding monitoring annotation image based on the output image blocks. The target environmental monitoring annotation neural network includes a plurality of feature weight focusing components, and the context cache matrix includes the feature mapping binary groups corresponding to each feature weight focusing component in the plurality of feature weight focusing components.
3. The method according to claim 2, wherein The method further includes: Loading the initial environmental monitoring image into the target environmental monitoring annotation neural network to obtain a first image patch features, where the target environmental monitoring annotation neural network includes r feature weight focusing components, the first image patch feature is the output result of the r-th feature weight focusing component, the target environmental monitoring annotation neural network outputs a image patches one by one through repeated iteration in the scanning order of a image patches, the (s + 1)-th image patch is determined by loading the s-th image patch output by the target environmental monitoring annotation neural network in the s-th image patch scanning order and the s-th context cache matrix into the target environmental monitoring annotation neural network in the (s + 1)-th image patch scanning order, the s-th context cache matrix is used to represent r feature map binary tuples corresponding to the s-th image patch scanning order, the r feature map binary tuples are in one-to-one correspondence with the r feature weight focusing components, the a first image patch features are used to determine the unoptimized monitoring annotation image constructed by the a image patches, where r≥2, a≥s + 1, s≥1; Loading the a first image patch features into the target generation result evaluation component respectively to obtain a evaluation results, and optimizing the target context cache matrix, where the target context cache matrix represents the context cache matrix corresponding to the target image patch output by the target environmental monitoring annotation neural network after being loaded into the target environmental monitoring annotation neural network, and the target image patch includes the image patches that do not meet the target monitoring annotation information determined by the a evaluation results; Loading the initial environmental monitoring image into the target environmental monitoring annotation neural network, determining a second image patch features based on the optimized target context cache matrix, and generating the optimized monitoring annotation image, where the acquisition process of the second image patch feature is the same as that of the first image patch feature.
4. The method according to claim 3, wherein The step of loading the a first image patch features into the target generation result evaluation component respectively to obtain a evaluation results, and optimizing the target context cache matrix includes: Loading the a first image patch features into the target generation result evaluation component respectively to determine a error metrics, and the a evaluation results are in one-to-one correspondence with the a error metrics; Determining the target error metric corresponding to the x-th image patch scanning order among the a error metrics, where the target error metric is the error metric that does not meet the requirements of the set error metric, and x≤a; Optimizing the target context cache matrix based on the target error metric.
5. The method according to claim 4, characterized in that, The step of optimizing the target context cache matrix based on the target error metric includes: Obtaining an adjustable original parameter matrix; Determining the parameter change rate of the target generation result evaluation component according to the target error metric, and debugging the original parameter matrix based on the parameter change rate until a target parameter matrix is obtained, and the element values in the target parameter matrix are the values for optimizing the target context cache matrix; Add the target context cache matrix to the target parameter matrix, and determine the result of optimizing the target context cache matrix as the added matrix.
6. The method according to claim 5, wherein The determining the parameter change rate of the target generation result evaluation component according to the target error index, and debugging the original parameter matrix based on the parameter change rate until a target parameter matrix is obtained includes: Obtain a pre-determined debugging learning rate, where the magnitude of the debugging learning rate is positively correlated with the magnitude of the parameter update amount of the iterative debugging of the original parameter matrix; Iteratively debug the original parameter matrix according to the debugging learning rate. In each generation, optimize the target context cache matrix with the parameter matrix after debugging. At the same time, after each optimization of the target context cache matrix, obtain the error index and parameter change rate corresponding to the scanning order of the x-th image block in the corresponding generation again until the error index corresponding to the scanning order of the x-th image block meets the requirements of the set error index. Determine the parameter matrix after debugging in the corresponding generation as the target parameter matrix, where the parameter matrix after debugging in each generation is determined together with the parameter matrix before debugging in each generation, the corresponding parameter change rate in each generation, and the debugging learning rate.
7. The method according to claim 4, wherein The loading the initial environmental monitoring image into the target environmental monitoring annotation neural network to generate an image to be optimized, and outputting an optimized monitoring annotation image includes: Load the initial environmental monitoring image into the target environmental monitoring annotation neural network to generate the a second image block features, and the context cache matrix corresponding to the scanning order of the x-th image block has been optimized; Load the a second image block features into the target generation result evaluation component respectively to obtain a evaluation results again, where the a evaluation results obtained again all represent that the corresponding image blocks meet the target monitoring annotation information; Perform sampling processing on the a second image block features to obtain the optimized monitoring annotation image.
8. The method according to claim 3, wherein The method further includes: Perform normalization processing on the a first image block features and the a second image block features respectively to obtain a first probability density and a second probability density; Determine a target distribution deviation based on the first probability density and the second probability density, where the target distribution deviation is used to evaluate the deviation between the first probability density and the second probability density; Optimize the target context cache matrix along the direction that minimizes the magnitude of the target distribution deviation.
9. The method according to claim 1, characterized in that, The method further includes: Obtain an image training sample without a label; Perform pre-debugging on the initial environmental monitoring annotation neural network based on the image training sample to obtain the target environmental monitoring annotation neural network, where the initial environmental monitoring annotation neural network and the target environmental monitoring annotation neural network execute a global feature weight focusing strategy on the loaded image and a local feature weight focusing strategy on the output image.
10. A monitoring system, characterized in that, Including: One or more processors; and one or more memories, wherein computer-readable code is stored in the memories, and when the computer-readable code is run by the one or more processors, causes the one or more processors to perform the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Data processing method and device, attitude prediction method and device and storage medium
CN113807150A
Robust machine learning for imperfect labeled image segmentation
US20220067940A1