A method for preprocessing input grayscale images to improve sub-pixel offset tracking accuracy
Feature extraction and loss function optimization for multi-band images through deep neural networks, solving the problem of low efficiency of manual input grayscale images in the prior art, and achieving more efficient and accurate input image selection with subpixel offset tracking algorithm.
Patent Information
- Application Number
- CN202311455674.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-03
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-11-03
AI Technical Summary
In the subpixel offset tracking algorithm, the prior art relies on manual trial and error when selecting input grayscale images, which is inefficient and cannot use semantic information in the image, resulting in unstable quality of the calculation results.
Feature extraction is performed on multi-band images through deep neural networks (such as ResNet backbone network and SoftMax activation function), and the loss function (including mutual information and two-dimensional entropy loss terms) is calculated to optimize network parameters to generate optimized input grayscale images.
Automatic input grayscale image feature extraction is realized, which avoids the huge workload of manual trial and error, improves the accuracy of the subpixel offset tracking algorithm, and can quickly and conveniently select high-quality input images in the research area.
Smart Images

Figure CN117710798B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of surface deformation monitoring, and in particular to an input grayscale image preprocessing method for improving sub-pixel offset tracking accuracy. Background Art
[0002] When using the sub-pixel offset tracking algorithm for surface deformation analysis, a key influencing factor is how to select the appropriate input image band. Generally, the drone aerial photography data set processed by the SfM algorithm can export multiple one-dimensional grayscale images such as Red Band, Green Band, Blue Band, DSM, hillshade, etc. Usually one of the grayscale images is selected to input the sub-pixel offset tracking algorithm for processing. Different grayscale images have no obvious difference in calculation time, but will produce significant differences in the quality of the calculation results (i.e., displacement field).
[0003] Usually, in order to select an optimal grayscale image for inputting the sub-pixel offset tracking algorithm, a manual trial and error method under small window conditions is often adopted. That is, under a smaller window parameter, different grayscale images are input in groups, and the accuracy of the results obtained by the sub-pixel offset tracking algorithm under the condition that the grayscale image is used as the input image is evaluated by visual inspection or calculation of the average gradient, signal-to-noise ratio, etc. This process mainly relies on manual work, which is time-consuming and laborious. In addition, the quality of the input grayscale image is judged only by the quality of the output result after the algorithm calculation is completed, ignoring the semantic information contained in the input grayscale image itself, which is essentially equivalent to the operation of feature selection, that is, only the best one can be selected from the existing grayscale images (RGB, DSM, hillshade, etc.), and the effect of feature extraction cannot be achieved. Summary of the invention
[0004] The embodiment of the present application provides an input grayscale image preprocessing method for improving sub-pixel offset tracking accuracy, and according to the preprocessing method, a displacement field result with better quality can be finally output.
[0005] The input grayscale image preprocessing method provided in the embodiment of the present application comprises the following steps:
[0006] 1) Provide 2k one-dimensional grayscale images derived from drone aerial photography dataset;
[0007] 2) Crop the 2k one-dimensional grayscale images to a uniform size of N×M and superimpose them to obtain two sets of N×M×k multi-band images;
[0008] 3) The matrices generated by the two sets of N×M×k multi-band images are respectively input into a ResNet backbone network to form a dual backbone network, and then the network output results are input into the SoftMax activation function, and the activation function output results are mapped to the unsigned integer range of 0 to 255, and finally an N×M grayscale image is obtained as the final output result;
[0009] 4) Calculate the loss function based on the two N×M grayscale images output to obtain the loss value;
[0010] 5) Back-propagation is performed based on the calculated loss value to correct the ResNet network parameters, and finally a trained sub-pixel offset tracking input grayscale image feature extraction neural network model is obtained.
[0011] In addition, the input grayscale image preprocessing method provided in the embodiment of the present application may also have the following additional technical features:
[0012] In an optional solution, the calculation of the loss function in step 4) is given by the following formula:
[0013] loss = αL MI +βL E +γ
[0014] In the formula, loss represents the loss value, L MI represents the mutual information loss term, L E represents the two-dimensional entropy loss term, and α, β, and γ are hyperparameters.
[0015] In an optional solution, the mutual information loss term is calculated by the following formula:
[0016]
[0017] In the formula, A and B represent two grayscale images, A i represents the number of pixels with gray value i in the gray image A, B j Represents the number of pixels with pixel value j in the grayscale image, A i ∩B j It represents a pixel whose gray value is i in A and j in B. The operator |·| represents the total number, and N is the total number of pixels.
[0018] In an optional solution, the two-dimensional entropy loss term is calculated by the following formula:
[0019]
[0020]
[0021]
[0022] Where i represents the grayscale value of a pixel, j represents the average grayscale value of all neighboring pixels of a pixel, j(k) is the grayscale value of a neighboring pixel of a pixel, n is the number of neighboring pixels of a pixel, the operator |·| represents the total number, and N is the total number of pixels.
[0023] In an optional solution, the 2k one-dimensional grayscale images in step 2) include Red Band, Green Band, Blue Band, DSM, and hillshade; one of the two groups of N×M×k multi-band images corresponds to “pre-event images” and the other corresponds to “post-event images”.
[0024] The beneficial effects of the embodiments of the present application are:
[0025] 1) The input grayscale image preprocessing method uses a deep neural network to automatically extract features from the input grayscale image, which not only avoids the huge workload caused by manual trial and error, but also can achieve a better combination of multiple input grayscale images; 2) The mutual information and two-dimensional entropy are introduced into the network loss calculation, and the optimal input is directly predicted and selected from the perspective of the semantic information of the input grayscale image, without relying on the calculation results of sub-pixel offset tracking; 3) For a study area, only one data set training process is required to obtain a stable network model, and then the input grayscale image data set of the study area can be directly preprocessed using the model, thereby conveniently and quickly realizing the selection of input grayscale images for the sub-pixel offset tracking algorithm and improving the accuracy of sub-pixel offset tracking.
[0026] It should be understood that the foregoing general description and the following detailed description are exemplary only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 Schematic diagram of the network training steps of the input grayscale image preprocessing method provided in this application. DETAILED DESCRIPTION
[0028] In order to better understand the technical solution of the present application, the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0029] It should be clear that the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.
[0030] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms "a", "said" and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.
[0031] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0032] By planning flight routes and PPK positioning, drone aerial photography can produce a large range of surface image data sets. Subsequently, through the Structure-from-Motion (SfM) algorithm, the landslide three-dimensional model based on the point cloud and the derived orthophoto and DSM images can be quickly generated. The sub-pixel offset tracking algorithm (sPOT) provides a convenient method for processing surface images to generate surface deformation fields: it synchronously takes reference windows and search windows on multi-temporal images, iteratively traverses the search window with the reference window and calculates the correlation coefficient of each iteration (normalized cross-correlation is generally used), and takes the largest correlation coefficient in the traversal process as the displacement vector representing the center of the window. The displacement vector of the center of the window is determined by continuously obtaining the window in the entire landslide image with a specified step size. After completion, interpolation is performed based on methods such as C1 continuous approximation of quadratic B-spline curves to generate a global displacement field at the sub-pixel level. At present, such a technical route has been effectively applied in fields such as landslide surface deformation monitoring.
[0033] In the related art, the quality of the input grayscale image is judged only by the quality of the output result after the algorithm calculation is completed, ignoring the semantic information contained in the input grayscale image itself. In fact, the input grayscale image itself contains rich semantic information, which can be associated with the final output result to a certain extent. Simply put, a worse output displacement field result is more likely to correspond to a "worse" input grayscale image. Therefore, the key issue is how to measure the quality of the input grayscale image.
[0034] In order to conveniently and quickly implement the selection of the input grayscale image of the sub-pixel offset tracking algorithm and improve the accuracy of the sub-pixel offset tracking, the embodiment of the present application provides an input grayscale image preprocessing method for improving the sub-pixel offset tracking accuracy, and the input grayscale image preprocessing method specifically includes the following steps:
[0035] 1) Prepare the Red Band, Green Band, Blue Band, DSM,
[0036] There are k kinds of one-dimensional grayscale images such as hillshade (a total of 2k one-dimensional grayscale images, corresponding to k "pre-event images" and k "post-event images" respectively).
[0037] 2) All one-dimensional grayscale images are cropped to a uniform size of N×M and superimposed to obtain two sets of N×M×k multi-band images (corresponding to the “pre-event image” and “post-event image”, respectively).
[0038] 3) Input the two sets of N×M×k matrices into a ResNet backbone network (forming a dual backbone network), then input the network output results into the SoftMax activation function, and then map the activation function output results to the range of unsigned integers (0 to 255), and finally obtain an N×M grayscale image as the final output result.
[0039] 4) Calculate the loss function based on the two N×M grayscale images output. The loss function mainly considers mutual information and
[0040] Entropy is calculated by the following formula:
[0041] loss = αL MI +βL E +γ
[0042] Among them: loss represents the loss value; L MI represents the mutual information loss term; L E represents the two-dimensional entropy loss term; α, β, γ are hyperparameters.
[0043] The quality of the input grayscale image can be observed from the perspective of information theory and entropy. The concept of mutual information in information theory can quantitatively give the degree of loss of correlation of the input grayscale image, while information entropy (two-dimensional entropy) can quantitatively give the information richness of the input grayscale image. In addition, the ideal input grayscale image selection method should be able to realize the combination of multiple input grayscale images, that is, to obtain the optimal solution while allowing the grayscale images to be combined, which requires the implementation of a cyclic iterative algorithm. Deep neural networks are very suitable for this requirement.
[0044] Specifically, for the mutual information loss term L in the above formula MI , calculated by the following formula:
[0045]
[0046] Among them: A and B represent two grayscale images; Ai represents the number of pixels with gray value i in the gray image A; B j Represents the number of pixels with pixel value j in the grayscale image; A i ∩B j It represents a pixel whose gray value is i in A and j in B. The operator |·| represents the total number. N is the total number of pixels.
[0047] For the two-dimensional entropy loss term L E , calculated by the following formula:
[0048]
[0049]
[0050]
[0051] Where: i represents the gray value of a pixel; j represents the average gray value of all the neighboring pixels of a pixel; j(k) is the gray value of a neighboring pixel of a pixel; n is the number of neighboring pixels of a pixel; the operator |·| represents taking the total number; N is the total number of pixels.
[0052] 5) Back-propagation is performed according to the calculated loss value to correct the ResNet network parameters, and finally a trained sub-pixel offset tracking input grayscale image feature extraction neural network model is obtained.
[0053] 6) Next, for any one-dimensional grayscale image data set organized as required in step 1), perform grayscale image preprocessing by performing the operations of steps 2) to 4) above, and you can get a set of M×N preprocessed grayscale images (representing the feature extraction results of "pre-event images" and "post-event images" respectively). This set of preprocessed grayscale images is used as the input image of the sub-pixel offset tracking algorithm, and a displacement field result with good quality can be output. Figure 1 As shown, the above steps 3) to 5) correspond to the training steps of a dual-backbone neural network.
[0054] Example:
[0055] (1) Four phases of drone image data of a landslide in Aba Prefecture, Sichuan Province were processed by SfM software to generate four groups of five grayscale images (including Red Band, Green Band, Blue Band, DSM, and hillshade).
[0056] (2) Two of the four data phases were selected, and the five grayscale images of each phase were stacked using ArcGIS software to obtain two multi-band images. The two multi-band images were then cropped synchronously using ArcGIS software to obtain 625 sets of multi-band image pairs with a size of (25×25) and the same range as the study area for training.
[0057] (3) The 625 sets of multi-band effects are input into the dual backbone network for training, and the loss becomes stable after about 85 epochs.
[0058] (4) Select the remaining 2 phases of data from the 4 phases and repeat the operation in step (2). Input the obtained 625 sets of multi-band images into the network to obtain 625 sets of grayscale images. Input the 625 sets of grayscale images into the sub-pixel migration tracking algorithm and obtain 625 displacement field results. Finally, the 625 displacement field results are stitched and restored according to the geographic coordinates to obtain the displacement field results of the original study area.
[0059] (5) In order to verify the quality of the displacement field result obtained in step (4) of this method, the average gradient is calculated for the displacement field result and the displacement field result obtained by inputting the original 5 grayscale images into the sub-pixel offset tracking algorithm (Table 1 below). It can be seen that the average gradient of the displacement field result obtained by this method is smaller, which proves that the result has less noise and is better.
[0060] Table 1 Average gradient comparison table of output displacement field results
[0061]
[0062] The above are only preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for preprocessing an input grayscale image to improve sub-pixel offset tracking accuracy, characterized in that: The steps include: 1) Provide 2k one-dimensional grayscale images derived from drone aerial photography dataset; 2) Crop the 2k one-dimensional grayscale images to a uniform size of N×M and superimpose them to obtain two sets of N×M×k multi-band images; 3) The matrices generated by the two sets of N×M×k multi-band images are respectively input into a ResNet backbone network to form a dual backbone network, and then the network output results are input into the SoftMax activation function, and the activation function output results are mapped to the unsigned integer range of 0 to 255, and finally an N×M grayscale image is obtained as the final output result; 4) Calculate the loss function based on the two N×M grayscale images output to obtain the loss value; the loss function includes a mutual information loss term and a two-dimensional entropy loss term; The mutual information loss term is calculated by the following formula: In the formula, A and B represent two grayscale images, A i represents the number of pixels with gray value i in the gray image A, B j Represents the number of pixels with pixel value j in the grayscale image, A i ∩B j It represents a pixel whose gray value in A is i and its gray value in B is j. The operator |·| represents the total number, where N is the total number of pixels. The two-dimensional entropy loss term is calculated by the following formula: In the formula, i represents the gray value of a pixel, j represents the average gray value of all neighboring pixels of a pixel, j(k) is the gray value of a neighboring pixel of a pixel, n is the number of neighboring pixels of a pixel, the operator |·| represents taking the total number, and N is the total number of pixels; 5) Back-propagation is performed to correct the ResNet network parameters according to the calculated loss value, and finally a trained sub-pixel offset tracking input grayscale image feature extraction neural network model is obtained; 6) For any one-dimensional grayscale image data set organized as required in step 1), the grayscale image preprocessing is performed by performing the operations of steps 2) to 3) above, and a set of M×N preprocessed grayscale images can be obtained.
2. The input grayscale image preprocessing method for improving sub-pixel offset tracking accuracy according to claim 1, characterized in that: The calculation of the loss function in step 4) is given by the following formula: loss=αL MI +βL E +g In the formula, loss represents the loss value, L MI represents the mutual information loss term, L E represents the two-dimensional entropy loss term, and α, β, and γ are hyperparameters.
3. The input grayscale image preprocessing method for improving sub-pixel offset tracking accuracy according to claim 1, characterized in that: The 2k one-dimensional grayscale images in step 2) include RedBand, GreenBand, BlueBand, DSM, and hillshade; one of the two groups of N×M×k multi-band images corresponds to "pre-event images" and the other group corresponds to "post-event images".
Citation Information
Patent Citations
Earth surface deformation monitoring method under complex terrain
CN115031674A
Small sample remote sensing image water body information extraction method
CN116486273A