A remote sensing image distortion region identification method based on self-supervised learning
Patent Information
- Application Number
- CN202611016472.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-09
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2046-07-09
AI Technical Summary
本发明通过待检测DOM影像自身携带的地理坐标信息、内方位元素和正射纠正模型参数反算原始影像位置,并在此基础上生成原始影像位置连续性场,使拉花变形识别不再仅依赖DOM影像表面的灰度或纹理变化,而是将像素拉伸、方向变化、距离变化和位置连续断裂关系共同纳入判断过程,从而能够更准确地反映拉花变形区域在正射生产过程中的几何异常特征。
Smart Images

Figure CN122530276B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of remote sensing image processing technology, and in particular to a method for identifying distortion regions in remote sensing images based on self-supervised learning. Background Technology
[0002] During the orthorectification process of remote sensing images, due to the influence of imaging geometry, terrain undulation, orthorectification parameter errors, and image resampling processing, DOM images may exhibit local pixel stretching, abrupt changes in texture direction, and broken positional relationships between adjacent areas, resulting in distorted regions. Current remote sensing image quality inspection typically identifies distorted regions in DOM images through manual visual inspection, fixed threshold screening, texture feature analysis, or anomaly detection based on inverse calculations of relationships from the original image, and outputs corresponding raster or vector detection results.
[0003] The methods described above can detect some geometric or textural anomalies. However, in actual production scenarios, the shape, scale, and texture of the deformed areas vary greatly. Relying solely on a single geometric inverse calculation feature or a single texture feature makes them susceptible to changes in ground texture, shadow edges, and local grayscale abrupt changes. When relying on manually labeled samples to train the recognition model, the sample construction cost is high, and it is difficult to cover the deformed appearance in different regions and under different imaging conditions. At the same time, existing methods do not make sufficient use of the spatial continuity between adjacent image blocks, and the production review results are difficult to feed back into the model training process in a timely manner. As a result, there is still room for improvement in the detection results in terms of boundary localization, false detection correction, and false negative supplementation.
[0004] Therefore, how to provide a method for identifying distortion regions in remote sensing images based on self-supervised learning is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] One objective of this invention is to propose a remote sensing image distortion region identification method based on self-supervised learning. This invention fully utilizes the original image location inverse calculation of DOM images, self-supervised sample construction, shallow endogenous feature extraction, deep spatial continuity analysis, and satellite image production review and re-injection technology. It details the complete implementation process from the generation of the original image location continuity field, the extraction of endogenous anomaly features of image blocks, the construction of self-supervised positive and negative sample sets, the training of the distortion recognition model, to the output of distortion region results. It has the advantages of reducing dependence on manual annotation, improving the accuracy of distortion recognition, enhancing adaptability to complex texture scenes, and improving the utilization efficiency of production review results.
[0006] A method for identifying distortion regions in remote sensing images based on self-supervised learning according to an embodiment of the present invention includes the following steps: Acquire the DOM image to be detected and generate the original image position continuity field based on the geometric parameters of the DOM image itself; An endogenous anomaly feature set of image blocks is generated based on the positional continuity field of the original image and the texture information of the DOM image. Generate a self-supervised positive and negative sample set based on the endogenous anomaly feature set of image blocks; A paper-cutting deformation recognition network is constructed based on a self-supervised positive and negative sample set, which is a fusion of shallow endogenous feature branches and deep spatial continuity branches; A network for recognizing paper-cutting deformations was trained using a self-supervised positive and negative sample set and neighborhood continuity constraints to generate an initial paper-cutting deformation recognition model. The initial pattern deformation recognition model is embedded into the satellite image production workflow to generate a self-reinforcing pattern deformation recognition model. By utilizing a self-reinforcing latte art deformation recognition model and an endogenous anomaly feature set of image blocks, a comprehensive latte art discrimination map of the DOM image to be identified is generated; Based on the integrated latte art discrimination map and the positional continuity field of the original image, the latte art deformation area results are output.
[0007] Optionally, the step of acquiring the DOM image to be detected and generating the original image position continuity field based on the DOM image's own geometric parameters specifically includes: Obtain the DOM image to be detected and generate a set of geometric parameters of the DOM image itself from the image metadata; Based on the geometric parameter set of the DOM image itself, the pixel row and column positions of each DOM pixel are converted into corresponding ground coordinates to obtain the ground coordinate field; Input the ground coordinates in the ground coordinate field into the inverse calculation relationship corresponding to the orthorectification model parameters, and combine the interior orientation elements to inversely calculate the original image position corresponding to each DOM pixel point, and obtain the original image position coordinate field. Read the original image position coordinates corresponding to adjacent DOM pixels, calculate the displacement direction and displacement distance between adjacent original image positions, and obtain the set of inverse displacement relationships; By comparing the directional and distance differences between adjacent back-calculated displacement relationships, the directional change and distance change are obtained respectively; The changes in direction and distance are normalized, and the normalized changes in direction and distance are fused into a continuous break, thus constructing the positional continuity field of the original image.
[0008] Optionally, the generation of the image patch endogenous anomaly feature set based on the original image location continuity field and DOM image texture information specifically includes: The image block side length is determined according to the resolution of the DOM image to be detected, and the positional continuity field of the DOM image to be detected and the original image is divided by the boundary of the same image block to obtain a set of image blocks at the same position. Read the continuous fracture data within the coverage area of each image block in the image block set at the same location, and statistically obtain the set of continuous fracture intensities at the location. Read the direction change data and distance change data within the coverage area of each image block in the image block set at the same location, and calculate the pixel stretching ratio set by combining the location continuous fracture intensity set. Read the grayscale data of the DOM images of each image block in the image block set at the same location, count the grayscale change relationship in different directions, and combine it with the set of pixel stretching ratios to obtain the texture mutation feature set. Perform gray-level co-occurrence matrix statistics on DOM image gray-level data and extract a set of abnormal features from the gray-level co-occurrence matrix; The set of location-continuous fracture intensity, pixel stretching ratio, texture abrupt change feature set, and gray-level co-occurrence matrix anomaly feature set are combined according to the spatial location of the image block to obtain the image block endogenous anomaly feature set.
[0009] Optionally, the generation of a self-supervised positive and negative sample set based on the endogenous anomaly feature set of image blocks specifically includes: Geometric continuity anomaly information and texture mutation information corresponding to each image block are extracted from the endogenous anomaly feature set of the image blocks to form a candidate sample set; Normalize and statistically analyze the geometric continuity anomaly information and texture mutation information in the candidate sample set to obtain a set of joint distribution characterization values; The sample separation boundary is determined based on the distribution differences of each candidate sample in the joint distribution characterization value set; By using the sample separation boundary to screen the endogenous anomaly feature set of image blocks, an initial self-supervised positive and negative sample set is obtained; The initial self-supervised positive and negative sample set is subjected to rotation, translation, and grayscale perturbation to obtain the enhanced sample set; By integrating the initial self-supervised positive and negative sample set and the augmented sample set, a self-supervised positive and negative sample set is obtained.
[0010] Optionally, the construction of the paper-cutting deformation recognition network based on the self-supervised positive and negative sample set, which fuses shallow endogenous feature branches and deep spatial continuity branches, specifically includes: Extract the central image block, neighboring image blocks, endogenous anomaly features of image blocks, and sample labels from the self-supervised positive and negative sample set and the image block set at the same location to form the network input sample group; Input the network input sample group into the shallow endogenous feature branch, extract the local geometric anomaly features and texture anomaly features of the central image block, and obtain the shallow endogenous feature vector; Input the network input sample group into the deep spatial continuity branch, extract the spatial continuity difference between the central image block and the neighboring image blocks, and obtain the deep spatial continuity feature vector. The shallow endogenous feature vector and the deep spatial continuity feature vector are input into the feature fusion layer to obtain the discriminative features for latte art deformation. Input the discriminant features of the latte art deformation into the classification output layer to obtain the latte art deformation recognition output; By connecting the shallow endogenous feature branch, the deep spatial continuity branch, the feature fusion layer, and the classification output layer, a paper-cutting deformation recognition network is obtained.
[0011] Optionally, the step of inputting the network input sample group into the deep spatial continuity branch, extracting the spatial continuity difference between the central image patch and the neighboring image patches, and obtaining the deep spatial continuity feature vector specifically includes: Read the central image block and neighboring image blocks from the network input sample group, and arrange them according to their spatial location in the neighborhood to obtain the central neighboring image block sequence; Embedding mapping is performed on each image block in the central neighborhood image block sequence to obtain the image block embedding feature sequence; According to the spatial position of each image block relative to the central image block, the adjacent direction position coding sequence is obtained; The image patch embedding feature sequence and the adjacent direction position encoding sequence are fused together to obtain the Transformer input feature sequence; Input the Transformer input feature sequence into the multi-head attention layer to obtain the center neighborhood attention features; Calculate texture continuity, boundary connection, and neighborhood spatial continuity differences based on the attention features of the central neighborhood; The texture continuity relationship, boundary connection relationship and neighborhood spatial continuity difference are combined into a deep spatial continuity feature vector.
[0012] Optionally, the step of training the latte art deformation recognition network using a self-supervised positive and negative sample set and neighborhood continuity constraints to generate an initial latte art deformation recognition model specifically includes: The model training sample sequence is obtained by organizing the supervised positive and negative sample set and the network input sample set according to the sample labels. Input the training sample sequence of the model into the latte art deformation recognition network, output the latte art deformation recognition result, and calculate the classification loss according to the difference between the latte art deformation recognition result and the sample label; The distance between the discriminant features of the same type of samples corresponding to the latte art deformation is statistically analyzed to obtain the feature aggregation loss of the same type of samples; The continuous breakage amount in the location continuity field of the original image is read and correspondingly constrained with the deep spatial continuity feature vector to obtain the neighborhood continuity constraint loss. The classification loss, the feature aggregation loss of similar samples, and the neighborhood continuity constraint loss are weighted and summed to obtain the joint loss; The network parameters of the shallow endogenous feature branch, feature fusion layer, and classification output layer are updated according to the joint loss, and the initial paper-cutting deformation recognition model is obtained using the updated network parameters.
[0013] Optionally, the step of embedding the initial latte art deformation recognition model into the satellite image production workflow to generate a self-reinforcing latte art deformation recognition model specifically includes: The initial pattern deformation recognition model is embedded into the satellite image production workflow, and pattern deformation recognition is performed on DOM images during the production process to obtain a set of production recognition results. The set of production identification results is compared with the results of manual review to obtain a set of review difference samples; Extract the DOM image grayscale data and endogenous anomaly features of the corresponding image blocks from the set of review difference samples, and assign sample labels according to the manual review results to obtain a new training sample set; The confidence bias is obtained by calculating the difference between the recognition probability of each newly added training sample in the newly added training sample set and the manually verified label. The training weights of the samples are obtained by combining the confidence bias and the sample types of verification differences; The sample training weights are added to the fine-tuning training process of the newly added training sample set, and the confidence-weighted fine-tuning loss is calculated. The initial latte art deformation recognition model is updated by adjusting the loss using confidence weighting, resulting in a self-reinforcing latte art deformation recognition model.
[0014] Optionally, the step of generating a comprehensive discriminant map of the DOM image to be identified using a self-reinforcing latte art deformation recognition model and an endogenous anomaly feature set of image blocks specifically includes: Perform original image position inverse calculation and same-position block processing on the DOM image to be identified to obtain the original image position continuity field and the set of same-position image blocks to be identified. Based on the original image location continuity field of the DOM image to be identified and the set of image blocks at the same location to be identified, the endogenous anomaly feature set of the image blocks to be identified is obtained. The set of image blocks at the same location to be identified, the set of endogenous abnormal features of the image blocks to be identified, and the self-reinforcing lace deformation identification model are input into the identification calculation process to obtain the lace deformation probability and the difference in the continuity of the neighborhood space. The grayscale data of DOM images in the set of image blocks at the same location to be identified are statistically analyzed to obtain the texture information entropy; Based on the validation sample set, determine the fusion weights corresponding to the probability of latte art deformation, pixel stretching ratio, positional continuity break strength, texture information entropy, and neighborhood spatial continuity difference, and calculate the comprehensive latte art discriminant value. The comprehensive stylus discriminant value is written into the layer according to the spatial position of the set of image blocks at the same location to be identified, so as to obtain the comprehensive stylus discriminant image of the DOM image to be identified.
[0015] Optionally, the step of outputting the result of the latte art deformation region based on the comprehensive latte art discrimination map and the positional continuity field of the original image includes: filtering latte art deformation image blocks according to the comprehensive latte art discrimination value in the comprehensive latte art discrimination map and the positional continuity fracture intensity in the positional continuity field of the original image; performing connectivity merging and boundary optimization on the latte art deformation image blocks; generating a binary raster mask and vector boundary file to obtain the result of the latte art deformation region.
[0016] The beneficial effects of this invention are: This invention calculates the original image position by using the geographic coordinate information, interior orientation elements, and orthorectification model parameters carried by the DOM image to be detected. Based on this, it generates a continuity field of the original image position, so that the recognition of latte art deformation no longer depends solely on the grayscale or texture changes on the surface of the DOM image. Instead, it incorporates pixel stretching, orientation changes, distance changes, and positional continuity breakage relationships into the judgment process, thereby more accurately reflecting the geometric anomaly characteristics of the latte art deformation area during the orthorectification production process.
[0017] This invention generates a self-supervised positive and negative sample set without manual annotation based on the endogenous anomaly feature set of image blocks. The sample separation boundary is determined by the joint distribution of positional continuous fracture intensity, pixel stretching ratio, texture anisotropy, and pixel gradient mutation value. It also combines active masking, rotation, translation, and grayscale perturbation to form an enhanced sample set, reducing the dependence on manually annotated samples. This allows the model training samples to come from the geometric and texture anomaly distribution of the DOM image to be detected, which is beneficial for adapting to the needs of pattern deformation recognition in different regions, different ground textures, and under different imaging conditions.
[0018] This invention constructs a pattern deformation recognition network that fuses a shallow endogenous feature branch and a deep spatial continuity branch. The shallow endogenous feature branch is used to extract local geometric anomalies and texture anomalies, while the deep spatial continuity branch is used to extract the texture continuity relationship, boundary connection relationship, and spatial continuity difference between the central sample image block and neighboring image blocks. This allows the recognition result to consider both the anomalies within a single image block and the spatial continuity relationship between adjacent image blocks, thereby reducing misjudgments caused by local shadows, ground object edges, and gray-scale abrupt changes.
[0019] This invention uses the results of satellite image production verification to perform confidence-weighted fine-tuning on the initial pattern deformation recognition model, transforming falsely detected image blocks, missed image blocks, boundary-corrected image blocks, and normal image blocks into new training samples. It also generates sample training weights based on the difference between the recognition probability and the manual verification label, enabling the production verification results to be fed back into the model update process, which is beneficial to improving the model's continuous adaptability in the actual satellite image production process.
[0020] This invention outputs the results of the deformed area based on the comprehensive discriminant image and the positional continuity field of the original image. It determines the deformation boundary by the maximum segmentation result of the interclass variance of the comprehensive discriminant value, and filters the deformed image blocks by combining the judgment condition formed by the median of the positional continuity fracture intensity and the interquartile range. Then, it performs connectivity merging, isolated block removal, boundary smoothing and vectorization processing, which can output a binary raster mask and vector boundary file that are consistent with the spatial range of the DOM image to be identified, which is convenient for subsequent quality review and production result management. Attached Figure Description
[0021] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a remote sensing image distortion region identification method based on self-supervised learning proposed in this invention; Figure 2 This is a schematic diagram of the pattern deformation recognition network structure of the present invention; Figure 3 This is a schematic diagram of the result of the latte art deformation area under a complex shadow background in an embodiment of the present invention, wherein the left image is a local area of the original DOM image, and the right image is an overlay of the result of the latte art deformation area; Figure 4 This is a schematic diagram of the result of the latte art deformation area under the background of complex terrain texture in an embodiment of the present invention. The left image is a local area of the original DOM image, and the right image is an overlay of the result of the latte art deformation area. Detailed Implementation
[0022] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0023] refer to Figure 1 A method for identifying distortion regions in remote sensing images based on self-supervised learning includes the following steps: Acquire the DOM image to be detected and generate the original image position continuity field based on the geometric parameters of the DOM image itself; An endogenous anomaly feature set of image blocks is generated based on the positional continuity field of the original image and the texture information of the DOM image. Generate a self-supervised positive and negative sample set based on the endogenous anomaly feature set of image blocks; A paper-cutting deformation recognition network is constructed based on a self-supervised positive and negative sample set, which is a fusion of shallow endogenous feature branches and deep spatial continuity branches; A network for recognizing paper-cutting deformations was trained using a self-supervised positive and negative sample set and neighborhood continuity constraints to generate an initial paper-cutting deformation recognition model. The initial pattern deformation recognition model is embedded into the satellite image production workflow to generate a self-reinforcing pattern deformation recognition model. By utilizing a self-reinforcing latte art deformation recognition model and an endogenous anomaly feature set of image blocks, a comprehensive latte art discrimination map of the DOM image to be identified is generated; Based on the integrated latte art discrimination map and the positional continuity field of the original image, the latte art deformation area results are output.
[0024] In this embodiment, acquiring the DOM image to be detected and generating the original image position continuity field based on the DOM image's own geometric parameters specifically includes: Obtain the DOM image to be detected and generate a set of geometric parameters of the DOM image itself from the image metadata, specifically including: Read the pixel row and column information, geographic coordinate transformation information, pixel resolution, interior orientation elements and orthorectification model parameters of the DOM image to be detected; The above information is organized according to the correspondence between pixel position, ground position and original imaging position to obtain the set of geometric parameters of the DOM image itself that can support the inverse calculation of the original image position; The ground coordinates of each DOM pixel are calculated based on the geometric parameter set of the DOM image itself, and a ground coordinate field is generated. Specifically, this includes: The pixel position of each DOM pixel is read according to the pixel row and column order of the DOM image to be detected, and the planar coordinates of each DOM pixel are determined by combining the geographic coordinate transformation information and the pixel resolution. By combining the elevation reduction information in the orthorectification model parameters, the elevation coordinates of each DOM pixel are determined, and the ground coordinate field corresponding to the pixel row and column positions of the DOM image to be detected is obtained. Based on the ground coordinate field and interior orientation elements, the original image position corresponding to each DOM pixel is calculated and the original image position coordinate field is generated. Specifically, this includes: Input the ground coordinates corresponding to each DOM pixel into the inverse calculation relationship corresponding to the orthorectification model parameters; By combining the principal point coordinates, focal length parameters, and imaging coordinate reference points of the original image, the position coordinates of each DOM pixel in the original image are determined; Generate the original image position coordinate field according to the pixel row and column positions of the DOM image to be detected; The inverse displacement relationship between adjacent DOM pixels is calculated based on the original image position coordinate field, and a set of inverse displacement relationships is generated, specifically including: Read the original image position coordinates corresponding to adjacent DOM pixels in the horizontal and vertical directions respectively; Determine the displacement direction and displacement distance between adjacent original image position coordinates, and generate a set of inverse displacement relationships according to the pixel row and column positions of the DOM image to be detected; The calculation of directional and distance changes between adjacent inverse displacement relationships is based on the set of inverse displacement relationships, specifically including: By comparing the angle changes between the inverse displacement directions of adjacent positions, the change in direction is obtained; by comparing the length differences between the inverse displacement distances of adjacent positions, the change in distance is obtained. The orientation change and distance change are written into the orientation change channel and distance change channel respectively according to the pixel row and column positions of the DOM image to be detected; Based on the fusion of direction change and distance change, continuous fracture data is generated and the original image location continuity field is constructed, specifically including: The changes in direction and distance are normalized within the same image. The normalized direction change and distance change are weighted and synthesized according to the corresponding pixel positions to obtain the continuity break amount, which represents the degree of destruction of the positional continuity of the original image. The inter-class separability of the deformed latte art samples and normal samples in the sample set was statistically verified, and the separability of the direction change was obtained. The inter-class separability of the deformed latte art samples and normal samples in the sample set was statistically verified by measuring the normalized distance change, and the distance change separability was obtained. Divide the separability of the direction change by the sum of the separability of the direction change and the separability of the distance change to obtain the weight of the direction change. Divide the separability of distance change by the sum of the separability of direction change and the separability of distance change to obtain the distance change weight; Multiply the normalized direction change by the direction change weight, multiply the normalized distance change by the distance change weight, and sum the two to obtain the continuous breakage amount of the corresponding pixel. By writing the orientation change channel, distance change channel, and continuous breakage amount into the multi-channel raster data within the same spatial range, the original image location continuity field is generated.
[0025] In this embodiment, generating an image patch endogenous anomaly feature set based on the original image location continuity field and DOM image texture information specifically includes: Based on the resolution of the DOM image to be detected and the positional continuity field of the original image, a set of image patches at the same location is generated, specifically including: The image block side length is determined based on the ground resolution of the DOM image to be detected, and the DOM image to be detected, the orientation change channel, the distance change channel and the continuous break channel are divided according to the boundary of the same image block; The side length of the image block is determined based on the ground resolution of the DOM image to be detected and the width of the minimum identifiable graffiti deformation area; Divide the width of the smallest recognizable lace deformation area by the ground resolution of the DOM image to be detected to obtain the minimum pixel side length; The minimum pixel side length is adjusted to an integer side length that facilitates convolution processing and neighborhood partitioning, and used as the image block side length. In one implementation, the ground resolution of the DOM image to be detected is 0.5m, the width of the minimum identifiable lace deformation area is 250m to 260m, and the side length of the image block is 512 pixels, so that the ground area covered by a single image block is approximately 256m × 256m. Each image block has corresponding DOM image grayscale data, orientation change data, distance change data, and continuous break data; Based on the continuous fracture data in the same image block set, a set of location continuous fracture intensity for each image block is generated, specifically including: For any image block, read the continuous breakage of all pixels within its coverage area, and calculate the average value of the continuous breakage within the same image block. The obtained positional continuous fracture intensity is written into the positional continuous fracture intensity set according to the spatial location of the image block. A set of pixel stretching ratios for each image patch is generated based on orientation change data, distance change data, and location continuity break intensity data from the set of image patches at the same location. Specifically, this includes: For any image block, calculate the inverse displacement distance of the horizontally adjacent pixels and the inverse displacement distance of the vertically adjacent pixels within its coverage area; Calculate the average displacement distance between horizontally adjacent pixels and the average displacement distance between vertically adjacent pixels, respectively. Divide the larger of the two averages by the smaller of the two averages to obtain the pixel stretching ratio of the image block; Write the pixel stretching ratio and the positional continuous fracture intensity of the image block into the pixel stretching ratio set. Based on the grayscale data of DOM images and the set of pixel stretching ratios in the set of image patches at the same location, a set of texture abrupt change features for each image patch is generated, specifically including: For any image block, the degree of grayscale difference divergence is calculated along the horizontal texture direction, the vertical texture direction, and the two diagonal texture directions respectively. Texture anisotropy is obtained by the maximum difference in the degree of dispersion of grayscale difference between each texture direction; The Sobel operator is used to calculate the horizontal and vertical gradient values of each pixel within the image block. The average gradient magnitude of each pixel is calculated to obtain the pixel gradient abrupt value. The texture anisotropy and pixel gradient mutation values are written into the texture mutation feature set in correspondence with the pixel stretching ratio of the image block; Based on the grayscale data of DOM images and the texture mutation feature set in the same location image patch set, a grayscale co-occurrence matrix anomaly feature set for each image patch is generated, specifically including: For any image block, its grayscale value is graded and quantized, and the probability distribution of different grayscale levels co-occurring in adjacent pixel positions is statistically analyzed to form a grayscale co-occurrence matrix. Calculate the contrast, correlation, and energy of the gray-level co-occurrence matrix based on the gray-level co-occurrence matrix; The gray-level co-occurrence matrix contrast, gray-level co-occurrence matrix correlation, and gray-level co-occurrence matrix energy are written into the gray-level co-occurrence matrix abnormal feature set in correspondence with the texture anisotropy of image blocks and pixel gradient abrupt values. An endogenous anomaly feature set for image patches is generated based on a set of location-continuous fracture intensity, a set of pixel stretching ratios, a set of texture abrupt change features, and a set of gray-level co-occurrence matrix anomaly features. Specifically, it includes: For any image block, the spatial location, location continuity break intensity, pixel stretching ratio, texture anisotropy, pixel gradient abrupt value, gray-level co-occurrence matrix contrast, gray-level co-occurrence matrix correlation, and gray-level co-occurrence matrix energy of the image block are combined to form the endogenous anomalous feature record of the image block. All endogenous anomaly feature records of all image blocks are written into the same data set according to the spatial location of the image blocks to obtain the endogenous anomaly feature set of the image blocks.
[0026] In this embodiment, generating a self-supervised positive and negative sample set based on the endogenous anomaly feature set of image blocks specifically includes: A candidate sample set carrying information on geometric continuity anomalies and texture mutations is generated based on the endogenous anomaly feature set of image patches, specifically including: For any image block, the grayscale data of the DOM image corresponding to the image block is used as a candidate sample image block; The continuous fracture intensity, pixel stretching ratio, texture anisotropy and pixel gradient abrupt value of the corresponding image block are combined into the candidate sample feature vector. Write the candidate sample image block, candidate sample feature vector, and corresponding spatial location of the image block into the candidate sample set; A set of joint distribution characterization values reflecting the degree of joint deviation from the anomaly in latte art is generated based on the candidate sample set, specifically including: Statistical analysis of the overall mean and feature co-variation relationship of feature vectors of all candidate samples within the same DOM image to be detected; Based on the degree of deviation of each candidate sample feature vector from the overall mean and the synergistic change relationship between the features, a joint distribution characterization value is generated that simultaneously reflects positional continuity breaks, pixel stretching, texture anisotropy, and pixel gradient abrupt changes. All joint distribution characterization values are written into the joint distribution characterization value set according to the spatial location of the candidate samples; Based on the joint distribution characterization value set, a sample separation boundary is generated to distinguish between normal samples and latte art deformation samples, specifically including: The candidate segmentation positions are traversed in ascending order of the joint distribution representation values; At each candidate segmentation location, the candidate sample set is divided into a low-abnormality candidate sample group and a high-abnormality candidate sample group; Calculate the sample proportions of the low-anomaly candidate sample group and the high-anomaly candidate sample group, as well as the concentration of the joint distribution characterization value; The candidate segmentation position that makes the greatest difference between the low-abnormality candidate sample group and the high-abnormality candidate sample group is determined as the sample separation boundary. The candidate segmentation positions are traversed in ascending order of joint distribution characterization value. At each candidate segmentation position, the candidate samples with joint distribution characterization value lower than that candidate segmentation position are classified into the low anomaly candidate sample group. Candidate samples whose joint distribution characterization value is not lower than the candidate segmentation position are classified into the high anomaly candidate sample group; Calculate the sample proportion of the low-abnormality candidate sample group, the sample proportion of the high-abnormality candidate sample group, the average value of the joint distribution characterization value of the low-abnormality candidate sample group, and the average value of the joint distribution characterization value of the high-abnormality candidate sample group respectively; Multiply the sample proportions of the low-abnormality candidate sample group, the sample proportions of the high-abnormality candidate sample group, and the squared difference between the average values of the two sample groups to obtain the inter-class variance corresponding to the candidate segmentation position. After traversing all candidate segmentation positions, the candidate segmentation position with the largest inter-class variance is determined as the sample separation boundary; According to the joint distribution characterization value sorted from small to large, the candidate samples in the low distribution segment are used as normal candidate samples, the candidate samples in the high distribution segment are used as latte art deformation candidate samples, and the candidate samples in the middle distribution segment are excluded from the initial training samples. An initial self-supervised set of positive and negative samples with dual constraints on geometric texture is generated based on the sample separation boundary and the endogenous anomaly feature set of image patches. Specifically, it includes: For any candidate sample, if the joint distribution characterization value is lower than the sample separation limit, the positional continuity fracture intensity is lower than the median value of the positional continuity fracture intensity in the same DOM image to be detected, and the texture anisotropy is lower than the median value of the texture anisotropy in the same DOM image to be detected, the candidate sample is labeled as a normal sample. When the joint distribution characterization value is not lower than the sample separation limit, the positional continuous fracture intensity is not lower than the median value of the positional continuous fracture intensity within the same DOM image to be detected, the pixel stretching ratio is not lower than the median value of the pixel stretching ratio within the same DOM image to be detected, and the texture anisotropy is not lower than the median value of the texture anisotropy within the same DOM image to be detected, the candidate sample is labeled as a lace deformation sample. All normal samples and latte art deformation samples are written into the initial self-supervised positive and negative sample set; An enhanced set of classes-preserving samples is generated based on the initial self-supervised positive and negative sample set, specifically including: For any normal sample and any deformed latte art sample, a local area is selected in the sample image block for active masking to obtain an active masked sample. Perform a rotation transformation on the sample image block to obtain a rotated sample; Perform a horizontal or vertical translation transformation on the sample image block to obtain a translated sample; Brightness adjustment and grayscale shift are performed on the sample image block to obtain grayscale perturbation sample; The active masking occlusion area is determined based on the side length of the sample image block and the location of local texture abrupt changes, with the occlusion area ranging from 5% to 20% of the sample image block area. The occlusion location is selected first from local areas with high pixel gradient abrupt values, and the occlusion area is filled with the average gray value within the same image block. The rotation transformation uses three angles: 90°, 180°, and 270°, without changing the category attributes of normal samples and latte art deformed samples; The translation transformation is performed in the horizontal or vertical direction, and the translation distance is 5% to 15% of the side length of the sample image block. The blank area after translation is filled with the boundary pixel mirror. The brightness adjustment range is determined based on the average gray value of the same DOM image to be detected. The gray value offset range is 5% to 10% of the standard deviation of gray values of the same DOM image to be detected. The pixel values after gray value perturbation are limited to the gray value range of the DOM image. The active mask samples, rotation samples, translation samples, and grayscale perturbation samples are written into the enhanced sample set after retaining the original sample calibration results; Based on the augmented sample set, a self-supervised positive and negative sample set with controlled sample ratios and no manual annotation is generated, specifically including: Normal samples and their corresponding active masked samples, rotated samples, translated samples, and grayscale perturbation samples are grouped into the negative sample subset. The deformed samples and their corresponding active mask samples, rotation samples, translation samples and grayscale perturbation samples are classified into the positive sample subset; The augmented sample set is repeatedly extracted or removed based on the ratio of the number of positive sample subsets to negative sample subsets; The negative sample subset, the positive sample subset, and the corresponding sample labels are written together into a self-supervised positive and negative sample set that does not require manual annotation.
[0027] refer to Figure 2In this embodiment, the construction of a paper-cutting deformation recognition network based on a self-supervised positive and negative sample set, which integrates shallow endogenous feature branches and deep spatial continuity branches, specifically includes: Among them, the shallow endogenous feature branch includes a self-learning convolutional network, which is used to receive the DOM image grayscale data and endogenous abnormal features of the image block of the center sample image block. Local geometric and textural anomaly features of the central sample image block are extracted through convolution, pooling, and feature mapping, and shallow endogenous feature vectors are output. Based on a self-supervised positive and negative sample set that requires no manual annotation and a set of image patches at the same location, a network input sample set carrying neighborhood spatial relationships is generated, specifically including: For any central sample image block, read the DOM image grayscale data, image block spatial location, sample label, and neighboring image blocks that have a spatial adjacency relationship with the central sample image block. Read the continuous fracture intensity, pixel stretching ratio, texture anisotropy, pixel gradient abrupt value, gray-level co-occurrence matrix contrast, gray-level co-occurrence matrix correlation and gray-level co-occurrence matrix energy corresponding to the central sample image block; The central sample image block, neighboring image blocks, image block spatial location, sample label, and the aforementioned endogenous anomaly features are all written into the network input sample group. A deep spatial continuity branch is constructed based on the network input sample group to generate a deep spatial continuity feature vector; A deep spatial continuity branch is constructed based on the network input sample set to extract the differences in neighborhood spatial continuity, specifically including: Input the central sample image block and its neighboring image blocks into the pre-trained visual Transformer feature extraction model in spatial adjacency order; The central sample image block and the neighboring image blocks are converted into image block feature sequences through block embedding processing; The central sample image block, together with its upper neighbor, lower neighbor, left neighbor, right neighbor, and four diagonal neighbor image blocks, forms the central neighborhood image block sequence, which includes a total of 9 image blocks. The grayscale data of the DOM image of each image block is first adjusted to a uniform input size, and then divided into blocks according to a fixed block size. After flattening the grayscale data of each block, it is converted into image block embedding features through block embedding mapping. The adjacent direction position encoding is generated based on the position type of the image block relative to the center sample image block. The position types include center position, upper adjacent position, lower adjacent position, left adjacent position, right adjacent position and diagonal adjacent position. The image patch embedding features are added to the adjacent direction position encoding features to form the Transformer input feature sequence; The backbone parameters of the pre-trained visual Transformer feature extraction model are used for forward feature computation and remain unchanged during initial training and confidence-weighted fine-tuning. The self-learning convolutional network, feature fusion layer, and classification output layer in the shallow endogenous feature branch are trained parts that participate in parameter updates. The vertical, horizontal, and diagonal adjacency relationships between the central sample image block and its neighboring image blocks are written into the image block feature sequence through positional encoding processing. In multi-head attention processing, image patch feature sequences are mapped to query features, key features, and value features, respectively. Attention allocation results are generated based on the similarity between query features and key features; The value features are weighted and converged using the attention allocation results to obtain attention features that reflect the differences in texture continuity, boundary connection, and spatial continuity between the central sample image block and its neighboring image blocks; The attention features output by each attention head are concatenated and fed forward, and then processed through feature convergence to generate a deep spatial continuity feature vector. A feature fusion layer driven by the intensity of endogenous anomalies is constructed based on shallow endogenous feature vectors and deep spatially continuous feature vectors, specifically including: The positional continuity fracture intensity, pixel stretching ratio, texture anisotropy, and pixel gradient abrupt value of the central sample image block are normalized according to the value range in the same batch of network input sample groups. The shallow fusion weights are calculated based on the normalized positional continuity fracture strength and pixel stretching ratio. The average of the normalized positional continuous fracture intensity and the normalized pixel stretching ratio is used to obtain the shallow anomaly score. The deep anomaly score is obtained by averaging the normalized texture anisotropy, the normalized pixel gradient abruptness value, and the normalized neighborhood spatial continuity difference. The shallow anomaly score is divided by the sum of the shallow anomaly score and the deep anomaly score to obtain the shallow fusion weight. The deep fusion weights are calculated based on the normalized texture anisotropy, pixel gradient abrupt change values, and differences in neighborhood spatial continuity. The shallow fusion weights and deep fusion weights are normalized so that their sum is a single unit weight. If the sum of the shallow anomaly score and the deep anomaly score is zero, then both the shallow fusion weight and the deep fusion weight are set to 0.5. The shallow endogenous feature vector is multiplied by the shallow fusion weight, the deep spatial continuity feature vector is multiplied by the deep fusion weight, and then the vectors are spliced and mapped to obtain the latte art deformation discrimination feature. Multiply the shallow endogenous feature vector by the shallow fusion weight, and multiply the deep spatial continuity feature vector by the deep fusion weight. Then, the weighted shallow endogenous feature vector and the deep spatial continuity feature vector are concatenated and mapped to generate a pattern deformation discrimination feature that combines local anomalies and neighborhood continuity. A classification output layer is constructed based on the discriminative features of latte art deformation, and a latte art deformation recognition output is generated, specifically including: The discriminant features for latte art deformation are input into the classification output layer, and classification scores for normal samples and latte art deformation samples are generated through fully connected mapping. Normalize the classification scores of normal samples and latte art deformed samples to obtain the output probability that the central sample image block belongs to normal samples and the output probability that it belongs to latte art deformed samples. Write the output probability of normal samples, the output probability of deformed latte art samples, and the sample label of the central sample image block into the latte art deformation recognition output. A network for recognizing paper-cut patterns is generated based on shallow endogenous feature branches, deep spatial continuity branches, feature fusion layers, and classification output layers. Specifically, it includes: The shallow endogenous feature branch is set to receive the central sample image block and endogenous anomaly features and output the shallow endogenous feature vector; The deep spatial continuity branch is set to receive the central sample image block and its neighboring image blocks and output the deep spatial continuity feature vector. The feature fusion layer is set to perform weighted fusion of shallow endogenous feature vectors and deep spatial continuity feature vectors based on the intensity of endogenous anomalies and output the latte art deformation discrimination feature. The classification output layer is set to receive the discriminative features of the latte art deformation and output the latte art deformation recognition output, resulting in a latte art deformation recognition network that is fused from the shallow endogenous feature branch and the deep spatial continuity branch.
[0028] In this embodiment, the construction of a deep spatial continuity branch based on the network input sample group to generate a deep spatial continuity feature vector specifically includes: Generate a sequence of central neighborhood image patches carrying central neighborhood spatial relationships based on the network input sample group, specifically including: For any central sample image block, read the central sample image block and its spatially adjacent neighboring image blocks; The central sample image block and the neighboring image block are sorted according to the spatial order of the center position, the upper adjacent position, the lower adjacent position, the left adjacent position, the right adjacent position, and the diagonal adjacent position to obtain the central neighboring image block sequence used for inputting the deep spatial continuity branch. Image patch embedding feature sequences are generated based on the central neighborhood image patch sequence, specifically including: For any image block in the central neighborhood image block sequence, read the DOM image grayscale data of the image block; Convert the grayscale data of the DOM image into grayscale feature data according to the pixel arrangement order; By linearly mapping and biasing the grayscale feature data through block embedding mapping, the image block embedding features corresponding to the image block are obtained. All image block embedding features are written into the image block embedding feature sequence according to the spatial order of the central neighborhood image block sequence; Based on the image patch embedding feature sequence and the spatial location of the image patch, an adjacent direction position encoding sequence is generated, specifically including: Read the spatial location of the central sample image block and the spatial location of each neighboring image block; Determine the spatial offset relationship and adjacency direction type of each neighboring image block relative to the central sample image block; Map spatial offset relationships and adjacency direction types to adjacency direction position encoding features; Generate an adjacent direction position coding sequence according to the spatial order of the central neighborhood image block sequence; The Transformer input feature sequence is generated based on the image patch embedding feature sequence and the adjacent orientation position encoding sequence, specifically including: The image block embedding feature corresponding to the same image block is added to the adjacent direction position encoding feature so that the image block embedding feature carries the spatial adjacency information of the image block relative to the central sample image block. The summed features are arranged in spatial order according to the central neighborhood image block sequence to obtain the Transformer input feature sequence for inputting the pre-trained visual Transformer feature extraction model. Multi-head attention processing is performed on the input feature sequence of Transformer to generate center neighborhood attention features, specifically including: In each Transformer encoding layer, the central neighborhood feature sequence output from the previous layer is mapped to query features, key features, and value features, respectively, and the similarity between query features and key features is calculated. Attention allocation results are generated based on similarity, and the value features are weighted and aggregated using the attention allocation results to obtain the central neighborhood attention features that reflect the degree of spatial correlation between the central sample image block and the neighboring image blocks. Based on the attention features of the central neighborhood, spatial continuity encoding features of the central neighborhood are generated, specifically including: The center neighborhood attention features output from multiple attention heads in the same Transformer encoding layer are concatenated; Output mapping and feedforward network mapping are performed on the stitched attention features to transmit the texture continuity information, boundary connection information and spatial continuity difference information between the central sample image block and the neighboring image blocks layer by layer between the coding layers, so as to obtain the spatial continuity coding features of the central neighborhood. The texture continuity relationship, boundary connection relationship and neighborhood spatial continuity difference are calculated based on the coding features of the central neighborhood space. Read the coding features of the central sample image block and the coding features of each neighboring image block from the output of the last Transformer coding layer; The neighborhood spatial continuity difference is generated based on the overall feature difference between the coding features of the central sample image block and the coding features of the neighboring image blocks; Boundary connection relationships are generated based on the similarity of encoded features between the central sample image block and neighboring image blocks near the common boundary. Texture continuity relationships are generated based on the similarity of encoded features between the central sample image block and neighboring image blocks along the adjacent direction; Deep spatial continuity feature vectors are generated based on texture continuity, boundary connection, and neighborhood spatial continuity differences, specifically including: The image block coding features of the central sample, the average convergence results of the coding features of the neighboring image blocks, the texture continuity relationship, the boundary connection relationship, and the differences in the continuity of the neighborhood space are stitched together; Deep spatial continuity feature vectors are generated through feature mapping.
[0029] In this embodiment, the process of training a paper-cutting deformation recognition network using a self-supervised positive and negative sample set and neighborhood continuity constraints to generate an initial paper-cutting deformation recognition model specifically includes: The model training sample sequence carrying neighborhood relationships is generated based on the self-supervised positive and negative sample set and the network input sample set, specifically including: Read the positive sample subset, negative sample subset and corresponding sample labels, and input any central sample image block and neighboring image blocks into the paper pattern deformation recognition network; Obtain the discriminant features of latte art deformation recognition network, the output probability of normal samples, and the output probability of latte art deformation samples; The classification output layer receives the latte art deformation discrimination features output by the feature fusion layer, and generates classification scores for normal samples and latte art deformation samples respectively through fully connected mapping. Softmax normalization is applied to the classification scores of normal samples and latte art deformation samples to obtain the output probabilities of normal samples and latte art deformation samples. Among them, the normal sample output probability is used to represent the confidence level that the central sample image block belongs to the normal sample, and the latte art deformation sample output probability is used to represent the confidence level that the central sample image block belongs to the latte art deformation sample. The sum of the two is one unit probability. Write the central sample image block, neighboring image blocks, sample labels, latte art deformation discrimination features, normal sample output probability, and latte art deformation sample output probability into the model training sample sequence. The classification loss is calculated based on the output probability of normal samples, the output probability of latte art deformed samples, and the sample labels, specifically including: For any central sample image block, when the sample label is normal sample, read the normal sample output probability; When the sample label is a latte art deformation sample, read the output probability of the latte art deformation sample, and take the negative logarithm of the output probability that matches the sample label. The average of the negative logarithmic results corresponding to all central sample image blocks in the current training batch is used to obtain the classification loss used to constrain the classification results of normal samples and latte art deformed samples; The loss of feature aggregation for similar samples is calculated based on the discriminative features of latte art deformation and sample labels, specifically including: Based on the sample labels, the discriminant features of latte art deformation in the current training batch are divided into normal sample feature group and latte art deformation sample feature group; Calculate the feature centers of the normal sample feature group and the feature centers of the latte art deformation sample feature group respectively; Calculate the feature distance between each pattern deformation discrimination feature and its corresponding feature center within the same category; The average of the squared feature distances is used to obtain the feature aggregation loss of similar samples, which is used to constrain the degree of feature clustering of similar samples; The neighborhood continuity constraint loss is calculated based on the original image location continuity field and deep spatial continuity feature vectors, specifically including: For any central sample image block, read the differences in positional continuity intensity, texture continuity, boundary connection, and spatial continuity between it and its neighboring image blocks; Within the same training batch, the differences in positional continuity fracture intensity between all central sample image blocks and neighboring image blocks are statistically analyzed, and their median value and interquartile range are calculated. The differences in all texture continuity relationships, boundary connection relationships, and neighborhood spatial continuity are statistically analyzed, and the median and interquartile range are calculated respectively. When the difference in positional continuity intensity between the central sample image block and the neighboring image block is lower than the median value of the difference in positional continuity intensity, and the texture continuity relationship is not lower than the median value of the texture continuity relationship and the boundary connection relationship is not lower than the median value of the boundary connection relationship, the central sample image block and the neighboring image block are determined as a continuous neighborhood pair. When the difference in positional continuity intensity between the central sample image block and the neighboring image block is not less than the sum of the median value of the positional continuity intensity difference and the interquartile range, and the difference in spatial continuity between the neighboring blocks is not less than the sum of the median value of the spatial continuity difference and the interquartile range, the central sample image block and the neighboring image block are determined as a breakage neighborhood pair. Calculate the discriminant feature distance of the lace-pattern deformation between consecutive neighborhood pairs and average their squared values; Calculate the penalty value generated when the distance between the discriminant feature of the broken neighborhood pair is lower than the interval distance; The average feature distance of continuous neighborhood pairs and the average penalty value of broken neighborhood pairs are added together to obtain the neighborhood continuity constraint loss used to constrain the continuity difference between adjacent image blocks. The joint loss is calculated based on classification loss, feature aggregation loss of similar samples, and neighborhood continuity constraint loss, specifically including: The feature aggregation loss weight and neighborhood continuity constraint loss weight of the same type of samples are determined based on the inter-class separability change results of normal samples and latte art deformation samples in the validation samples. After each round of training, the inter-class separability of the discriminant features of the latte art deformation and the inter-class separability of the neighborhood space continuity difference are calculated using the validation sample set. The inter-class separability of the discriminative features of latte art deformation is determined based on the distance between the feature centers of normal samples and latte art deformation samples, as well as the average distance of each of the two classes of samples to the feature centers of their respective classes. The greater the distance between the feature center of the normal sample and the feature center of the latte art deformation sample, and the smaller the average distance between the two types of samples to their respective category feature centers, the higher the inter-class separability of the latte art deformation discrimination feature; The inter-class separability of neighborhood space continuity differences is determined based on the difference in class mean and class dispersion between normal samples and latte art deformation samples in neighborhood space continuity differences. Divide the inter-class separability of the latte art deformation discrimination feature by the sum of the two inter-class separability values to obtain the feature aggregation loss weight of the same type of samples; Divide the inter-class separability of the neighborhood space continuity difference by the sum of the two inter-class separability to obtain the neighborhood continuity constraint loss weight. The joint loss is obtained by summing the classification loss, the aggregated loss of similar sample features multiplied by the weight of the aggregated loss of similar sample features, and the neighborhood continuity constraint loss multiplied by the weight of the neighborhood continuity constraint loss. The network parameters of the shallow endogenous feature branch, feature fusion layer, and classification output layer are updated based on joint loss, specifically including: The pre-trained visual Transformer feature extraction model is used as the fixed feature extraction backbone of the deep spatial continuity branch; The pre-trained parameters are retained for forward feature computation, and the gradient of the joint loss with respect to the trainable parameters of the self-learning convolutional network, feature fusion layer, and classification output layer in the shallow endogenous feature branch is calculated. Update the network parameters of the shallow endogenous feature branch, feature fusion layer, and classification output layer in the direction that reduces the joint loss; An initial model for recognizing latte art deformations is generated based on the updated network parameters, specifically including: During training, the classification accuracy of validation samples, the recall rate of deformed samples, and the decrease in neighborhood continuity constraint loss are continuously calculated. Training is stopped when the classification accuracy of the validation samples and the recall of the latte art deformation samples remain stable for multiple consecutive iterations, and the decrease in the neighborhood continuity constraint loss is less than the training convergence limit. During the training process, after each iteration, the classification accuracy of the validation samples, the recall rate of the latte art deformation samples, and the neighborhood continuity constraint loss are statistically analyzed. If the classification accuracy of the validation samples does not change by more than 0.5% for five consecutive iterations, the recall rate of the latte art deformation samples does not change by more than 0.5% for five consecutive iterations, and the neighborhood continuity constraint loss decreases by less than 1% for five consecutive iterations, then the training is considered to have reached convergence and training is stopped. The current shallow endogenous feature branch, deep spatial continuity branch, feature fusion layer, and classification output layer are combined into an initial paper-cutting deformation recognition model. The current shallow endogenous feature branch, deep spatial continuity branch, feature fusion layer, and classification output layer are combined into an initial paper-cutting deformation recognition model.
[0030] In this embodiment, embedding the initial latte art deformation recognition model into the satellite image production workflow to generate a self-reinforcing latte art deformation recognition model specifically includes: Based on the initial pattern deformation recognition model and satellite imagery production workflow, a set of production recognition results is generated, specifically including: The DOM images in the production process are divided into production image blocks using the same position image block division method that is consistent with the spatial position of the DOM image to be detected; Each production image block is input into the initial latte art deformation recognition model to obtain the normal sample output probability, latte art deformation sample output probability, latte art deformation discrimination features, and latte art deformation recognition results for each production image block. Write the above results into the production recognition result set according to the spatial location of the production image blocks; A set of verification difference samples is generated based on the production identification result set and the manual review result set, specifically including: Read the manual review labels and boundaries from the manual review results; The results of the pattern deformation recognition of each production image block are compared with the manually verified labels; Production image blocks whose latte art deformation identification results are latte art deformation samples and whose manual review labels are normal samples are classified as false detection image blocks; Production image blocks whose pull-out deformation identification results are normal samples and whose manual verification labels are pull-out deformation samples are classified into the missed detection image blocks. Production image blocks whose latte art deformation recognition results are consistent with the manually verified labels and whose recognition boundaries are spatially offset from the manually verified boundaries are classified into boundary correction image blocks. Production image blocks whose latte art deformation recognition results are normal samples and whose manual verification labels are normal samples are classified into normal image blocks. False detection image blocks, missed detection image blocks, boundary-corrected image blocks, and normal image blocks are written into the review difference sample set; A new training sample set is generated based on the verified difference sample set, specifically including: For any sample with discrepancies, read the grayscale data of the DOM image, the spatial location of the image block, the endogenous anomaly features of the image block, the texture continuity relationship between the central sample image block and the neighboring image blocks, the boundary connection relationship, and the differences in the spatial continuity of the neighborhood. Use the manually verified labels as the updated sample labels for the verified difference samples; The verified difference sample types, updated sample labels, and corresponding feature data are then written into the newly added training sample set. Among them, the types of images with differences to be reviewed include falsely detected image blocks, missed image blocks, boundary-corrected image blocks, and normal image blocks; The confidence bias is generated based on the difference between the recognition probability of each newly added training sample in the newly added training sample set and the manually verified label, specifically including: For any new training sample, when the label is manually verified as a latte art deformation sample, calculate the absolute difference between the output probability of the latte art deformation sample output by the initial latte art deformation recognition model and the label value of the latte art deformation sample. When the labels are manually verified as normal samples, calculate the absolute difference between the output probability of the deformed sample of the initial latte art deformation recognition model and the label value of the normal sample. Set the label value of manually reviewed samples labeled as deformed to 1, and set the label value of manually reviewed samples labeled as normal to 0. For any new training sample, read the output probability of the latte art deformation recognition model and calculate the absolute difference between the output probability and the manually verified label value to obtain the confidence bias. If the output probability of the new training sample with the deformed latte art pattern is 0.82, and the manually verified label is a deformed latte art pattern sample, then the confidence bias is 0.18. If the output probability of the new training sample with the deformed latte art pattern is 0.76, and the manually verified label is a normal sample, then the confidence bias is 0.76. The absolute difference is used as the confidence bias of the newly added training sample, and the confidence bias is written into the set of newly added training samples along with the spatial location of the newly added training sample and the type of the verification difference sample. Sample training weights are generated based on confidence bias and verification difference sample types, specifically including: For any new training sample, with the basic training weights as the initial value, the weight contribution corresponding to the confidence bias, the weight contribution corresponding to the false detection image block, the weight contribution corresponding to the false detection image block, and the weight contribution corresponding to the boundary correction image block are superimposed on the basic training weights to obtain the unnormalized training weights. The adjustment range of each type of weight contribution is determined based on the proportion of the number of falsely detected image blocks, missed image blocks, boundary-corrected image blocks, and normal image blocks in the newly added training sample set. Calculate the average of all unnormalized training weights in the newly added training sample set, and divide the unnormalized training weight of each newly added training sample by the average to obtain the sample training weight of the newly added training sample. Set the basic training weight to 1. For any new training sample, multiply the confidence bias by the confidence bias adjustment coefficient to obtain the weight contribution corresponding to the confidence bias. If the newly added training sample is a false detection image block, then the weight contribution of the false detection image block is added. If the newly added training sample is a missed image block, then the weight contribution of the missed image block is added. If the newly added training sample is a boundary-corrected image patch, then the weight contribution of the boundary-corrected image patch is superimposed. If the newly added training sample is a normal image patch, the difference type weight contribution will not be superimposed. The weight contribution of false detection image blocks, false detection image blocks, and boundary correction image blocks is determined based on the proportion of the corresponding type of samples in the newly added training sample set. Types with a lower proportion of samples and a stronger effect on model correction correspond to higher weight contributions. After obtaining the unnormalized training weights of each new training sample, calculate the average of all unnormalized training weights, and divide each unnormalized training weight by the average to obtain the sample training weight, so that the average sample training weight of the new training sample set remains at 1. The confidence-weighted fine-tuning loss is calculated based on the sample training weights and the newly added training sample set, specifically including: Each newly added training sample is input into the initial latte art deformation recognition model to obtain the latte art deformation discrimination features, the output probability of normal samples, and the output probability of latte art deformation samples. The classification loss of each newly added training sample is calculated based on the difference between the updated sample label and the classification output probability. The classification loss is then weighted and averaged according to the sample training weights to obtain the weighted classification loss. The feature aggregation loss of similar samples is calculated based on the distance between the discriminant feature of the same updated sample label and the feature center of the corresponding category, and then weighted and averaged according to the sample training weights to obtain the weighted feature aggregation loss of similar samples. The neighborhood continuity constraint loss is calculated based on the differences in positional continuity fracture intensity, texture continuity, boundary connection, and neighborhood spatial continuity between the newly added training samples and their neighboring image blocks. The weighted average is then calculated according to the training weights of the samples to obtain the weighted neighborhood continuity constraint loss. The confidence-weighted fine-tuning loss is obtained by summing the weighted classification loss, the weighted feature aggregation loss of similar samples, and the weighted neighborhood continuity constraint loss according to their respective loss weights. The initial latte art deformation recognition model is updated based on confidence-weighted fine-tuning loss, and a self-reinforcing latte art deformation recognition model is generated, specifically including: The backbone parameters of the pre-trained visual Transformer feature extraction model in the deep spatial continuity branch are retained to participate in the forward feature calculation; Calculate the gradient of the confidence-weighted fine-tuning loss with respect to the trainable parameters of the shallow endogenous feature branch, the feature fusion layer, and the classification output layer; Update the network parameters of the shallow endogenous feature branch, feature fusion layer, and classification output layer in the direction that reduces the confidence-weighted fine-tuning loss; The finely tuned shallow endogenous feature branch, deep spatial continuity branch, feature fusion layer, and classification output layer are combined into a self-reinforcing paper-cutting deformation recognition model.
[0031] In this embodiment, the generation of a comprehensive pattern discrimination map for the DOM image to be identified, using a self-reinforcing pattern deformation recognition model and an endogenous anomaly feature set of image blocks, specifically includes: Based on the DOM image to be identified, a field of original image location continuity and a set of image blocks at the same location to be identified are generated, specifically including: Read the geographic coordinate information, interior orientation elements, and orthorectification model parameters carried by the DOM image to be identified; The pixel coordinates in the DOM image to be identified are inversely calculated from the original image position to obtain the original image position coordinate field of the DOM image to be identified. Calculate the directional change, distance change, and continuity break of the original image position coordinates of adjacent pixels to generate the original image position continuity field of the DOM image to be identified. The image block side length is determined according to the ground resolution of the DOM image to be identified. The DOM image to be identified, the orientation change channel, the distance change channel and the continuous break channel are divided into the same position to obtain the set of image blocks to be identified at the same position. Based on the original image location continuity field of the DOM image to be identified and the set of image blocks at the same location to be identified, an endogenous anomaly feature set of the image blocks to be identified is generated, specifically including: For any image block to be identified, calculate the positional continuity fracture intensity, pixel stretching ratio, texture anisotropy, pixel gradient abrupt change value, gray-level co-occurrence matrix contrast, gray-level co-occurrence matrix correlation, and gray-level co-occurrence matrix energy of the image block to be identified. The above features are combined with the spatial location of the image block to be identified to form an endogenous anomalous feature record of the image block to be identified. All endogenous anomaly features of the image blocks to be identified are recorded and written into the endogenous anomaly feature set of the image blocks to be identified according to the spatial location of the image blocks to be identified. Based on the set of image blocks at the same location to be identified, the set of endogenous anomaly features of the image blocks to be identified, and the self-reinforcing pattern deformation recognition model, the probability of pattern deformation and the difference in neighborhood spatial continuity of the image blocks to be identified are generated, specifically including: For any image block to be identified, read the image block to be identified and its neighboring image blocks to construct the corresponding network input sample group; Input the network input sample group into the self-reinforcing paper pattern deformation recognition model to obtain the paper pattern deformation discrimination features, normal sample output probability and paper pattern deformation sample output probability of the image block to be recognized; The probability of the pull-out deformation sample output is used as the probability of the pull-out deformation of the image block to be identified, and the difference in neighborhood spatial continuity output of the deep spatial continuity branch is read. The texture information entropy of each image patch to be identified is generated based on the grayscale data of the DOM images in the set of image patches to be identified at the same location. Specifically, this includes: For any image block to be identified, the proportion of each gray level pixel within the image block is counted. Calculate the product of the percentage of pixels at each gray level and its logarithmic result; The texture information entropy of the image block to be identified is obtained by summing the product results of all gray levels and taking the negative number. Based on the separability between classes in the validation sample set, fusion weights are determined for the probability of wreath deformation, pixel stretching ratio, positional continuity break strength, texture information entropy, and neighborhood spatial continuity differences. Specifically, these weights include: The class mean and class dispersion of normal samples and deformed samples in the validation sample set were statistically analyzed in terms of deformed probability, pixel stretching ratio, positional continuity break strength, texture information entropy and neighborhood spatial continuity difference. For any feature, calculate the difference between the mean of the latte art deformation sample category and the mean of the normal sample category, and divide the square of the difference by the sum of the dispersion of the latte art deformation sample category and the dispersion of the normal sample category to obtain the inter-class separability of the feature; Divide the inter-class separability of each feature by the sum of the inter-class separability of all features to obtain the fusion weight corresponding to the feature; The comprehensive stylus discriminant value for each image block to be identified is generated based on the fusion weights, specifically including: The probability of the graffiti deformation, the pixel stretching ratio, the positional continuity break intensity, the texture information entropy, and the difference in the continuity of the neighborhood space of each image block to be identified are normalized within the same image to obtain the corresponding normalization results. The normalized result of the deformation probability of the latte art is multiplied by the deformation probability fusion weight, the normalized result of the pixel stretching ratio is multiplied by the pixel stretching ratio fusion weight, the normalized result of the positional continuity break intensity is multiplied by the positional continuity break intensity fusion weight, the normalized result of the texture information entropy is multiplied by the texture information entropy fusion weight, and the normalized result of the neighborhood spatial continuity difference is multiplied by the neighborhood spatial continuity difference fusion weight. The five weighted results are then summed to obtain the comprehensive deformation discriminant value of the image block to be identified. A comprehensive pattern discrimination map of the DOM image to be identified is generated based on the comprehensive pattern discrimination value and the set of image blocks at the same location to be identified. Specifically, it includes: The comprehensive stylus discriminant value of each image block to be identified is written into the corresponding raster area according to the spatial position of the image block to be identified, so as to obtain the image block-level discriminant result; Neighborhood smoothing is performed on the comprehensive stagger discriminant values between spatially adjacent image blocks to be identified to generate a comprehensive stagger discriminant map that is consistent with the spatial range of the DOM image to be identified.
[0032] In this embodiment, based on the comprehensive latte art discrimination map and the positional continuity field of the original image, the output of the latte art deformation area results specifically includes: Based on the comprehensive latte art discriminant map, a comprehensive latte art discriminant value distribution sequence is generated, and the latte art judgment boundary is determined. Specifically, this includes: Read the comprehensive stylus discriminant value of all image blocks to be identified within the same DOM image to be identified; The overall latte art discriminant values are sorted from smallest to largest to form an overall latte art discriminant value distribution sequence; Each candidate segmentation value in the comprehensive latte art discriminant value distribution sequence is selected as the candidate latte art judgment boundary; The sample proportions and average values of the low discriminant value group (below the candidate latte art judgment threshold) and the high discriminant value group (not below the candidate latte art judgment threshold) were calculated separately. Multiply the sample proportions of the low discriminant value group, the sample proportions of the high discriminant value group, and the squared difference between the means of the two groups to obtain the inter-class variance corresponding to the candidate latte art decision boundary. The candidate latte art decision boundary that maximizes the inter-class variance is determined as the latte art decision boundary. The boundary for determining the intensity of fracture continuity is generated based on the location continuity field of the original image, specifically including: Read the positional continuity fracture intensity of all image blocks to be identified within the same DOM image to be identified; Arrange the continuous fracture strength at each location in ascending order; Determine the median, first quartile, and third quartile of the continuous fracture strength at the location, and define the difference between the third quartile and the first quartile as the interquartile range. The median of the location-continuous fracture strength is added to the interquartile range to obtain the determination threshold of the location-continuous fracture strength. Based on the set of image blocks to be identified at the same location and the positional continuity field of the original image, the main stretching direction of each image block to be identified is generated, specifically including: For any image block to be identified, read the inverse displacement relationship of adjacent pixels within the coverage area of the image block to be identified; The cumulative results of the horizontal and vertical inverse displacement components are calculated separately. The main stretching direction of the image block to be identified is determined based on the directional relationship between the cumulative displacement results in the vertical direction and the cumulative displacement results in the horizontal direction. Write the spatial correspondence between the main stretching direction and the image block to be identified into the main stretching direction set; Based on the criteria for determining the quality of the pull-out pattern, the criteria for determining the continuous fracture strength at the location, and the main tensile direction, a set of pull-out deformation image blocks is generated, specifically including: For any image block to be identified, read the comprehensive stylization discrimination value, positional continuous fracture strength, and main tensile direction; When the comprehensive discriminant value of the pull-out pattern is not lower than the pull-out pattern judgment limit and the positional continuous fracture strength is not lower than the positional continuous fracture strength judgment limit, the image block to be identified is determined as the pull-out pattern deformed image block. Write all the latte art deformation image blocks, their corresponding main stretching directions, and the spatial positions of the image blocks into the latte art deformation image block set. Based on the set of latte art deformable image blocks and the main stretching direction, a set of connected latte art deformable regions is generated, specifically including: For spatially adjacent distorted image blocks, perform an adjacency search. For any two spatially adjacent distorted image blocks, calculate the directional difference between their main stretching directions. When the directional difference is not greater than the directional consistency limit, the corresponding two graffiti image blocks are classified into the same connected graffiti region. Perform connectivity merging on all the latte art deformation image blocks that satisfy spatial connectivity and have the same main stretching direction to generate a set of connected latte art deformation regions. Based on the set of connected latte art deformation regions, a set of boundary-optimized latte art deformation regions is generated, specifically including: Count the number of image blocks and the area of each connected lattice deformation region; Connected ripple deformation regions with fewer image blocks than the isolated block number limit or with a region area smaller than the isolated block area limit are identified as isolated blocks and removed. Perform boundary smoothing on the preserved connected lattice deformation region for any boundary point on the region boundary; Read the x-coordinate and y-coordinate of the previous boundary point, the current boundary point, and the next boundary point respectively; The smoothed x-coordinates are obtained by averaging the three x-coordinates and the smoothed y-coordinates are obtained by averaging the three y-coordinates. The set of boundary-optimized latte art deformation regions is then generated according to the smoothed boundary point coordinates. Based on the boundary-optimized set of deformed regions, a binary raster mask and vector boundary files are generated, specifically including: According to the raster row and column position of the DOM image to be identified, the pixels belonging to the set of boundary-optimized latte art deformation regions are assigned a value of one; Pixels that do not belong to the set of deformed regions after boundary optimization are assigned a value of zero, and a binary raster mask of the deformed region is generated. Extract the sequence of closed boundary points for each optimized latte art deformation region, generate polygonal boundaries according to the spatial connection order of the boundary points, and write the vertex coordinates, region number, and spatial reference information of each polygonal boundary into a vector data structure to output the vector boundary file of the latte art deformation region.
[0033] refer to Figure 3 and Figure 4Example 1: To verify the feasibility of this invention in practice, it was applied to the high-resolution satellite DOM image quality inspection process of a surveying and mapping production unit. This process is mainly used to identify distorted areas in orthorectified DOM results. The image to be processed covers urban building areas, farmland areas, road intersections, bare areas, and hilly edges. The ground resolution of the image is 0.5m, and the size of a single DOM image is 16000 pixels × 16000 pixels. Because hilly edges and areas with elevation changes are prone to local pixel stretching, abrupt changes in texture direction, and breaks in adjacent positional relationships during orthorectification, traditional manual visual inspection requires repeated zooming in, zooming out, and switching image layers. Fixed threshold methods are prone to misjudging road edges, building shadows, and abrupt changes in farmland texture as distorted areas. Therefore, this scenario effectively demonstrates the problems that this invention aims to solve regarding the accuracy of distorted area identification, boundary positioning, and production review efficiency.
[0034] In this scenario, the system first reads the geographic coordinates, interior orientation elements, and orthorectification model parameters carried by the DOM image to be identified. It then performs inverse calculation of the original image position for the pixel coordinates in the DOM image, forming an original image position continuity field. During implementation, the image is divided into 512×512 pixel image blocks at the same location. For each image block, the system extracts the positional continuity fracture intensity, pixel stretching ratio, texture anisotropy, pixel gradient abrupt change value, gray-level co-occurrence matrix contrast, gray-level co-occurrence matrix correlation, and gray-level co-occurrence matrix energy. Taking one DOM image as an example, the system generates 976 valid image blocks, of which 58 were manually verified to have distorted patterns, and 918 were normal. Based on the endogenous anomaly feature set of the image blocks, the system generates a self-supervised positive and negative sample set that requires no manual annotation. Initially, 64 image blocks were labeled as distorted patterns, and 812 were initially labeled as normal. After active masking, rotation, translation, and gray-level perturbation enhancement, 3480 training samples are formed.
[0035] In the model application phase, the system inputs each central sample image block and its neighboring image blocks into the self-reinforcing pattern deformation recognition model to obtain the pattern deformation probability and the difference in neighborhood spatial continuity. It then combines pixel stretching ratio, positional continuous fracture strength, and texture information entropy to generate a comprehensive pattern discrimination map. For a pattern deformation area confirmed by manual verification, the manual verification record shows a continuous deformation zone length of approximately 368.5m and an average width of approximately 18.7m, mainly manifested as road edge misalignment, slope texture stretching, and boundary fractures of adjacent features. When the traditional fixed threshold method identifies this area, it simultaneously includes nearby building shadows and farmland stripes in the abnormal area, resulting in an output area of 10320m²; the manually verified area is 6910m², with an area deviation of 3410m². The present invention outputs an area of 7245m², with an area deviation of 335m² and an average boundary offset distance of 1.8m. For another small-scale pattern deformation area, the manually verified area was 1260m². Traditional methods failed to detect it completely due to the lack of obvious texture contrast, resulting in a missed detection area of 520m². This invention, through joint discrimination based on the positional continuity of fracture strength and the difference in the continuity of the neighborhood space, resulted in a missed detection area of 96m².
[0036] To verify that the results were not accidental from a single scene, the implementer applied the invention to 12 DOM images from the same production batch. Each image underwent manual review to create a control result. The total processed area of the 12 DOM images was 768 km², with a total of 11,712 image blocks. 684 image blocks with pattern distortion were confirmed by manual review. The traditional fixed threshold method detected 612 image blocks with pattern distortion, including 138 false positives and 210 false negatives. The invention detected 701 image blocks with pattern distortion, including 46 false positives and 29 false negatives. The average boundary offset distance of the pattern distortion area using the traditional fixed threshold method was 5.6 m, while the average boundary offset distance using the invention was 1.9 m. With traditional manual review combined with fixed threshold screening, the average processing time for a single DOM image was 46.8 min, with manual review taking 31.5 min. Using the invention, the average processing time for a single DOM image was 18.6 min, with manual review taking 9.4 min. After production review and re-feedback, the system writes the false detection image blocks, missed detection image blocks, boundary correction image blocks, and normal image blocks into the new training sample set. It also generates sample training weights based on the difference between the recognition probability and the manual review label. After three rounds of confidence weighted fine-tuning, the recall rate of the latte art deformation sample increased from 90.8% to 95.8%, and the false detection rate decreased from 7.3% to 4.1%.
[0037] Table 1 Comparison of the Implementation Effects of Remote Sensing Image Deformation Recognition
[0038] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for identifying distortion regions in remote sensing images based on self-supervised learning, characterized in that, The steps include the following: Acquire the DOM image to be detected and generate the original image position continuity field based on the geometric parameters of the DOM image itself; An endogenous anomaly feature set for image patches is generated based on the location continuity field of the original image and the texture information of the DOM image, specifically including: The image block side length is determined according to the resolution of the DOM image to be detected, and the positional continuity field of the DOM image to be detected and the original image is divided by the boundary of the same image block to obtain a set of image blocks at the same position. Read the continuous fracture data within the coverage area of each image block in the image block set at the same location, and statistically obtain the set of continuous fracture intensities at the location. Read the direction change data and distance change data within the coverage area of each image block in the image block set at the same location, and calculate the pixel stretching ratio set by combining the location continuous fracture intensity set. Read the grayscale data of the DOM images of each image block in the image block set at the same location, count the grayscale change relationship in different directions, and combine it with the set of pixel stretching ratios to obtain the texture mutation feature set. Perform gray-level co-occurrence matrix statistics on DOM image gray-level data and extract a set of abnormal features from the gray-level co-occurrence matrix; The set of location continuous fracture intensity, set of pixel stretching ratio, set of texture abrupt change features, and set of gray-level co-occurrence matrix anomaly features are combined according to the spatial location of the image block to obtain the set of endogenous anomaly features of the image block; Generate a self-supervised positive and negative sample set based on the endogenous anomaly feature set of image blocks; A network for recognizing paper-cut patterns is constructed based on a self-supervised positive and negative sample set, which fuses shallow endogenous feature branches and deep spatial continuity branches. Specifically, it includes: Extract the central image block, neighboring image blocks, endogenous anomaly features of image blocks, and sample labels from the self-supervised positive and negative sample set and the image block set at the same location to form the network input sample group; Input the network input sample group into the shallow endogenous feature branch, extract the local geometric anomaly features and texture anomaly features of the central image block, and obtain the shallow endogenous feature vector; Input the network input sample group into the deep spatial continuity branch, extract the spatial continuity difference between the central image block and the neighboring image blocks, and obtain the deep spatial continuity feature vector. The shallow endogenous feature vector and the deep spatial continuity feature vector are input into the feature fusion layer to obtain the discriminative features for latte art deformation. Input the discriminant features of the latte art deformation into the classification output layer to obtain the latte art deformation recognition output; By connecting the shallow endogenous feature branch, the deep spatial continuity branch, the feature fusion layer, and the classification output layer, a paper-cutting deformation recognition network is obtained. A network for recognizing paper-cutting deformations was trained using a self-supervised positive and negative sample set and neighborhood continuity constraints to generate an initial paper-cutting deformation recognition model. The initial pattern deformation recognition model is embedded into the satellite image production workflow to generate a self-reinforcing pattern deformation recognition model. By utilizing a self-reinforcing latte art deformation recognition model and an endogenous anomaly feature set of image blocks, a comprehensive latte art discrimination map of the DOM image to be identified is generated; Based on the integrated latte art discrimination map and the positional continuity field of the original image, the latte art deformation area results are output.
2. The remote sensing image distortion region identification method based on self-supervised learning according to claim 1, characterized in that, The process of acquiring the DOM image to be detected and generating the original image position continuity field based on the DOM image's own geometric parameters specifically includes: Obtain the DOM image to be detected and generate a set of geometric parameters of the DOM image itself from the image metadata; Based on the geometric parameter set of the DOM image itself, the pixel row and column positions of each DOM pixel are converted into corresponding ground coordinates to obtain the ground coordinate field; Input the ground coordinates in the ground coordinate field into the inverse calculation relationship corresponding to the orthorectification model parameters, and combine the interior orientation elements to inversely calculate the original image position corresponding to each DOM pixel point, and obtain the original image position coordinate field. Read the original image position coordinates corresponding to adjacent DOM pixels, calculate the displacement direction and displacement distance between adjacent original image positions, and obtain the set of inverse displacement relationships; By comparing the directional and distance differences between adjacent back-calculated displacement relationships, the directional change and distance change are obtained respectively; The changes in direction and distance are normalized, and the normalized changes in direction and distance are fused into a continuous break, thus constructing the positional continuity field of the original image.
3. The remote sensing image distortion region identification method based on self-supervised learning according to claim 1, characterized in that, The generation of a self-supervised positive and negative sample set based on the endogenous anomaly feature set of image blocks specifically includes: Geometric continuity anomaly information and texture mutation information corresponding to each image block are extracted from the endogenous anomaly feature set of the image blocks to form a candidate sample set; Normalize and statistically analyze the geometric continuity anomaly information and texture mutation information in the candidate sample set to obtain a set of joint distribution characterization values; The sample separation boundary is determined based on the distribution differences of each candidate sample in the joint distribution characterization value set; By using the sample separation boundary to screen the endogenous anomaly feature set of image blocks, an initial self-supervised positive and negative sample set is obtained; The initial self-supervised positive and negative sample set is subjected to rotation, translation, and grayscale perturbation to obtain the enhanced sample set; By integrating the initial self-supervised positive and negative sample set and the augmented sample set, a self-supervised positive and negative sample set is obtained.
4. The remote sensing image distortion region identification method based on self-supervised learning according to claim 1, characterized in that, The step of inputting the network input sample group into the deep spatial continuity branch, extracting the spatial continuity difference between the central image block and neighboring image blocks, and obtaining the deep spatial continuity feature vector specifically includes: Read the central image block and neighboring image blocks from the network input sample group, and arrange them according to their spatial location in the neighborhood to obtain the central neighboring image block sequence; Embedding mapping is performed on each image block in the central neighborhood image block sequence to obtain the image block embedding feature sequence; According to the spatial position of each image block relative to the central image block, the adjacent direction position coding sequence is obtained; The image patch embedding feature sequence and the adjacent direction position encoding sequence are fused together to obtain the Transformer input feature sequence; Input the Transformer input feature sequence into the multi-head attention layer to obtain the center neighborhood attention features; Calculate texture continuity, boundary connection, and neighborhood spatial continuity differences based on the attention features of the central neighborhood; The texture continuity relationship, boundary connection relationship and neighborhood spatial continuity difference are combined into a deep spatial continuity feature vector.
5. The remote sensing image distortion region identification method based on self-supervised learning according to claim 1, characterized in that, The process of training the paper-cutting deformation recognition network using a self-supervised positive and negative sample set and neighborhood continuity constraints to generate an initial paper-cutting deformation recognition model specifically includes: The model training sample sequence is obtained by organizing the supervised positive and negative sample set and the network input sample set according to the sample labels. Input the training sample sequence of the model into the latte art deformation recognition network, output the latte art deformation recognition result, and calculate the classification loss according to the difference between the latte art deformation recognition result and the sample label; The distance between the discriminant features of the same type of samples corresponding to the latte art deformation is statistically analyzed to obtain the feature aggregation loss of the same type of samples; The continuous breakage amount in the location continuity field of the original image is read and correspondingly constrained with the deep spatial continuity feature vector to obtain the neighborhood continuity constraint loss. The classification loss, the feature aggregation loss of similar samples, and the neighborhood continuity constraint loss are weighted and summed to obtain the joint loss; The network parameters of the shallow endogenous feature branch, feature fusion layer, and classification output layer are updated according to the joint loss, and the initial paper-cutting deformation recognition model is obtained using the updated network parameters.
6. The remote sensing image distortion region identification method based on self-supervised learning according to claim 1, characterized in that, The step of embedding the initial latte art deformation recognition model into the satellite image production workflow to generate a self-reinforcing latte art deformation recognition model specifically includes: The initial pattern deformation recognition model is embedded into the satellite image production workflow, and pattern deformation recognition is performed on DOM images during the production process to obtain a set of production recognition results. The set of production identification results is compared with the results of manual review to obtain a set of review difference samples; Extract the DOM image grayscale data and endogenous anomaly features of the corresponding image blocks from the set of review difference samples, and assign sample labels according to the manual review results to obtain a new training sample set; The confidence bias is obtained by calculating the difference between the recognition probability of each newly added training sample in the newly added training sample set and the manually verified label. The training weights of the samples are obtained by combining the confidence bias and the sample types of verification differences; The sample training weights are added to the fine-tuning training process of the newly added training sample set, and the confidence-weighted fine-tuning loss is calculated. The initial latte art deformation recognition model is updated by adjusting the loss using confidence weighting, resulting in a self-reinforcing latte art deformation recognition model.
7. The remote sensing image distortion region identification method based on self-supervised learning according to claim 1, characterized in that, The process of generating a comprehensive pattern discrimination map for the DOM image to be identified by utilizing a self-reinforcing pattern deformation recognition model and an endogenous anomaly feature set of image blocks specifically includes: Perform original image position inverse calculation and same-position block processing on the DOM image to be identified to obtain the original image position continuity field and the set of same-position image blocks to be identified. Based on the original image location continuity field of the DOM image to be identified and the set of image blocks at the same location to be identified, the endogenous anomaly feature set of the image blocks to be identified is obtained. The set of image blocks at the same location to be identified, the set of endogenous abnormal features of the image blocks to be identified, and the self-reinforcing lace deformation identification model are input into the identification calculation process to obtain the lace deformation probability and the difference in the continuity of the neighborhood space. The grayscale data of the DOM images in the set of image blocks at the same location to be identified are statistically analyzed to obtain the texture information entropy; Based on the validation sample set, determine the fusion weights corresponding to the probability of latte art deformation, pixel stretching ratio, positional continuity break strength, texture information entropy, and neighborhood spatial continuity difference, and calculate the comprehensive latte art discriminant value. The comprehensive stylus discriminant value is written into the layer according to the spatial position of the set of image blocks at the same location to be identified, so as to obtain the comprehensive stylus discriminant image of the DOM image to be identified.
8. The remote sensing image distortion region identification method based on self-supervised learning according to claim 1, characterized in that, The process of outputting the deformation region result based on the integrated pattern discrimination map and the positional continuity field of the original image includes: filtering the deformation image blocks according to the integrated pattern discrimination value in the integrated pattern discrimination map and the positional continuity fracture intensity in the positional continuity field of the original image; performing connectivity merging and boundary optimization on the deformation image blocks; generating a binary raster mask and vector boundary file to obtain the deformation region result.
Citation Information
Patent Citations
Rapid satellite image stretching deformation detection method based on GPU-CPU (graphics processing unit-central processing unit) collaboration
CN108230326A
Rapid detection method for garland areas of panoramic images of aerial orthoimages
CN108257130A