A Method for Extracting Cultivated Land Plots from Remote Sensing Images and a Detection Model
By combining the two-path module of hollow convolution and the detection model of attention fusion mechanism, the problem of inaccurate extraction of cultivated land plots in complex agricultural environments is solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202411283753.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-13
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-09-13
AI Technical Summary
Traditional edge detection models based on machine learning algorithms have subjectivity and limitations in complex agricultural environments, resulting in inaccurate extraction results of fuzzy areas of arable land and other land objects, and contain too many irrelevant details.
The detection model of a dual-path module combining hollow convolution and an attention fusion mechanism is adopted. Through the encoder and decoder structure, the dual-path module is used to compress and integrate the feature map, and the attention fusion mechanism is combined to improve the accuracy and robustness of edge detection of cultivated land plots.
It improves the accuracy and robustness of edge detection of cultivated land plots, ensures that the extracted edge information is more accurate, reduces irrelevant details, and adapts to complex agricultural environments.
Smart Images

Figure CN119229287B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and particularly to a method for extracting cultivated land plots from remote sensing images and a detection model. Background Art
[0002] With the advancement of agricultural modernization, the social demand for land resources is continuously increasing, and the accuracy requirements for farmland management are getting higher and higher. Traditional edge detection models based on machine learning algorithms have certain subjectivity and limitations in feature selection and classifier design. For complex and variable agricultural environments, especially in areas where the boundaries between cultivated land and other land features are blurred, the extraction results are prone to errors. Therefore, although the prior art can extract edge information, the extracted edge information often contains too many unimportant details, affecting the accuracy of extraction. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a method for extracting cultivated land plots from remote sensing images and a detection model, which can improve the accuracy and robustness of edge detection of cultivated land plots.
[0004] The technical solution of the present invention for solving the above technical problem is as follows:
[0005] On the one hand, the present invention provides a method for extracting cultivated land plots from remote sensing images, including determining a remote sensing image to be detected. The remote sensing image to be detected includes at least one cultivated land plot. Inputting the remote sensing image to be detected into a detection model to obtain edge information of the at least one cultivated land plot.
[0006] Among them, the detection model includes an encoder, a middle part, and a decoder. The middle part includes a dual-path module. The encoder is used to output at least one feature map corresponding to the remote sensing image to be detected based on the remote sensing image to be detected. The dual-path module is used to output a spliced feature map corresponding to the feature map output by the encoder; the spliced feature map includes multi-scale context information of the corresponding feature map; the dual-path module includes N layer module structures; the first layer module structure includes a first convolutional unit, a first splicing unit, and a dilated convolutional unit connected in sequence; the jth layer module structure includes a first convolutional unit, a first splicing unit, a dilated convolutional unit, a second splicing unit, and a second convolutional unit connected in sequence; 1 < j < N; the Nth layer module structure includes a first convolutional unit and a first splicing unit connected in sequence. The dilation rates of the dilated convolutional units in different layers are different. The first splicing unit included in the jth layer module structure is connected to the dilated convolutional unit included in the (j - 1)th layer module structure; the first splicing unit included in the Nth layer module structure is respectively connected to the dilated convolutional unit included in the (N - 1)th layer module structure and the second convolutional unit included in the (N - 1)th layer module structure; the second splicing unit included in the second layer module structure is connected to the dilated convolutional unit included in the first layer module structure; the second splicing unit included in the fth layer module structure is connected to the second convolutional unit included in the (f - 1)th layer module structure; 2 < f < N. The first convolutional unit is used to compress the input feature map to reduce the number of channels of the input feature map; the second convolutional unit is used to fuse the input feature map to extract features while reducing the number of channels of the input feature map. The dilated convolutional unit is used to obtain multi-scale context information corresponding to the input spliced feature map; both the first splicing unit and the second splicing unit are used to splice the input content and output the corresponding spliced feature map. The decoder is used to output the edge information of at least one cultivated land plot based on the spliced feature maps corresponding to the respective feature maps in the at least one feature map output by the dual-path module.
[0007] The beneficial effects of the present invention are: by combining a dual-path module based on dilated convolution and an attention fusion mechanism network, the accuracy and robustness of cultivated land plot edge detection are improved.
[0008] On the basis of the above technical solutions, the present invention can also be improved as follows.
[0009] Further, the value by which the number of channels of the feature map passing through the first convolutional unit is reduced is a first value. The value by which the number of channels of the feature map passing through the second convolutional unit is reduced is a second value. The second value is greater than the first value.
[0010] Further, for any one of the at least one feature map, the first splicing unit included in the first-layer module structure is configured to splice the any one of the feature maps and the output of the first convolutional unit included in the first-layer module structure, and output a first spliced map corresponding to the any one of the feature maps in the first-layer module structure; the number of channels of the first spliced map corresponding to the any one of the feature maps in the first-layer module structure is equal to the sum of the number of channels of the any one of the feature maps and the number of channels of the feature map output by the first convolutional unit included in the first-layer module structure. The first splicing unit included in the j-th layer module structure is configured to splice the output of the first convolutional unit included in the j-th layer module structure and the output of the dilated convolutional unit included in the (j - 1)-th layer module structure, and output a first spliced map corresponding to the any one of the feature maps in the j-th layer module structure; the number of channels of the first spliced map corresponding to the any one of the feature maps in the j-th layer module structure is equal to the sum of the number of channels of the feature map output by the first convolutional unit included in the j-th layer module structure and the number of channels of the multi-scale context information output by the dilated convolutional unit included in the (j - 1)-th layer module structure. The first splicing unit included in the N-th layer module structure is configured to splice the output of the first convolutional unit included in the N-th layer module structure, the output of the dilated convolutional unit included in the (N - 1)-th layer module structure, and the output of the second convolutional unit included in the (N - 1)-th layer module structure, and output a spliced feature map corresponding to the any one of the feature maps; the number of channels of the spliced feature map corresponding to the any one of the feature maps is equal to the sum of the number of channels of the feature map output by the first convolutional unit included in the N-th layer module structure, the number of channels of the multi-scale context information output by the dilated convolutional unit included in the (N - 1)-th layer module structure, and the number of channels of the feature map output by the second convolutional unit included in the (N - 1)-th layer module structure.
[0011] Further, the second splicing unit included in the second-layer module structure is configured to splice the output of the dilated convolution unit included in the first-layer module structure and the output of the dilated convolution unit included in the second-layer module structure, and output a second spliced map corresponding to any one of the feature maps in the second-layer module structure; the number of channels of the second spliced map corresponding to any one of the feature maps in the second-layer module structure is equal to the sum of the number of channels of the multi-scale context information output by the dilated convolution unit included in the first-layer module structure and the number of channels of the multi-scale context information output by the dilated convolution unit included in the second-layer module structure. The second splicing unit included in the f-th layer module structure is configured to splice the output of the dilated convolution unit included in the f-th layer module structure and the output of the second convolution unit included in the (f-1)-th layer module structure, and output a second spliced map corresponding to any one of the feature maps in the f-th layer module structure; the number of channels of the second spliced map corresponding to any one of the feature maps in the f-th layer module structure is equal to the sum of the number of channels of the multi-scale context information output by the dilated convolution unit included in the f-th layer module structure and the number of channels of the feature map output by the second convolution unit included in the (f-1)-th layer module structure.
[0012] Further, perform segmentation processing on the remote sensing image to be detected to obtain a first number of detection regions. The sizes of the detection regions are the same. Input each of the detection regions into the detection model; the patch partitioning unit splices the values at the same positions in each of the detection regions to obtain a second number of spliced patches. The second number is the same as the number of positions included in the detection region. The patch partitioning unit stacks the spliced patches in the dimension of the number of channels to obtain stacked patches. The patch partitioning unit inputs the stacked patches into the encoder.
[0013] Further, the encoder includes U coding units; for any one of the U coding units, the any one coding unit includes a preprocessing unit and a processing unit. For the first coding unit among the U coding units, the preprocessing unit in the first coding unit performs linear embedding processing on the stacked patches to reduce the dimension of the stacked patches, and converts the stacked patches after the linear embedding processing into a patch sequence corresponding to the stacked patches based on a preset language conversion rule, and inputs the patch sequence corresponding to the stacked patches into the processing unit in the first coding unit. The processing unit in the first coding unit extracts features from the patch sequence corresponding to the stacked patches, and outputs a feature map corresponding to the stacked patches in the first coding unit. For the u-th coding unit among the U coding units, the preprocessing unit in the u-th coding unit splices the values at the same positions in each patch feature output by the processing unit in the (u - 1)-th coding unit to obtain a third number of spliced features, and performs feature transformation processing on each of the spliced features based on a preset feature transformation rule to obtain transformed features corresponding to each of the spliced features, and inputs the transformed features corresponding to each of the spliced features into the processing unit in the u-th coding unit; where 1 < u ≤ U, and the third number is the same as the number of positions included in the patch features. The processing unit in the u-th coding unit extracts features from the transformed features corresponding to each of the spliced features, and outputs a feature map corresponding to the stacked patches in the u-th coding unit.
[0014] Further, the decoder includes U depthwise separable convolutional network layers connected in sequence; one depthwise separable convolutional network layer is connected to one processing unit; the middle part further includes U attention fusion mechanism networks; the output of the processing unit is input into the corresponding depthwise separable convolutional network layer through the corresponding attention fusion mechanism network; the first depthwise separable convolutional network layer among the U depthwise separable convolutional network layers connected in sequence is connected to the dual-path module.
[0015] Among them, the attention fusion mechanism includes a channel attention module, a spatial attention module, a first edge attention module, a second edge attention module, and a fusion module; the first edge attention module includes G i global attention branch, L iLocal attention branch and splicing processing module. Determine the feature map output by the processing unit as the target feature map. The channel attention module determines the channel attention weight matrix Mc corresponding to the target feature map based on the target feature map, and determines the first feature map corresponding to the target feature map based on the channel attention weight matrix Mc. The channel attention weight matrix Mc includes the weights of the target feature map under different numbers of channels; the first feature map is the target feature map under the number of channels with the highest weight. The spatial attention module outputs the attention weight matrix Ms corresponding to the first feature map based on the first feature map, and determines the second feature map corresponding to the first feature map based on the attention weight matrix Ms. The attention weight matrix Ms includes the weights of each feature region in the target feature map. The second feature map includes a preset number of feature regions with the highest weights in the target feature map. The G i The global attention branch obtains the global features in the second feature map in the global dimension, and the L i The local attention branch obtains the local features in the second feature map in the local dimension. The splicing processing module splices the global features in the second feature map and the local features in the second feature map, and determines the edge features in each feature after the splicing processing as the first edge features. The second edge attention module locates and models the boundaries in the second feature map in a preset noise environment to extract the edge features in the second feature map, and determines the edge features in the second feature map as the second edge features. The fusion module fuses the first edge features and the second edge features and outputs at least one feature map corresponding to the remote sensing image to be detected.
[0016] On the other hand, the present invention provides a detection model, including an encoder, an intermediate part, and a decoder; the intermediate part includes a dual-path module. The encoder is configured to output at least one feature map corresponding to the remote sensing image to be detected based on the remote sensing image to be detected; the remote sensing image to be detected includes at least one cultivated land plot. The dual-path module is configured to output a spliced feature map corresponding to the feature map output by the encoder; the spliced feature map includes multi-scale context information of the corresponding feature map; the dual-path module includes N layer module structures; the first layer module structure includes a first convolutional unit, a first splicing unit, and a dilated convolutional unit connected in sequence; the j-th layer module structure includes a first convolutional unit, a first splicing unit, a dilated convolutional unit, a second splicing unit, and a second convolutional unit connected in sequence; 1 < j < N; the N-th layer module structure includes a first convolutional unit and a first splicing unit connected in sequence. The dilation rates of the dilated convolutional units in different layers are different. The first splicing unit included in the j-th layer module structure is connected to the dilated convolutional unit included in the (j - 1)-th layer module structure; the first splicing unit included in the N-th layer module structure is connected to the dilated convolutional unit included in the (N - 1)-th layer module structure and the second convolutional unit included in the (N - 1)-th layer module structure respectively; the second splicing unit included in the second layer module structure is connected to the dilated convolutional unit included in the first layer module structure; the second splicing unit included in the f-th layer module structure is connected to the second convolutional unit included in the (f - 1)-th layer module structure; 2 < f < N. The first convolutional unit is configured to perform compression processing on the input feature map to reduce the number of channels of the input feature map; the second convolutional unit is configured to perform fusion on the input feature map to extract features while reducing the number of channels of the input feature map. The dilated convolutional unit is configured to obtain multi-scale context information corresponding to the input spliced feature map; both the first splicing unit and the second splicing unit are configured to perform splicing processing on the input content and output the corresponding spliced feature map. The decoder is configured to output the edge information of the at least one cultivated land plot based on the spliced feature maps corresponding to the respective feature maps in the at least one feature map output by the dual-path module.
[0017] Based on the above technical solutions, the present invention can also be improved as follows.
[0018] Further, for any one of the at least one feature map, the first splicing unit included in the first-layer module structure is configured to splice the any one of the feature maps and the output of the first convolutional unit included in the first-layer module structure, and output a first spliced map corresponding to the any one of the feature maps in the first-layer module structure; the number of channels of the first spliced map corresponding to the any one of the feature maps in the first-layer module structure is equal to the sum of the number of channels of the any one of the feature maps and the number of channels of the feature map output by the first convolutional unit included in the first-layer module structure. The first splicing unit included in the j-th layer module structure is configured to splice the output of the first convolutional unit included in the j-th layer module structure and the output of the dilated convolutional unit included in the j-1-th layer module structure, and output a first spliced map corresponding to the any one of the feature maps in the j-th layer module structure; the number of channels of the first spliced map corresponding to the any one of the feature maps in the j-th layer module structure is equal to the sum of the number of channels of the feature map output by the first convolutional unit included in the j-th layer module structure and the number of channels of the multi-scale context information output by the dilated convolutional unit included in the j-1-th layer module structure. The first splicing unit included in the N-th layer module structure is configured to splice the output of the first convolutional unit included in the N-th layer module structure, the output of the dilated convolutional unit included in the N-1-th layer module structure, and the output of the second convolutional unit included in the N-1-th layer module structure, and output a spliced feature map corresponding to the any one of the feature maps; the number of channels of the spliced feature map corresponding to the any one of the feature maps is equal to the sum of the number of channels of the feature map output by the first convolutional unit included in the N-th layer module structure, the number of channels of the multi-scale context information output by the dilated convolutional unit included in the N-1-th layer module structure, and the number of channels of the feature map output by the second convolutional unit included in the N-1-th layer module structure.
[0019] Further, the second splicing unit included in the second-layer module structure is used to splice the output of the dilated convolution unit included in the first-layer module structure and the output of the dilated convolution unit included in the second-layer module structure, and output the second spliced graph corresponding to any one of the feature maps in the second-layer module structure; the number of channels of the second spliced graph corresponding to any one of the feature maps in the second-layer module structure is equal to the sum of the number of channels of the multi-scale context information output by the dilated convolution unit included in the first-layer module structure and the number of channels of the multi-scale context information output by the dilated convolution unit included in the second-layer module structure. The second splicing unit included in the f-th layer module structure is used to splice the output of the dilated convolution unit included in the f-th layer module structure and the output of the second convolution unit included in the (f - 1)-th layer module structure, and output the second spliced graph corresponding to any one of the feature maps in the f-th layer module structure; the number of channels of the second spliced graph corresponding to any one of the feature maps in the f-th layer module structure is equal to the sum of the number of channels of the multi-scale context information output by the dilated convolution unit included in the f-th layer module structure and the number of channels of the feature map output by the second convolution unit included in the (f - 1)-th layer module structure.
[0020] In a third aspect, the present invention provides an electronic device, including: a memory, one or more processors; the memory and the processor are coupled; wherein, computer program code is stored in the memory, and the computer program code includes computer instructions, when the computer instructions are executed by the processor, the electronic device is enabled to execute a method for extracting cultivated land plots from remote sensing images according to any one of the above first aspects.
[0021] In a fourth aspect, a computer-readable storage medium is provided, including computer instructions, when the computer instructions are run on an electronic device, the electronic device is enabled to execute a method for extracting cultivated land plots from remote sensing images according to any one of the above first aspects.
[0022] In a fifth aspect, a computer program product is provided, when the computer program product is run on a computer, the computer is enabled to execute a method for extracting cultivated land plots from remote sensing images according to any one of the above first aspects.
[0023] It can be understood that the beneficial effects that can be achieved by the detection model in the second aspect, the electronic device in the third aspect, the computer-readable storage medium in the fourth aspect, and the computer program product in the fifth aspect provided above can refer to the beneficial effects in the first aspect and any of its possible design manners, which will not be elaborated here. Description of the Drawings
[0024] Figure 1 It is a schematic flowchart of a method for extracting cultivated land plots from remote sensing images provided by the present invention;
[0025] Figure 2 Schematic structural diagram of the detection model provided by the present invention;
[0026] Figure 3 Schematic diagram of the remote sensing image and edge detection label of the cultivated land plot provided by the present invention;
[0027] Figure 4 Schematic structural diagram of the dual-path module provided by the present invention;
[0028] Figure 5 Schematic flow diagram of processing the remote sensing image and cultivated land label included in the data set D into the format required by the detection model provided by the present invention;
[0029] Figure 6 Schematic structural diagram of the attention fusion mechanism network provided by the present invention. Detailed implementation manners
[0030] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. Among them, in the description of the present application, unless otherwise specified, " / " means that the objects associated before and after are in an "or" relationship. For example, A / B may represent A or B; the "and / or" in the present application is only a description of the association relationship of the associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. These three situations, where A and B may be singular or plural. And, in the description of the present application, unless otherwise specified, "a plurality of" means two or more than two. "At least one (piece)" or its similar expression below refers to any combination of these items, including any combination of single item (piece) or plural items (pieces). For example, at least one (piece) of a, b, or c may represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c may be single or multiple. In addition, in order to clearly describe the technical solutions in the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that the words such as "first" and "second" do not limit the quantity and execution order, and the words such as "first" and "second" do not necessarily limit to be different. At the same time, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate as an example, illustration or explanation. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions.
[0031] With the advancement of agricultural modernization, the society's demand for land resources is constantly rising, and the precision requirements for farmland management are getting higher and higher. This requires more advanced technologies to achieve the automatic recognition and edge extraction of farmland plots. Among them, the rapid development of high-resolution remote sensing image technology provides a more abundant and detailed data source for the edge extraction of cultivated land plots, making it possible to improve the extraction accuracy of farmland management.
[0032] Currently, the edge detection models constructed based on traditional machine learning algorithms have certain subjectivity and limitations in feature selection and classifier design. For complex and variable agricultural environments, especially in areas where the boundary between cultivated land and other land features is blurred, the extraction results are prone to errors. That is, although the existing technologies can extract edge information, they often contain too many irrelevant details, affecting the accuracy. When adding structures to fuse features, it may cause low-level features to completely mask local details, or the adaptability to complex agricultural environments is not high enough. There is currently no good solution to optimize the edge position of cultivated land plots accurately perceived by the edge detection task for remote sensing image extraction results.
[0033] To address the above problems, the present invention proposes a method for extracting cultivated land plots from remote sensing images, aiming to improve the accuracy and robustness of edge detection of cultivated land plots by combining a detection model based on a dilated convolutional dual-path module and an attention fusion mechanism.
[0034] See Figure 1 The method for extracting cultivated land plots from remote sensing images provided by the present invention includes the following steps S101 - S102.
[0035] S101: Determine the remote sensing image to be detected.
[0036] Among them, the remote sensing image to be detected includes at least one cultivated land plot.
[0037] S102: Input the remote sensing image to be detected into the detection model to obtain the edge information of at least one cultivated land plot.
[0038] See Figure 2 which is a schematic structural diagram of the detection model provided by the present invention. The following will detail the method for constructing the Figure 2 detection model shown.
[0039] S210: Construct the edge detection of cultivated land plots from high-resolution remote sensing images.
[0040] Among them, the data set D includes the remote sensing images (referred to as images) of cultivated land plots as Figure 3 shown and edge detection labels.
[0041] S220: Based on a deep convolutional neural network, construct a detection model for the task of detecting the edges of cultivated land plots from remote sensing images.
[0042] As Figure 2 shown, the detection model provided by the present invention includes an encoder, an intermediate part, and a decoder. Among them, the intermediate part includes a dual-path module.
[0043] In some embodiments, the BiFormer dynamic sparse attention mechanism is introduced into the encoder, and the encoder is used to output at least one feature map corresponding to the input remote sensing image based on the input remote sensing image.
[0044] The dual-path module is used to output a spliced feature map corresponding to the feature map output by the encoder. Among them, the spliced feature map includes multi-scale context information of the corresponding feature map.
[0045] In some embodiments, referring to Figure 4 , the dual-path module includes N layer module structures. The first layer module structure includes a first convolutional unit, a first splicing unit, and a dilated convolutional unit connected in sequence. The j-th layer module structure includes a first convolutional unit, a first splicing unit, a dilated convolutional unit, a second splicing unit, and a second convolutional unit connected in sequence. 1 < j < N, and both N and J are integers. The N-th layer module structure includes a first convolutional unit and a first splicing unit connected in sequence.
[0046] The first splicing unit included in the j-th layer module structure is connected to the dilated convolutional unit included in the j-1-th layer module structure. The first splicing unit included in the N-th layer module structure is respectively connected to the dilated convolutional unit included in the N-1-th layer module structure and the second convolutional unit included in the N-1-th layer module structure. The second splicing unit included in the second layer module structure is connected to the dilated convolutional unit included in the first layer module structure. The second splicing unit included in the f-th layer module structure is connected to the second convolutional unit included in the f-1-th layer module structure. 2 < f < N, and f is an integer.
[0047] Among them, the first convolutional unit is used to compress the input feature map to reduce the number of channels of the input feature map. The second convolutional unit is used to fuse the input feature map to extract features while reducing the number of channels of the input feature map. The dilated convolutional unit is used to obtain multi-scale context information corresponding to the input spliced feature map. Both the first splicing unit and the second splicing unit are used to splice the input content and output the corresponding spliced feature map.
[0048] In some embodiments, the value by which the number of channels of the feature map passing through the first convolutional unit is reduced is the first value. The value by which the number of channels of the feature map passing through the second convolutional unit is reduced is the second value. The second value is greater than the first value.
[0049] In some embodiments, the first splicing unit included in the first-layer module structure is configured to splice any input feature map (for short, any feature map) and the output of the first convolutional unit included in the first-layer module structure, and output the first spliced map corresponding to any feature map in the first-layer module structure. The number of channels of the first spliced map corresponding to any feature map in the first-layer module structure is equal to the sum of the number of channels of any feature map and the number of channels of the feature map output by the first convolutional unit included in the first-layer module structure. The first splicing unit included in the j-th layer module structure is configured to splice the output of the first convolutional unit included in the j-th layer module structure and the output of the dilated convolutional unit included in the (j - 1)-th layer module structure, and output the first spliced map corresponding to any feature map in the j-th layer module structure. Wherein, the number of channels of the first spliced map corresponding to any feature map in the j-th layer module structure is equal to the sum of the number of channels of the feature map output by the first convolutional unit included in the j-th layer module structure and the number of channels of the multi-scale context information output by the dilated convolutional unit included in the (j - 1)-th layer module structure. The first splicing unit included in the N-th layer module structure is configured to splice the output of the first convolutional unit included in the N-th layer module structure, the output of the dilated convolutional unit included in the (N - 1)-th layer module structure, and the output of the second convolutional unit included in the (N - 1)-th layer module structure, and output the spliced feature map corresponding to any feature map. The number of channels of the spliced feature map corresponding to any feature map is equal to the sum of the number of channels of the feature map output by the first convolutional unit included in the N-th layer module structure, the number of channels of the multi-scale context information output by the dilated convolutional unit included in the (N - 1)-th layer module structure, and the number of channels of the feature map output by the second convolutional unit included in the (N - 1)-th layer module structure.
[0050] In some embodiments, the second splicing unit included in the second-layer module structure is configured to splice the output of the dilated convolutional unit included in the first-layer module structure and the output of the dilated convolutional unit included in the second-layer module structure, and output the second spliced map corresponding to any feature map in the second-layer module structure. Wherein, the number of channels of the second spliced map corresponding to any feature map in the second-layer module structure is equal to the sum of the number of channels of the multi-scale context information output by the dilated convolutional unit included in the first-layer module structure and the number of channels of the multi-scale context information output by the dilated convolutional unit included in the second-layer module structure.
[0051] In some embodiments, the second splicing unit included in the f-th layer module structure is used to splice the output of the dilated convolution unit included in the f-th layer module structure and the output of the second convolution unit included in the (f - 1)-th layer module structure, and output the corresponding second spliced map of any feature map in the f-th layer module structure. Among them, the number of channels of the corresponding second spliced map of any feature map in the f-th layer module structure is equal to the sum of the number of channels of the multi-scale context information output by the dilated convolution unit included in the f-th layer module structure and the number of channels of the feature map output by the second convolution unit included in the (f - 1)-th layer module structure.
[0052] It can be seen that the dual-path module can compress and integrate each input feature to improve the information content and effectiveness of the edge features extracted by the encoder.
[0053] In some embodiments, in order to avoid the problem of losing edge detail information when directly using dilated convolution, a feature hierarchical compression path is introduced in the dual-path module, and the deep features extracted by the encoder are provided with compressed features for each layer of dilated convolution through multiple first convolution units. Among them, the first convolution unit can include a 1x1 convolution. The last node before the output of the first convolution unit does not pass through the dilated convolution unit, so that the edge information extracted by the encoder can be retained to the greatest extent, ensuring that the original information and the information processed by the dilated convolution provide sufficient information for the decoder.
[0054] In some embodiments, the dilated convolution unit can include a 3x3 dilated convolution. And, the dilation rates of different dilated convolutions are different. For example, the dilation rate of the first layer of dilated convolution is 1, the dilation rate of the second layer of dilated convolution is 2, the dilation rate of the third layer of dilated convolution is 4, the dilation rate of the fourth layer of dilated convolution is 8, etc. Dilated convolutions with different dilation rates can capture multi-scale context information, enhance the perception ability of the detection model for the edges of cultivated land plots by means of a large receptive field, so as to capture richer context information and spatial relationships.
[0055] In some embodiments, the dual-path module includes a horizontal Bottleneck path and a vertical feature hierarchical compression and iterative fusion path. Among them, the main feature of the horizontal Bottleneck structure is that its number of channels first decreases and then increases in the middle layer. This structure realizes the reduction and increase of the number of channels by using 1x1 convolutions. First, a 1x1 convolution is used to reduce the number of channels (dimensionality reduction), which is connected to the dilated convolution unit through the first splicing unit, and then connected to a 1x1 convolution through the second splicing unit to restore to the original number of channels (dimensionality increase). By reducing the number of channels in the middle layer, the Bottleneck structure can significantly reduce the computational amount and the size of the model, while maintaining or even improving the performance of the model. The vertical feature hierarchical compression and iterative fusion path can use 1×1 convolutions to fuse the lateral output of the current stage and the result of the previous fusion to obtain a new fusion result.
[0056] In some embodiments, the decoder includes U depthwise separable convolutional network layers connected in sequence. A depthwise separable convolutional network layer is connected to a processing unit. The middle part may further include U attention fusion mechanism networks. The output of the processing unit is input to the corresponding depthwise separable convolutional network layer through the corresponding attention fusion mechanism network. The first depthwise separable convolutional network layer is connected to the dual-path module. The decoder is configured to output the edge information of at least one cultivated land plot based on the stitched feature maps corresponding to the feature maps in at least one feature map output by the dual-path module.
[0057] It should be noted that the middle part includes multiple skip connections between the encoder and the decoder. The bottom skip connection includes a dual-path module, and the other upper skip connections include attention fusion modules.
[0058] In some embodiments, the attention fusion mechanism includes a channel attention module, a spatial attention module, a first edge attention module, a second edge attention module, and a fusion module. The first edge attention module includes a G i global attention branch, an L i local attention branch, and a stitching processing module. The feature map output by the processing unit can be determined as the target feature map. The channel attention module can determine the channel attention weight matrix Mc corresponding to the target feature map based on the target feature map, and determine the first feature map corresponding to the target feature map based on the channel attention weight matrix Mc. The channel attention weight matrix Mc includes the weights of the target feature map at different numbers of channels. The first feature map is the target feature map at the number of channels with the highest weight. The spatial attention module outputs the attention weight matrix Ms corresponding to the first feature map based on the first feature map, and determines the second feature map corresponding to the first feature map based on the attention weight matrix Ms. The attention weight matrix Ms includes the weights of the respective feature regions in the target feature map. The second feature map includes a preset number of feature regions with the highest weights in the target feature map. The G i global attention branch obtains the global features in the second feature map in the global dimension. The L iThe local attention branch obtains local features in the second feature map in the local dimension. The splicing processing module splices the global features in the second feature map and the local features in the second feature map, and determines the edge features in each feature after splicing processing as the first edge features. The second edge attention module locates and models the boundaries in the second feature map under a preset noise environment to extract the edge features in the second feature map, and determines the edge features in the second feature map as the second edge features. Specifically, the second edge attention module gradually optimizes a variable field through a series of dense and repeated local attention operations, and this field contains local boundary information around each pixel. The fusion module fuses the first edge features and the second edge features and outputs at least one feature map corresponding to the remote sensing image to be detected.
[0059] S230: Based on the dataset D, train the detection model until the detection model reaches the preset accuracy.
[0060] In some embodiments, the remote sensing images and cultivated land labels included in the dataset D can be input into the detection model to train the detection model. Specifically, the remote sensing images and cultivated land labels included in the dataset D can be cut into pieces to obtain each detection area. Among them, the widths and heights of the segmented remote sensing images and cultivated land labels are the same. Input multiple groups of training data into the detection model to complete the training.
[0061] Exemplarily, refer to Figure 5 that processing the remote sensing images and cultivated land labels included in the dataset D into the format required by the detection model may include the following steps S501 - S504:
[0062] S501: Assume the remote sensing images (which can also be called original images in the embodiments of the present application) and cultivated land labels included in the dataset D.
[0063] Among them, the height of the remote sensing image is h, the width is w, and the number of channels is c; the height of the cultivated land label is h, the width is w, and the number of channels is 1.
[0064] S502: Calculate whether each original image needs to be expanded in size according to the set size and overlapping size. If expansion is required, calculate the pixels of the expanded image at the same time, and expand the original image and the corresponding cultivated land plot edge label to the specified size at the same time.
[0065] S503: Perform sliding cropping on the original image and the cultivated land plot edge label image according to the set size and overlapping size.
[0066] S504: After the cropping work is completed, obtain the original image building cultivated land plot dataset, and divide the cultivated land dataset into a training set and a test set according to a ratio of 8:2 based on the original remote sensing image.
[0067] Among them, 80% of the data can be used for training, and 20% of the data can be used to test the performance after training.
[0068] In some embodiments, the detection model further includes a patch partitioning unit. After inputting the remote sensing images included in the dataset D into the detection model, the patch partitioning unit can splice the values at the same positions in each detection region to obtain a second number of spliced patches. Among them, the second number is the same as the number of positions included in the detection region. The patch partitioning unit stacks the spliced patches according to the dimension of the number of channels to obtain stacked patches. The patch partitioning unit inputs the stacked patches into the encoder.
[0069] In some embodiments, the encoder includes U encoding units. Any encoding unit includes a pre-unit and a processing unit. After the patch partitioning unit inputs the stacked patches into the encoder, the pre-unit in the first encoding unit performs linear embedding processing on the stacked patches to reduce the dimension of the stacked patches, and converts the stacked patches after the linear embedding processing into a patch sequence corresponding to the stacked patches based on a preset language conversion rule, and inputs the patch sequence corresponding to the stacked patches into the processing unit in the first encoding unit. The processing unit in the first encoding unit extracts features from the patch sequence corresponding to the stacked patches and outputs the feature map corresponding to the stacked patches in the first encoding unit. The pre-unit in the u-th encoding unit splices the values at the same positions in each patch feature output by the processing unit in the (u - 1)-th encoding unit to obtain a third number of spliced features, and performs feature transformation processing on each spliced feature based on a preset feature transformation rule to obtain the transformed features corresponding to each spliced feature, and inputs the transformed features corresponding to each spliced feature into the processing unit in the u-th encoding unit. Where 1 < u ≤ U and u is an integer. The third number is the same as the number of positions included in the patch features. The processing unit in the u-th encoding unit extracts features from the transformed features corresponding to each spliced feature and outputs the feature map corresponding to the stacked patches in the u-th encoding unit.
[0070] In some embodiments, the processing unit can process the input content based on BiFormer. The process of BiFormer processing is described in detail below.
[0071] S701: Construct an inter-region directed graph routing to find the association relationship (i.e., the regions that each given region should be associated with) by constructing a directed graph. Specifically, first derive Q and K, Qr, by applying the average value of each region to Q and K respectively, and then derive the adjacency matrix, draw the region-to-region affinity graph A through the matrix multiplication between Qr and Kr r = Q r (K r) T :
[0072] S702: Measure the degree of semantic relatedness between two regions in the adjacency matrix Ar. Next, prune the affinity graph by retaining only the top k connections for each region. Specifically, use the row-wise topk operator to derive the index matrix, I r = topIndex(A r ), so the i-th row of Ir contains the k indices of the most relevant regions for the i-th region.
[0073] S703: Utilize the region-to-region routing index matrix I r , and apply fine-grained token-to-token attention. For each query token in region i, it will attend to all the K-V pairs residing in the union of the k routing regions indexed by I r (i,1) , I r (i,2) ,..., I r (i,k) index.
[0074] S704: Collect the K and V tensors, i.e.: V g = gather(V, I r ), K g = gather(K, I r ), where are the gathered K and V tensors, and then apply attention to the gathered K-V pairs, with the formula O = Attention(Q, K g , V g ) + LCE(V) to implement the calculation of attention. Here, a local context enhancement term LCE(V) is introduced. The function LCE(.) is parameterized by depth convolution and the kernel size is set to 5. Thus, the model can better understand the global context information in the image and make full use of the parallel computing power of the GPU to improve performance in tasks such as image classification, object detection, and semantic segmentation.
[0075] In some embodiments, the factors affecting the output feature map size of the processing unit are the size (H, W) of the input feature map, the size (FH, FW) of the convolutional kernel, the output size is (OH, OW), the padding is P, and the stride is S. The formula for the output size is:
[0076]
[0077]
[0078] In some embodiments, both the first convolutional unit and the second convolutional unit may include 1x1 convolutions. For each layer of the module structure in the dual-path module, first, a 1×1 convolution can be used to reduce the number of channels of the input feature map (dimensionality reduction), then a dilated convolution, and finally a 1x1 convolution to restore to the original number of channels (dimensionality increase). Computational efficiency: By reducing the number of channels in the intermediate layer, each layer of the module structure can significantly reduce the amount of computation and the size of the model, while maintaining or even improving the performance of the model. Regarding the calculation of the number of channels, the input is Set the output result of the 1×1 convolution for feature compression to Each concatenation operation adds the number of input channels. Through the processing of the dilated convolution, the size and number of channels of the output of the dilated convolution are the same as those of the input. For the 1×1 convolutions for iterative fusion of different layers, the set output results are different, and are respectively The finally processed output result is
[0079] In some embodiments, for the processing of the dilated convolution part, it is analyzed that the same 3×3 convolution can achieve the effects of 5×5, 7×7, etc. convolutions. The dilated convolution can increase the receptive field without increasing the number of parameters (number of parameters = convolution kernel size + bias). Assuming the convolution kernel size of the dilated convolution is k and the dilation rate is d, then its equivalent convolution kernel size k'. For example, for a 3×3 convolution kernel, k = 3, and the formula is k' = k + (k - 1)×(d - 1). The formula for the receptive field of the current layer is as follows: RF i+1 = RF i +(k′ - 1)×S i , where RF i+1 represents the receptive field of the current layer, RF i represents the receptive field of the previous layer, k′ represents the size of the convolution kernel, and S i represents the product of the strides of all previous layers (excluding this layer), and the formula is Similarly, the stride of the current layer does not affect the receptive field of the current layer, and the receptive field has nothing to do with padding.
[0080] In some embodiments, through the attention fusion mechanism network in this model, channel attention and spatial attention can be passed in sequence, and then through a parallel edge attention mechanism, such as Figure 6 . Among them, the channel attention weight matrix M c can be expressed as M c = R c×1×1 . To reduce the computational parameters, a dimensionality reduction coefficient r is adopted in the MLP. M c ∈R C / r×1×1 , so the channel attention calculation formula is:
[0081]
[0082] In the above formula, and respectively represent the global average pooling feature and the max pooling feature.
[0083] In some embodiments, the spatial attention weight matrix M s , can be expressed as M s (F) ∈ R H,W , similarly, two pooling methods are used in the channel dimension to generate a 2D feature map: Therefore, the calculation formula for spatial attention is:
[0084]
[0085]
[0086] The result is processed through two parallel edge attention mechanisms (such as Figure 6 ), the parallel edge attention 1 and edge attention 2), focusing on edge-related information, and using the attention mechanism to filter the information in the feature map transmitted from the spatial attention branch. Each filtering can filter out some edge-irrelevant information, so that the finally retained information only contains edge-related information.
[0087] In some embodiments, the design details of the edge attention module 1 are as Figure 6 shown. Use G i and L i to represent the parts of the input feature of the edge attention module from the global attention branch and the local attention branch. First, G and L i are concatenated to obtain a new feature map, and then the concatenated feature map is fed into a gated convolution to filter out edge-irrelevant information and obtain an intermediate feature map T i . After that, T i and L i are multiplied pixel by pixel once to suppress those non-edge responses in L, and finally a convolutional layer is used to smooth the result of the multiplication to obtain the final output O. G i will adopt bilinear interpolation to increase the resolution to be consistent with L i before being fed into the edge attention module.
[0088] In ordinary convolution, for all channels of the convolution kernel at a spatial position, the sizes are the same. The calculation method of the output value O y,x at the position (y, x) of the ordinary convolution operation is given as: where I y+i,x+j ∈ RC represents the input of the convolution operation, k h and k w represent the size of the convolution kernel, represents the convolution kernel, In the structure of this module, gated convolution is used instead of ordinary convolution. The calculation process of gated convolution can be described by the following formula: O y,x = φ(Feature y,x ) ⊙ σ(Gating y,x ), where ⊙ represents performing element-wise multiplication, Feature represents dynamic feature selection, Gating represents the output gating value, σ represents using the Sigmoid function to normalize the output gating value between 0 and 1, and φ refers to any activation function such as the rectified linear unit, exponential linear unit, or leaky rectified linear unit, etc. Feature y,x and Gating y,x are calculated as follows: Gating y,x = ∑∑W g ·I, Feature y,x = ∑∑W f ·I, where W g and W f represent different convolution kernels respectively, and I represents the input of the convolutional layer.
[0089] In some embodiments, in combination with Figure 6 , while the edge attention 1 is being processed, the edge attention 2 module is also operating. This structure implicitly uses the high-dimensional embedding of the field elements, γ t [k] ∈ R D γ and T t [k] ∈ R Dπ , which can be decoded using the learned linear mappings γ → g and π → p at any iteration. The model uses D γ = 64 and D π = 8. Given an input image, the network first applies a "neighborhood MLP-mixer", which is an improved MLP-Mixer that replaces the global spatial operation with a convolution with a kernel size of 3. Another variation is that the input pixels are mapped to the hidden state size through a pixel-level linear mapping instead of using the input patches. This block converts the input image into a pixel-resolution map channel with D γ , denoted by γ0[n], and refers to it as the initial "hidden state".
[0090] Among them, this hidden state can be refined through an iterative sequence of eight boundary attentions. This process is divided into two boundary attention blocks, each with learned weights. To process the input, a linear projection of the initial hidden state is first added. This is essentially a skip connection that allows the network to retain information from the initial hidden state estimate in the later stages of processing. Next, the hidden state is copied into two identical parts. A learned window of dimension 8 is embedded into one copy, and the input image plus the current estimate of the smoothed global feature is added to the other. Then, neighborhood cross-attention is performed, where each pixel inside, the first copy performs two cross-attention iterations with an 11-sized patch of the second copy. A learned 11×11 positional encoding is added to the patch, which enables the network to access relative positioning even without global position cues. A small MLP is used to track each self-attention layer. A simple linear mapping is used to transform the output or intermediate hidden state into the knot space and render the output image. The window embedding (last 8 dimensions) is separated from the knot embedding (first 64 dimensions) and projected through a linear layer to map the hidden state to 7 numbers, which represent g=(u,sin(θ),cos(θ),ω), and this data can be used as the input for the collection and slicing operators.
[0091] In some embodiments, the result processed by the attention mechanism is passed through a decoder based on depthwise separable convolution. Focusing on capturing spatial features such as edges and textures, it can effectively fuse the multi-scale feature maps received by the decoder from the encoder and gradually restore the spatial resolution of the image, contributing to the reconstruction of high-quality outputs.
[0092] In some embodiments, the performance of the model can be measured based on the result output by the detection model through the loss function of this network model. This loss function consists of three losses: the tracking (cross-entropy) loss l t 、the boundary tracking loss l bt and the texture suppression loss l txs . Therefore, the resulting loss function loss2(I) is calculated as l = l t +α bt ×l bt +α txs ×l txs , where α bt is the weight for regularizing the boundary tracking loss, and α txs is the texture suppression loss predicted by this network model. The final loss is the sum of the losses calculated according to each . For the tracking (cross-entropy) loss l t loss, it is defined where w is the weight of the tracking loss, and Y- and Y+ represent the negative and positive marginal samples in the given Y respectively. Regarding the boundary tracking loss (l bt ), it is defined as follows where E is the marginal point of the given Y, represents the edge tile containing the edge segment from , the center of which is P, and the marginal points in are represented by D p . Finally, the texture loss (l txs ) is defined as where, is the edge mapping block centered on the non-edge point p, while is the set including all the edges and the confused pixels used in l bt .
[0093] It can be seen that the detection model provided by the present invention adopts a depthwise separable convolution in the decoder stage, focusing on capturing spatial features such as edges and textures, and can effectively fuse the multi-scale feature maps received by the decoder from the encoder and gradually restore the spatial resolution of the image, which helps to obtain a high-quality output. At the same time, combined with the design of the loss function, it ensures that the task can be fully trained and optimized.
[0094] After the detection model training reaches the preset accuracy, the trained detection model can be used to extract cultivated land plots from actual remote sensing images, obtaining a cultivated land plot extraction result with more accurate edges and richer details. It not only improves the accuracy and robustness of cultivated land plot extraction, but also provides more reliable data support for subsequent applications such as agricultural resource management and planting planning.
[0095] In some solutions, multiple embodiments of the present application can be combined and the combined solution can be implemented. Optionally, some operations in the processes of the method embodiments are optionally combined, and / or the order of some operations is optionally changed. And, the execution order between the steps of each process is only exemplary and does not constitute a limitation on the execution order between the steps. The steps can also be in other execution orders. It is not intended to indicate that the execution order is the only order in which these operations can be performed. Those of ordinary skill in the art will think of various ways to reorder the operations described herein. In addition, it should be noted that the process details involved in a certain embodiment herein are equally applicable to other embodiments in a similar manner, or different embodiments can be combined and used.
[0096] In addition, some steps in the method embodiments can be equivalently replaced with other possible steps. Or, some steps in the method embodiments can be optional and can be deleted in some usage scenarios. Or, other possible steps can be added to the method embodiments. Moreover, the method embodiments can be implemented independently or in combination with each other.
[0097] From the description of the above embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of the system is divided into different functional modules to complete all or part of the functions described above.
[0098] In several embodiments provided in the present application, it should be understood that the disclosed system and method can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the system or unit can be in electrical, mechanical or other forms.
[0099] In addition, each functional unit in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0100] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that makes a contribution, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs and other various media that can store program codes.
[0101] The above content is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application shall be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A method for extracting cultivated land plots from remote sensing images, characterized in that, Including: Determine the remote sensing image to be detected; the remote sensing image to be detected includes at least one cultivated land plot; Input the remote sensing image to be detected into the detection model to obtain the edge information of the at least one cultivated land plot; Wherein, the detection model includes an encoder, a middle part and a decoder; the middle part includes a dual-path module; The encoder is used to output at least one feature map corresponding to the remote sensing image to be detected based on the remote sensing image to be detected; The dual-path module is used to output a spliced feature map corresponding to the feature map output by the encoder based on the feature map output by the encoder; the spliced feature map includes multi-scale context information of the corresponding feature map; the dual-path module includes N layer module structures; the first layer module structure includes a first convolution unit, a first splicing unit and a dilated convolution unit connected in sequence; the j-th layer module structure includes a first convolution unit, a first splicing unit, a dilated convolution unit, a second splicing unit and a second convolution unit connected in sequence; 1 < j < N; the N-th layer module structure includes a first convolution unit and a first splicing unit connected in sequence; the dilation rates of the dilated convolution units in different layers are different; The first splicing unit included in the j-th layer module structure is connected to the dilated convolution unit included in the j-1-th layer module structure; the first splicing unit included in the N-th layer module structure is respectively connected to the dilated convolution unit included in the N-1-th layer module structure and the second convolution unit included in the N-1-th layer module structure; the second splicing unit included in the second layer module structure is connected to the dilated convolution unit included in the first layer module structure; the second splicing unit included in the f-th layer module structure is connected to the second convolution unit included in the f-1-th layer module structure; 2 < f < N; Both the first convolution unit and the first splicing unit included in the first layer module structure are used to receive the feature map output by the encoder; the first convolution unit is further used to compress the input feature map to reduce the number of channels of the input feature map; the second convolution unit is used to fuse the input feature map to extract features while reducing the number of channels of the input feature map; the dilated convolution unit is used to obtain multi-scale context information corresponding to the input spliced feature map; both the first splicing unit and the second splicing unit are used to splice the input content and output the corresponding spliced feature map; The decoder is used to output the edge information of the at least one cultivated land plot based on the spliced feature maps corresponding to the respective feature maps in the at least one feature map output by the dual-path module.
2. The method according to claim 1, wherein The value by which the number of channels of the feature map passing through the first convolution unit is reduced is the first value; the value by which the number of channels of the feature map passing through the second convolution unit is reduced is the second value; the second value is greater than the first value.
3. The method according to claim 2, wherein For any one of the at least one feature map, the first splicing unit included in the first-layer module structure is configured to splice the any one of the feature maps and the output of the first convolutional unit included in the first-layer module structure, and output a first spliced map corresponding to the any one of the feature maps in the first-layer module structure; the number of channels of the first spliced map corresponding to the any one of the feature maps in the first-layer module structure is equal to the sum of the number of channels of the any one of the feature maps and the number of channels of the feature map output by the first convolutional unit included in the first-layer module structure; The first splicing unit included in the j-th layer module structure is configured to splice the output of the first convolutional unit included in the j-th layer module structure and the output of the dilated convolutional unit included in the (j - 1)-th layer module structure, and output a first spliced map corresponding to the any one of the feature maps in the j-th layer module structure; the number of channels of the first spliced map corresponding to the any one of the feature maps in the j-th layer module structure is equal to the sum of the number of channels of the feature map output by the first convolutional unit included in the j-th layer module structure and the number of channels of the multi-scale context information output by the dilated convolutional unit included in the (j - 1)-th layer module structure; The first splicing unit included in the N-th layer module structure is configured to splice the output of the first convolutional unit included in the N-th layer module structure, the output of the dilated convolutional unit included in the (N - 1)-th layer module structure, and the output of the second convolutional unit included in the (N - 1)-th layer module structure, and output a spliced feature map corresponding to the any one of the feature maps; the number of channels of the spliced feature map corresponding to the any one of the feature maps is equal to the sum of the number of channels of the feature map output by the first convolutional unit included in the N-th layer module structure, the number of channels of the multi-scale context information output by the dilated convolutional unit included in the (N - 1)-th layer module structure, and the number of channels of the feature map output by the second convolutional unit included in the (N - 1)-th layer module structure.
4. The method according to claim 3, wherein The second splicing unit included in the second-layer module structure is configured to splice the output of the dilated convolutional unit included in the first-layer module structure and the output of the dilated convolutional unit included in the second-layer module structure, and output a second spliced map corresponding to the any one of the feature maps in the second-layer module structure; the number of channels of the second spliced map corresponding to the any one of the feature maps in the second-layer module structure is equal to the sum of the number of channels of the multi-scale context information output by the dilated convolutional unit included in the first-layer module structure and the number of channels of the multi-scale context information output by the dilated convolutional unit included in the second-layer module structure; The second splicing unit included in the f-th layer module structure is used to splice the output of the dilated convolution unit included in the f-th layer module structure and the output of the second convolution unit included in the (f - 1)-th layer module structure, and output the second spliced graph corresponding to any feature map in the f-th layer module structure; the number of channels of the second spliced graph corresponding to any feature map in the f-th layer module structure is equal to the sum of the number of channels of the multi-scale context information output by the dilated convolution unit included in the f-th layer module structure and the number of channels of the feature map output by the second convolution unit included in the (f - 1)-th layer module structure.
5. The method according to claim 4, wherein The step of inputting the remote sensing image to be detected into the detection model includes: Performing segmentation processing on the remote sensing image to be detected to obtain a first number of detection regions; the sizes of all the detection regions are the same; Inputting all the detection regions into the detection model; The detection model further includes a patch partitioning unit. After inputting the remote sensing image to be detected into the detection model, the method further includes: The patch partitioning unit splices the values at the same positions in each of the detection regions to obtain a second number of spliced patches; the second number is the same as the number of positions included in the detection regions; The patch partitioning unit stacks the spliced patches according to the dimension of the number of channels to obtain stacked patches; The patch partitioning unit inputs the stacked patches into the encoder.
6. The method according to claim 5, wherein The encoder includes U encoding units; for any one of the U encoding units, any one of the encoding units includes a pre-unit and a processing unit; After the patch partitioning unit inputs the stacked patches into the encoder, the method includes: For the first encoding unit among the U encoding units, the pre-unit in the first encoding unit performs linear embedding processing on the stacked patches to reduce the dimension of the stacked patches, and converts the stacked patches after the linear embedding processing into a patch sequence corresponding to the stacked patches based on a preset language conversion rule, and inputs the patch sequence corresponding to the stacked patches into the processing unit in the first encoding unit; The processing unit in the first encoding unit extracts features from the patch sequence corresponding to the stacked patches and outputs the feature map corresponding to the stacked patches in the first encoding unit; For the u-th encoding unit among the U encoding units, the pre-unit in the u-th encoding unit splices the values at the same positions in each of the patch features output by the processing unit in the (u - 1)-th encoding unit to obtain a third number of spliced features, and performs feature transformation processing on each of the spliced features based on a preset feature transformation rule to obtain transformed features corresponding to each of the spliced features, and inputs the transformed features corresponding to each of the spliced features into the processing unit in the u-th encoding unit; where 1 < u ≤ U, and the third number is the same as the number of positions included in the patch features; The processing unit in the $u$-th coding unit extracts features from the transformation features corresponding to each of the splicing features, and outputs the feature map corresponding to the stacked patch in the $u$-th coding unit.
7. The method according to claim 6, characterized in that The decoder includes $U$ depthwise separable convolutional network layers connected in sequence; one depthwise separable convolutional network layer is connected to one processing unit; the middle part further includes $U$ attention fusion mechanism networks; the output of the processing unit is input into the corresponding depthwise separable convolutional network layer through the corresponding attention fusion mechanism network; the first depthwise separable convolutional network layer in the $U$ depthwise separable convolutional network layers connected in sequence is connected to the dual-path module; Among them, any one of the attention fusion mechanisms includes a channel attention module, a spatial attention module, a first edge attention module, a second edge attention module, and a fusion module; the first edge attention module includes G i a global attention branch, L i a local attention branch, and a splicing processing module; the output of the processing unit is input into the corresponding depthwise separable convolutional network layer through the attention fusion mechanism network, including: Determine the feature map output by the processing unit as the target feature map: The channel attention module determines the channel attention weight matrix $M_c$ corresponding to the target feature map based on the target feature map, and determines the first feature map corresponding to the target feature map based on the channel attention weight matrix $M_c$; the channel attention weight matrix $M_c$ includes the weights of the target feature map under different numbers of channels; the first feature map is the target feature map under the number of channels with the highest weight; The spatial attention module outputs the attention weight matrix $M_s$ corresponding to the first feature map based on the first feature map, and determines the second feature map corresponding to the first feature map based on the attention weight matrix $M_s$; the attention weight matrix $M_s$ includes the weights of each feature region in the target feature map; the second feature map includes a preset number of feature regions with the highest weights in the target feature map; The said G i The global attention branch obtains the global features in the second feature map in the global dimension; The L i The local attention branch obtains local features in the second feature map in the local dimension; The splicing processing module performs splicing processing on the global features and the local features in the second feature map, and determines the edge features in each feature after splicing processing as the first edge features; The second edge attention module locates and models the boundaries in the second feature map in a preset noise environment to extract the edge features in the second feature map, and determines the edge features in the second feature map as the second edge features; The fusion module fuses the first edge features and the second edge features, and outputs at least one feature map corresponding to the remote sensing image to be detected.
8. A detection model, characterized in that Including: An encoder, a middle part, and a decoder; the middle part includes a dual-path module; The encoder is used to output at least one feature map corresponding to the remote sensing image to be detected based on the remote sensing image to be detected; the remote sensing image to be detected includes at least one cultivated land plot; The dual-path module is used to output a spliced feature map corresponding to the feature map output by the encoder based on the feature map output by the encoder; the spliced feature map includes multi-scale context information of the corresponding feature map; the dual-path module includes N layer module structures; the first layer module structure includes a first convolutional unit, a first splicing unit, and an atrous convolutional unit connected in sequence; the j-th layer module structure includes a first convolutional unit, a first splicing unit, an atrous convolutional unit, a second splicing unit, and a second convolutional unit connected in sequence; 1 < j < N; the N-th layer module structure includes a first convolutional unit and a first splicing unit connected in sequence; the atrous rates of the atrous convolutional units in different layers are different; The first splicing unit included in the j-th layer module structure is connected to the atrous convolutional unit included in the (j - 1)-th layer module structure; the first splicing unit included in the N-th layer module structure is respectively connected to the atrous convolutional unit included in the (N - 1)-th layer module structure and the second convolutional unit included in the (N - 1)-th layer module structure; the second splicing unit included in the second layer module structure is connected to the atrous convolutional unit included in the first layer module structure; the second splicing unit included in the f-th layer module structure is connected to the second convolutional unit included in the (f - 1)-th layer module structure; 2 < f < N; The first convolutional unit and the first splicing unit included in the first layer module structure are both used to receive the feature map output by the encoder; the first convolutional unit is further used to perform compression processing on the input feature map to reduce the number of channels of the input feature map; the second convolutional unit is used to perform fusion on the input feature map to extract features while reducing the number of channels of the input feature map; the atrous convolutional unit is used to obtain multi-scale context information corresponding to the input spliced feature map; the first splicing unit and the second splicing unit are both used to perform splicing processing on the input content and output the corresponding spliced feature map; The decoder is used to output the edge information of the at least one cultivated land plot based on the spliced feature maps corresponding to the feature maps in the at least one feature map output by the dual-path module.
9. The detection model according to claim 8, wherein For any one of the at least one feature map, the first splicing unit included in the first layer module structure is used to perform splicing processing on the any one of the feature map and the output of the first convolutional unit included in the first layer module structure, and output a first spliced map corresponding to the any one of the feature map in the first layer module structure; the number of channels of the first spliced map corresponding to the any one of the feature map in the first layer module structure is equal to the sum of the number of channels of the any one of the feature map and the number of channels of the feature map output by the first convolutional unit included in the first layer module structure; The first splicing unit included in the j-th layer module structure is used to splice the output of the first convolutional unit included in the j-th layer module structure and the output of the dilated convolutional unit included in the (j - 1)-th layer module structure, and output the first spliced graph corresponding to the any feature map in the j-th layer module structure; the number of channels of the first spliced graph corresponding to the any feature map in the j-th layer module structure is equal to the sum of the number of channels of the feature map output by the first convolutional unit included in the j-th layer module structure and the number of channels of the multi-scale context information output by the dilated convolutional unit included in the (j - 1)-th layer module structure; The first splicing unit included in the N-th layer module structure is used to splice the output of the first convolutional unit included in the N-th layer module structure, the output of the dilated convolutional unit included in the (N - 1)-th layer module structure, and the output of the second convolutional unit included in the (N - 1)-th layer module structure, and output the spliced feature map corresponding to the any feature map; the number of channels of the spliced feature map corresponding to the any feature map is equal to the sum of the number of channels of the feature map output by the first convolutional unit included in the N-th layer module structure, the number of channels of the multi-scale context information output by the dilated convolutional unit included in the (N - 1)-th layer module structure, and the number of channels of the feature map output by the second convolutional unit included in the (N - 1)-th layer module structure.
10. The detection model according to claim 9, wherein The second splicing unit included in the second layer module structure is used to splice the output of the dilated convolutional unit included in the first layer module structure and the output of the dilated convolutional unit included in the second layer module structure, and output the second spliced graph corresponding to the any feature map in the second layer module structure; the number of channels of the second spliced graph corresponding to the any feature map in the second layer module structure is equal to the sum of the number of channels of the multi-scale context information output by the dilated convolutional unit included in the first layer module structure and the number of channels of the multi-scale context information output by the dilated convolutional unit included in the second layer module structure; The second splicing unit included in the f-th layer module structure is used to splice the output of the dilated convolutional unit included in the f-th layer module structure and the output of the second convolutional unit included in the (f - 1)-th layer module structure, and output the second spliced graph corresponding to the any feature map in the f-th layer module structure; the number of channels of the second spliced graph corresponding to the any feature map in the f-th layer module structure is equal to the sum of the number of channels of the multi-scale context information output by the dilated convolutional unit included in the f-th layer module structure and the number of channels of the feature map output by the second convolutional unit included in the (f - 1)-th layer module structure.
Citation Information
Patent Citations
Image segmentation method and device, terminal equipment and readable storage medium
CN114742700A
Defect detection system and method based on double-path attention network
CN117036290A