Terraced field extraction method and system based on deep learning semantic segmentation model
Through the terraced field extraction method based on deep learning semantic segmentation model, the problem of poor recognition ability of terraced field plots in complex geographical environments is solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510466190.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-11
AI Technical Summary
The prior art is difficult to accurately identify and divide terraced plots in complex geographical environments, such as mountainous, hilly and basin valley areas, resulting in poor recognition capabilities and insufficient accuracy.
The terraced field extraction method based on the deep learning semantic segmentation model is adopted. By constructing the terraced field data set, multi-stage convolution and parallel multi-branch processing are performed, multi-scale features are generated and fusion is performed, and the terraced field pixel probability distribution is obtained based on the category weights, thereby extracting terraced field plots.
It significantly improves the expression ability of crop characteristics of terraced plots and the recognition rate of terraced plots in complex geographical environments, and improves the accuracy and robustness of terraced plot extraction.
Smart Images

Figure CN119992345A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of terraced field image extraction, and in particular to a terraced field extraction method and system based on a deep learning semantic segmentation model. Background Art
[0002] Remote sensing technology has important applications in the agricultural field, and can achieve large-scale, high-efficiency monitoring of crop types, areas, and distribution. Terraced fields are important agricultural production sites, and accurate extraction of their area and distribution is of great significance for agricultural management and decision-making. Some terrace distribution areas are mainly composed of mountains, hills, and basins and valleys. Compared with large plains, the geographical environment of terrace distribution areas composed of mountains, hills, and basins and valleys is more complex, and higher requirements are placed on remote sensing intelligent extraction algorithms.
[0003] In recent years, with the development of remote sensing technology and deep learning technology, significant technological progress has been made, especially in the field of high-resolution remote sensing terraced field image recognition. Convolutional neural network (CNN) has shown superior performance in tasks such as image classification and segmentation. However, the existing deep learning-based cultivated land identification methods still have certain limitations. For example, in complex geographical environments such as mountains, hills and basins, the following situations often occur in these complex environments: 1. The shape and size of the planting area of the terraced fields are irregular, and they are often interlaced with other landforms, making it difficult for the model to accurately segment the terraced field targets; 2. Terraces are usually distributed in mountainous areas with large slopes, with undulating terrain and various slope directions. This will cause the shape and size of terraces in remote sensing images to be irregular, and the boundaries to be unclear, making it difficult for automatic extraction algorithms to accurately identify and segment terraces; 3. Due to terrain occlusion and changes in the sun's angle, shadows are prone to appear in the terraced area. Shadows reduce the contrast and clarity of the image, making it more difficult to distinguish the terraced fields from the surrounding objects.
[0004] In summary, this leads to problems such as poor terraced field recognition ability and insufficient accuracy. Summary of the invention
[0005] The purpose of the present invention is to provide a terrace extraction method and system based on a deep learning semantic segmentation model, aiming to solve the problem that traditional technologies have poor terrace plot recognition ability and insufficient accuracy due to the complex actual distribution environment of terraces.
[0006] In a first aspect, the present invention provides a method for extracting terraces based on a deep learning semantic segmentation model, the method comprising: Construct a dataset about terraced fields and train a semantic segmentation model based on the dataset, including: Performing a multi-level convolution operation on the data set to obtain a multi-level convolution feature, and performing parallel multi-branch processing on the convolution feature to obtain a first feature, a second feature, a third feature, and a fourth feature; Performing upsampling processing on the convolution feature to obtain a multi-level upsampling feature, generating a multi-scale feature according to the first feature, the second feature, the third feature, and the fourth feature, and fusing the multi-scale feature with the multi-level upsampling feature to obtain a fused feature; Splicing the fusion feature with the convolution feature to obtain a spliced feature, generating a category weight according to the spliced feature, and obtaining a terrace pixel probability distribution of each image in the data set according to the category weight; The image to be tested is input into the trained semantic segmentation model to obtain the target terrace pixel probability distribution corresponding to the image to be tested, and all terrace plots are extracted from the image to be tested according to the target terrace pixel probability distribution.
[0007] Furthermore, the step of constructing a data set about terraced fields includes: Acquire terraced field remote sensing images, and annotate the images in the terraced field remote sensing images frame by frame, wherein the annotated results include terraced field contours, terraced field surfaces, and backgrounds, and generate a data set based on the annotated images.
[0008] Furthermore, the step of performing a multi-level convolution operation on the data set to obtain a multi-level convolution feature includes: Multi-level convolution operations are performed according to the following formula: ; in, Indicates The composite operation of the stage includes convolution, normalization, and activation in sequence. , Respectively represent Level Level convolutional features, Represents the images in the dataset; The step of performing parallel multi-branch processing on the convolution features to obtain the first feature, the second feature, the third feature and the fourth feature comprises: The convolutional features are processed in parallel with multiple branches according to the following formula: ; in, , , , Respectively represent the first feature, the second feature, the third feature, and the fourth feature, The void rate is of Convolution, d represents the void rate, represents global average pooling, Indicates first Convolution, then upsampling, and adjusting the features in height and width to and , represents the height of the feature, Indicates the width of the feature.
[0009] Further, the step of performing upsampling processing on the convolution feature to obtain a multi-level upsampling feature, generating a multi-scale feature according to the first feature, the second feature, the third feature and the fourth feature, and fusing the multi-scale feature with the multi-level upsampling feature to obtain a fused feature includes: The upsampling process is performed according to the following formula: ; in, represents the adjusted i-th level convolution feature, , represents the i-th level up-sampled feature, , Represents bilinear times upsampling; The multi-scale features are generated according to the following formula: ; in, Represents multi-scale features, represents channel dimension splicing, express The feature tensor dimension of The height becomes , the width becomes , the number of channels becomes 256; The fusion is performed according to the following formula: ; in, represents the fusion feature, express convolution, , , They represent the up-sampled features of level 2, level 3, and level 4 respectively.
[0010] Furthermore, the step of splicing the fusion feature with the convolution feature to obtain a spliced feature, and generating a category weight according to the spliced feature includes: The class weights are generated according to the following formula: ; in, Represents the category weight vector, C represents the dimension of the vector, which is a positive integer. , represents a learnable weight matrix, represents an activation function, represents the rectified linear unit activation function, Represents the splicing feature, express The vector space where .
[0011] Furthermore, the step of obtaining the terrace pixel probability distribution of each image in the data set according to the category weight includes: Generate a spatial attention map based on the class weights: ; in, is the spatial attention map, Indicates size The matrix of all 1s, represents tensor product; Perform feature reweighting on the spatial attention map: ; in, represents the feature representation after feature reweighting by the spatial attention mechanism, represents element-wise multiplication; The probability distribution of terrace pixels is obtained according to the following formula: ; in, represents the probability distribution of terrace pixels, represents normalization processing, is a learnable weight matrix, Indicates bilinear 8x upsampling, is a learnable bias vector.
[0012] Furthermore, the step of inputting the image to be tested into the trained semantic segmentation model to obtain the target terrace pixel probability distribution corresponding to the image to be tested, and extracting all terrace plots from the image to be tested according to the target terrace pixel probability distribution includes: Acquire a terrace pixel point set and a background pixel point set according to the target terrace pixel probability distribution, wherein the terrace pixel point set includes a plurality of first pixel points with terrace attributes and a first pixel coordinate corresponding to each of the first pixel points, and the background pixel point set includes a plurality of second pixel points with background attributes and a second pixel coordinate corresponding to each of the second pixel points; Acquire at least one first connected area according to the first pixel point and the first pixel coordinates; Determine whether the area of the first connected region is greater than a preset area threshold; If the area of the first connected region is less than or equal to the preset area threshold, all the pixels in the first connected region are changed to second pixels; If the area of the first connected region is greater than a preset area threshold, the first connected region is converted into a line vector, and the line vector is converted into a surface vector, and the extraction vector result of the terrace is obtained according to the surface vector.
[0013] In a second aspect, the present invention provides a terrace extraction system based on a deep learning semantic segmentation model, the system comprising: The semantic segmentation model training module is used to construct a data set about terraced fields and train the semantic segmentation model according to the data set, specifically including: Performing a multi-level convolution operation on the data set to obtain a multi-level convolution feature, and performing parallel multi-branch processing on the convolution feature to obtain a first feature, a second feature, a third feature, and a fourth feature; Performing upsampling processing on the convolution feature to obtain a multi-level upsampling feature, generating a multi-scale feature according to the first feature, the second feature, the third feature, and the fourth feature, and fusing the multi-scale feature with the multi-level upsampling feature to obtain a fused feature; Splicing the fusion feature with the convolution feature to obtain a spliced feature, generating a category weight according to the spliced feature, and obtaining a terrace pixel probability distribution of each image in the data set according to the category weight; The terrace image detection module is used to input the image to be tested into the trained semantic segmentation model to obtain the target terrace pixel probability distribution corresponding to the image to be tested, and extract all terrace plots from the image to be tested according to the target terrace pixel probability distribution.
[0014] In a third aspect, the present invention provides a storage medium storing one or more programs, which, when executed by a processor, implement the above-mentioned terrace extraction method based on the deep learning semantic segmentation model.
[0015] In a fourth aspect, the present invention provides an electronic device, the electronic device comprising a memory and a processor, wherein: The memory is used to store computer programs; When the processor is used to execute the computer program stored in the memory, the above-mentioned terrace extraction method based on the deep learning semantic segmentation model is implemented.
[0016] Compared with the prior art, the present invention has the following advantages: According to the above-mentioned terrace extraction method based on deep learning semantic segmentation model, by introducing multi-level convolution and category weights, the expression ability of terrace plot crop characteristics and the model's recognition rate of terrace plots in complex terrace geographical environments are significantly improved. Increasing category weights can enhance category differentiation and have good robustness for a variety of geographic objects of different areas and locations. Even when the geographic objects are small or partially occluded, specific categories can be accurately identified, so that the model can better identify the type, area, and distribution of geographic objects when processing a variety of geographic objects (including terrace plots) in complex scenes, thereby optimizing the extraction effect of terrace plots and improving the accuracy of terrace plot extraction in complex scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A flowchart of a terrace extraction method based on a deep learning semantic segmentation model proposed in one embodiment of the present invention; Figure 2 A detailed diagram of step S101 proposed in one embodiment of the present invention; Figure 3 A schematic diagram of an image to be detected according to an example of an embodiment of the present invention; Figure 4 A schematic diagram of a first connected region according to an example of an embodiment of the present invention; Figure 5 A line vector diagram of an example of an embodiment of the present invention; Figure 6 A schematic diagram of a surface vector according to an example of an embodiment of the present invention; Figure 7 A schematic diagram of the structure of a terrace extraction system based on a deep learning semantic segmentation model proposed in one embodiment of the present invention.
[0018] The following specific implementation manner will further illustrate the present invention in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein should be understood by people with general skills in the field to which the present invention belongs. "Including" and similar words used in this article mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects.
[0020] like Figure 1 As shown, an embodiment of the present invention provides a terrace extraction method based on a deep learning semantic segmentation model, the method comprising steps S101 to S102, wherein: Step S101: constructing a data set about terraced fields, and training a semantic segmentation model according to the data set; It should be noted that if Figure 2 As shown, this step includes steps S1011 to S1014, wherein: Step S1011: Acquire a remote sensing image of a terraced field, and annotate the image in the remote sensing image of the terraced field frame by frame, wherein the annotated result includes the terraced field outline, terraced field surface and background, and generate a data set according to the annotated image.
[0021] Specifically, after acquiring the remote sensing image, the terraces need to be manually marked first. When manually marking the terrace plots, it is necessary to use image annotation software, open the remote sensing image of the terrace area in the software, and mark the features of the terrace plots in the image.
[0022] Since the spectral characteristics of terraced fields in terraced remote sensing images are similar to those of other landforms (such as vegetation, weeds, etc.), and in order to solve the core difficulty of terraced field extraction, that is, the diversity of terraced field characteristics in remote sensing images, the data used for annotation should include terraced fields of irregular shapes and sizes, with a total of no less than 8,000 images after cropping according to the actual remote sensing image resolution, in order to provide enough samples covering as many terraced features as possible for model training. When annotating, the texture, shape, location and other information of the terraced fields should be combined for judgment. The specific requirements are as follows: (1) For the outline of terraces, you need to use the polygon annotation tool in the software to outline the boundaries of the terraces point by point along the edge of the terraces. During the drawing process, you must ensure the accuracy of the boundaries, especially in areas where terraces and other landforms are interspersed. You need to carefully distinguish them to avoid mistakenly including other landforms in the terrace outline.
[0023] (2) When marking the terrace surface, use the marking tool to fill in the area representing the terrace surface within the terrace outline. Also pay attention to distinguishing it from the surrounding features to ensure that the marked terrace surface range is accurate.
[0024] (3) When annotating the background class, annotate the area except the terrace outline and terrace surface. During the annotation process, the consistency and completeness of the annotation should be maintained, and each terrace plot in the image should be accurately annotated. After the annotation is completed, the annotated data is saved in a software-specific format so that it can be subsequently converted into the data format required for model training.
[0025] The subsequent part is mainly operated by the computer program designed in the present invention. The program creates a clear data set directory structure on the storage device, and creates three subdirectories "train", "val" and "test" under the root directory, which are used to store the training set, verification set and test set data respectively. Under each subdirectory, two subdirectories "images" and "masks" are created respectively. The "images" directory is used to store remote sensing image data, and the "masks" directory is used to store the corresponding annotation data. And the data set is divided according to the ratio of 7:2:1, that is, 70% of the data is used as a training set for model training; 20% of the data is used as a verification set to evaluate the performance of the model and adjust the model parameters during the training process; 10% of the data is used as a test set for the final evaluation of the generalization ability of the model. In order to enable the terrace features to be evenly distributed in the data set, the present invention adopts a random sampling method when dividing the data. On the premise of having enough labeled data, the data set construction process of the present invention can ensure the subsequent training needs of the complex model for terrace extraction.
[0026] Step S1012: performing a multi-level convolution operation on the data set to obtain a multi-level convolution feature, and performing parallel multi-branch processing on the convolution feature to obtain a first feature, a second feature, a third feature, and a fourth feature; It should be pointed out that the multi-level convolution operation is first performed according to the following formula: ; in, Indicates The composite operation of the stage includes convolution, normalization, and activation in sequence. , Respectively represent Level Level convolutional features, Represents an image in a data set. Specifically, the second-level convolutional features and the third-level convolutional features contain rich spatial detail information such as edges and textures, and the fourth-level convolutional features and the fifth-level convolutional features contain high-order semantic information, such as object categories. In the present invention, by extracting these multi-level convolutional features, the model can simultaneously utilize detail and semantic information to improve the segmentation accuracy of terraces.
[0027] In addition, the convolutional features are processed in parallel with multiple branches according to the following formula: ; in, , , , Respectively represent the first feature, the second feature, the third feature, and the fourth feature, The void rate is of Convolution, d represents the void rate, represents global average pooling, Indicates first Convolution, then upsampling, and adjusting the features in height and width to and , represents the height of the feature, Indicates the width of the feature.
[0028] Step S1013: performing upsampling processing on the convolution feature to obtain a multi-level upsampling feature, generating a multi-scale feature according to the first feature, the second feature, the third feature and the fourth feature, and fusing the multi-scale feature with the multi-level upsampling feature to obtain a fused feature; It should be noted that the upsampling process is specifically performed according to the following formula: ; in, represents the adjusted i-th level convolution feature, , represents the i-th level up-sampled feature, , Represents bilinear times upsampling; The multi-scale features are generated according to the following formula: ; in, Represents multi-scale features, represents channel dimension splicing, express The feature tensor dimension of The height becomes , the width becomes , the number of channels becomes 256; The fusion is performed according to the following formula: ; in, represents the fusion feature, express convolution, , , They represent the up-sampled features of level 2, level 3, and level 4 respectively.
[0029] It should be pointed out that the above multi-scale features are obtained by capturing context features, and the fused features are obtained by fusing multi-level adjusted convolution features, thereby jointly solving the scale change problem of complex scenes. In the semantic segmentation model of the present invention, the algorithm for setting the fused features can improve the sensitivity of the model to complex terrace targets with undulating terrain, large slopes and diverse slope directions in mountainous areas. Compared with the original model, it has stronger specialization and is more suitable for the task of terrace extraction.
[0030] Step S1014: splicing the fusion feature with the convolution feature to obtain a spliced feature, generating a category weight according to the spliced feature, and obtaining a terrace pixel probability distribution of each image in the data set according to the category weight.
[0031] It should be noted that in this step, the category weight is generated according to the following formula: ; in, Represents the category weight vector, C represents the dimension of the vector, which is a positive integer. , represents a learnable weight matrix, The effect is to The feature vector processed by the activation function is linearly transformed and mapped to a space with the same dimension as the number of categories. represents an activation function, represents the rectified linear unit activation function, which is a nonlinear function defined as For each element in the input vector, if the element is greater than 0, it remains unchanged; if the element is less than 0, it is set to 0. By introducing Activation function, the model can learn more complex nonlinear relationships and enhance the expressiveness of the model. Represents the splicing feature, express The vector space where . The role of is to average pooling through the global average ( ) is linearly transformed, mapped to a new feature space, and the dimension of the feature is changed.
[0032] Generate a spatial attention map based on the class weights: ; in, is the spatial attention map, Indicates size The matrix of all 1s, represents tensor product; Perform feature reweighting on the spatial attention map: ; in, represents the feature representation after feature reweighting by the spatial attention mechanism, represents element-wise multiplication; The probability distribution of terrace pixels is obtained according to the following formula: ; in, represents the probability distribution of terrace pixels, represents normalization processing, is a learnable weight matrix, Indicates bilinear 8x upsampling, is a learnable bias vector.
[0033] It should be pointed out that the purpose of obtaining the spatial attention map in this step is to adjust features for different categories, improve the model's ability to distinguish difficult samples, and improve segmentation accuracy. In the present invention, due to the introduction of convolutional features with higher spatial resolution in the model, it is more sensitive to complex shapes in space, and can improve the accuracy of the model in separating irregular contours of terraces, which has a considerable improvement on the problem of difficulty in distinguishing terraces from surrounding objects.
[0034] In addition, it should be noted that the convolution features concatenated with the fusion features are generally one or more of the second-level convolution features, the third-level convolution features, and the fourth-level convolution features.
[0035] Step S102: inputting the image to be tested into the trained semantic segmentation model to obtain the target terrace pixel probability distribution corresponding to the image to be tested, and extracting all terrace plots from the image to be tested according to the target terrace pixel probability distribution.
[0036] It should be pointed out that if Figures 3 to 6As shown, after the semantic segmentation model is trained, the image to be identified is input into the trained semantic segmentation model, so that the target terrace pixel probability distribution can be obtained.
[0037] Further, a terrace pixel point set and a background pixel point set are obtained according to the target terrace pixel probability distribution, the terrace pixel point set includes a plurality of first pixel points with terrace attributes and a first pixel coordinate corresponding to each of the first pixel points, and the background pixel point set includes a plurality of second pixel points with background attributes and a second pixel coordinate corresponding to each of the second pixel points; at least one first connected region is obtained according to the first pixel points and the first pixel coordinates; it is determined whether the area of the first connected region is greater than a preset area threshold; if the area of the first connected region is less than or equal to the preset area threshold, all the pixels in the first connected region are changed to second pixel points, that is, small-area broken patches that may have errors are eliminated; If the area of the first connected region is greater than a preset area threshold, the first connected region is converted into a line vector, and the line vector is converted into a surface vector, and the extraction vector result of the terrace is obtained according to the surface vector.
[0038] Specifically, when converting line vectors, the first connected area in the raster image is narrowed to one pixel width to form a skeleton line, and unnecessary pixels on the edge are iteratively deleted. Starting from one point, adjacent pixels are connected according to the clockwise or counterclockwise rule, and the raster outline is converted into a line vector composed of coordinate points. Then, for closed lines, the surface is directly enclosed; for open lines, according to the connection relationship with other lines, the connection points are found to form closed polygons, and the correct topological relationship between surfaces and surfaces, and between surfaces and lines is reconstructed to ensure the accuracy and completeness of the data. Then, the Douglas Peucker algorithm is used to find the point farthest from the line connecting the first and last endpoints on the curve (surface boundary). If the distance is greater than the preset distance threshold, it is retained and recursively processed in segments; if it is less than the preset distance threshold, it is deleted and replaced with a straight line segment, thereby realizing the thinning and simplification of surface vectors.
[0039] Due to the complex texture features related to the shape and size of the terraces themselves, the differences in lighting conditions caused by the terrain, and the fact that the spectral features of the terraces themselves are not single, a small number of broken patches or overly complex edge contours will inevitably appear in the original prediction results of the model. If the results of the model output are directly used as the extracted terrace plots, although the difference in area is small, it will inevitably lead to poor usability in the process of identifying terraces one by one. The post-extraction processing flow designed by the present invention for the characteristics of the terraces can significantly improve such problems.
[0040] In addition, in some embodiments, the above extraction process is also evaluated, and the accuracy and error evaluation indicators are as follows: (1) Accuracy reflects the proportion of correctly predicted pixels in the total pixels. The formula is as follows: ; in Represents the number of pixels correctly predicted as positive (e.g., terrace pixels are correctly identified), represents the number of pixels correctly predicted as negative class (non-terraced pixels are correctly identified), is the number of pixels that are incorrectly predicted as positive, is the number of pixels incorrectly predicted as negative class.
[0041] (2) Intersection over Union (IoU) calculates the ratio of the intersection and union of the predicted result and the true label. For each category, the IoU is calculated and averaged to obtain the mean IoU (mIoU), which can fully reflect the segmentation performance of the model on different categories. The formula is as follows: ; (3) Mean Absolute Error (MAE) can intuitively reflect the average degree to which the prediction results deviate from the true value. The smaller the value, the more accurate the prediction. The formula is as follows: ; in, is the total number of pixels, is the predicted pixel class, These indicators together constitute the accuracy evaluation system of the present invention, and serve as the basis for evaluating the accuracy of the method in the longitude verification link.
[0042] In summary, according to the above-mentioned terrace extraction method based on deep learning semantic segmentation model, by introducing multi-level convolution and category weights, the expression ability of terrace plot crop characteristics and the model's recognition rate of terrace plots in complex terrace geographical environments are significantly improved. Increasing category weights can enhance category differentiation, have good robustness for a variety of geographical objects of different areas and locations, and accurately identify specific categories even when the geographical objects are small or partially occluded, so that the model can better identify the type, area, and distribution of geographical objects when processing a variety of geographical objects (including terrace plots) in complex scenes, thereby optimizing the extraction effect of terrace plots and improving the accuracy of terrace plot extraction in complex scenes.
[0043] like Figure 7 As shown, the present invention provides a terrace extraction system based on a deep learning semantic segmentation model, the system comprising: The semantic segmentation model training module 10 is used to construct a data set about terraced fields and train a semantic segmentation model according to the data set, specifically including: Performing a multi-level convolution operation on the data set to obtain a multi-level convolution feature, and performing parallel multi-branch processing on the convolution feature to obtain a first feature, a second feature, a third feature, and a fourth feature; Performing upsampling processing on the convolution feature to obtain a multi-level upsampling feature, generating a multi-scale feature according to the first feature, the second feature, the third feature, and the fourth feature, and fusing the multi-scale feature with the multi-level upsampling feature to obtain a fused feature; Splicing the fusion feature with the convolution feature to obtain a spliced feature, generating a category weight according to the spliced feature, and obtaining a terrace pixel probability distribution of each image in the data set according to the category weight; The terrace image detection module 20 is used to input the image to be tested into the trained semantic segmentation model, obtain the target terrace pixel probability distribution corresponding to the image to be tested, and extract all terrace plots from the image to be tested according to the target terrace pixel probability distribution.
[0044] On the other hand, the present invention further proposes a storage medium on which one or more programs are stored, which, when executed by a processor, implement the above-mentioned terrace extraction method based on the deep learning semantic segmentation model.
[0045] On the other hand, the present invention also proposes an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to implement the above-mentioned terrace extraction method based on the deep learning semantic segmentation model.
[0046] Those skilled in the art will appreciate that the logic and / or steps represented in the flowchart or otherwise described herein, for example, may be considered as an ordered list of executable instructions for implementing logical functions, and may be specifically implemented in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For purposes of this specification, "computer-readable medium" may be any device that can contain storage, communication, propagation or transmission of a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.
[0047] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.
[0048] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or a combination thereof: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0049] Although the embodiments of the present invention are described in detail above, it is obvious to those skilled in the art that various modifications and variations can be made to these embodiments. However, it should be understood that such modifications and variations are within the scope and spirit of the present invention as described in the claims. Moreover, the present invention described herein may have other embodiments and may be implemented or realized in a variety of ways.
Claims
1. A terrace extraction method based on a deep learning semantic segmentation model, characterized in that: The method comprises: Construct a dataset about terraced fields and train a semantic segmentation model based on the dataset, including: Performing a multi-level convolution operation on the data set to obtain a multi-level convolution feature, and performing parallel multi-branch processing on the convolution feature to obtain a first feature, a second feature, a third feature, and a fourth feature; Performing upsampling processing on the convolution feature to obtain a multi-level upsampling feature, generating a multi-scale feature according to the first feature, the second feature, the third feature, and the fourth feature, and fusing the multi-scale feature with the multi-level upsampling feature to obtain a fused feature; Splicing the fusion feature with the convolution feature to obtain a spliced feature, generating a category weight according to the spliced feature, and obtaining a terrace pixel probability distribution of each image in the data set according to the category weight; The image to be tested is input into the trained semantic segmentation model to obtain the target terrace pixel probability distribution corresponding to the image to be tested, and all terrace plots are extracted from the image to be tested according to the target terrace pixel probability distribution.
2. The terrace extraction method based on deep learning semantic segmentation model according to claim 1 is characterized in that: The steps of constructing a data set about terraces include: Acquire terraced field remote sensing images, and annotate the images in the terraced field remote sensing images frame by frame, wherein the annotated results include terraced field contours, terraced field surfaces, and backgrounds, and generate a data set based on the annotated images.
3. The terrace extraction method based on deep learning semantic segmentation model according to claim 1 is characterized in that: The step of performing a multi-level convolution operation on the data set to obtain a multi-level convolution feature comprises: Multi-level convolution operations are performed according to the following formula: ; in, Indicates The composite operation of the stage includes convolution, normalization, and activation in sequence. , Respectively represent Level Level convolutional features, Represents the images in the dataset; The step of performing parallel multi-branch processing on the convolution features to obtain the first feature, the second feature, the third feature and the fourth feature comprises: The convolutional features are processed in parallel with multiple branches according to the following formula: ; in, , , , Respectively represent the first feature, the second feature, the third feature, and the fourth feature, The void rate is of Convolution, d represents the void rate, represents global average pooling, Indicates first Convolution, then upsampling, and adjusting the features in height and width to and , represents the height of the feature, Indicates the width of the feature.
4. The terrace extraction method based on deep learning semantic segmentation model according to claim 3 is characterized in that: The step of performing upsampling processing on the convolution feature to obtain a multi-level upsampling feature, generating a multi-scale feature according to the first feature, the second feature, the third feature, and the fourth feature, and fusing the multi-scale feature with the multi-level upsampling feature to obtain a fused feature comprises: The upsampling process is performed according to the following formula: ; in, represents the adjusted i-th level convolution feature, , represents the i-th level up-sampled feature, , Represents bilinear times upsampling; The multi-scale features are generated according to the following formula: ; in, Represents multi-scale features, represents channel dimension splicing, express The feature tensor dimension of The height becomes , the width becomes , the number of channels becomes 256; The fusion is performed according to the following formula: ; in, represents the fusion feature, express convolution, , , They represent the up-sampled features of level 2, level 3, and level 4 respectively.
5. The terrace extraction method based on deep learning semantic segmentation model according to claim 4 is characterized in that: The step of splicing the fusion feature with the convolution feature to obtain a spliced feature, and generating a category weight according to the spliced feature comprises: The class weights are generated according to the following formula: ; in, Represents the category weight vector, C represents the dimension of the vector, which is a positive integer. , represents a learnable weight matrix, represents an activation function, represents the rectified linear unit activation function, Represents the splicing feature, express The vector space where .
6. The terrace extraction method based on deep learning semantic segmentation model according to claim 5 is characterized in that: The step of obtaining the probability distribution of terraced pixels of each image in the data set according to the category weight comprises: Generate a spatial attention map based on the class weights: ; in, is the spatial attention map, Indicates size The matrix of all 1s, represents tensor product; Perform feature reweighting on the spatial attention map: ; in, represents the feature representation after feature reweighting by the spatial attention mechanism, represents element-wise multiplication; The probability distribution of terrace pixels is obtained according to the following formula: ; in, represents the probability distribution of terrace pixels, represents normalization processing, is a learnable weight matrix, Indicates bilinear 8x upsampling, is a learnable bias vector.
7. The terrace extraction method based on deep learning semantic segmentation model according to claim 6 is characterized in that: The step of inputting the image to be tested into the trained semantic segmentation model to obtain the target terrace pixel probability distribution corresponding to the image to be tested, and extracting all terrace plots from the image to be tested according to the target terrace pixel probability distribution comprises: Acquire a terrace pixel point set and a background pixel point set according to the target terrace pixel probability distribution, wherein the terrace pixel point set includes a plurality of first pixel points with terrace attributes and a first pixel coordinate corresponding to each of the first pixel points, and the background pixel point set includes a plurality of second pixel points with background attributes and a second pixel coordinate corresponding to each of the second pixel points; Acquire at least one first connected area according to the first pixel point and the first pixel coordinates; Determine whether the area of the first connected region is greater than a preset area threshold; If the area of the first connected region is less than or equal to the preset area threshold, all the pixels in the first connected region are changed to second pixels; If the area of the first connected region is greater than a preset area threshold, the first connected region is converted into a line vector, and the line vector is converted into a surface vector, and the extraction vector result of the terrace is obtained according to the surface vector.
8. A terrace extraction system based on a deep learning semantic segmentation model, characterized in that: The system comprises: The semantic segmentation model training module is used to construct a data set about terraced fields and train the semantic segmentation model according to the data set, specifically including: Performing a multi-level convolution operation on the data set to obtain a multi-level convolution feature, and performing parallel multi-branch processing on the convolution feature to obtain a first feature, a second feature, a third feature, and a fourth feature; Performing upsampling processing on the convolution feature to obtain a multi-level upsampling feature, generating a multi-scale feature according to the first feature, the second feature, the third feature, and the fourth feature, and fusing the multi-scale feature with the multi-level upsampling feature to obtain a fused feature; Splicing the fusion feature with the convolution feature to obtain a spliced feature, generating a category weight according to the spliced feature, and obtaining a terrace pixel probability distribution of each image in the data set according to the category weight; The terrace image detection module is used to input the image to be tested into the trained semantic segmentation model to obtain the target terrace pixel probability distribution corresponding to the image to be tested, and extract all terrace plots from the image to be tested according to the target terrace pixel probability distribution.
9. A storage medium, characterized in that: The storage medium stores one or more programs, which, when executed by the processor, implement the terrace extraction method based on the deep learning semantic segmentation model as described in any one of claims 1 to 7.
10. An electronic device, characterized in that: The electronic device comprises a memory and a processor, wherein: The memory is used to store computer programs; When the processor is used to execute the computer program stored in the memory, it implements the terrace extraction method based on the deep learning semantic segmentation model as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Feature fusion coefficient learnable image semantic segmentation method
CN107766794A
Display panel appearance defect detection method
CN110097544A
Remote sensing image semantic segmentation method based on multi-scale attention fusion
CN113283435A
Semantic segmentation-based unstructured field road scene recognition method and device
CN114155481A
Paddy field segmentation method combining attention mechanism and spatial feature fusion algorithm
CN114419468A