Deep learning-driven camellia oleifera orchard topographic feature extraction method and system
Through deep learning combined with occlusion area recognition and feature reconstruction of RGB images and point cloud data, the accuracy of terrain feature extraction of oleifera orchards is solved, and efficient and accurate terrain feature extraction and management planning is achieved.
Patent Information
- Application Number
- CN202510511302.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-08-08
AI Technical Summary
The terrain feature extraction of the oil tea orchard is affected by the canopy obscuring and complex terrain, resulting in low extraction accuracy and limited production management planning.
Deep learning-driven method is adopted, combining RGB terrain-aware images and terrain point cloud perception data, and feature reconstruction of occlusion areas is realized through occlusion area recognition, feature particle size analysis and multiple reconstruction learning modules.
It improves the accuracy and efficiency of terrain feature extraction of oil tea orchards, can better adapt to complex shading and terrain changes, and provides high-fidelity terrain restoration.
Smart Images

Figure CN120451582A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to a deep learning-driven method and system for extracting terrain features from an oil-tea orchard. Background Art
[0002] Camellia orchards are a special agricultural production area, and the precise extraction of their terrain features is of great significance for the management, planning, and automated operations of the orchards. However, in practical applications, the extraction of terrain features in camellia orchards faces many challenges. Camellia orchards are different from general fields. Their dense tree canopies result in severe obstruction of ground terrain information. This obstruction affects the direct observation of terrain features and leads to a serious loss of ground terrain information, such as soil undulations and stone positions, which is not conducive to subsequent management and operations. Existing single-sensor methods, such as using only cameras or lidar, are often unable to effectively restore the terrain details of the obscured areas, thereby limiting the accuracy of terrain feature extraction. Secondly, camellia orchards are mostly located in complex hilly areas with large variations in terrain slope, which makes the extraction of terrain features more complicated. Simple image restoration or terrain reconstruction methods are often unable to adapt to such complex terrain changes, resulting in large deviations between the reconstruction results and the actual terrain, which in turn makes the terrain feature extraction results less accurate. Summary of the Invention
[0003] The present invention aims to solve the technical problem in the prior art that the extraction of terrain features in oil-tea orchards is affected by tree canopy occlusion and complex terrain, resulting in low extraction accuracy and limited production management planning. The present invention provides a deep learning-driven oil-tea orchard terrain feature extraction method and system to solve the problem.
[0004] The technical solution of the present invention to solve the above technical problems is as follows:
[0005] In a first aspect, the present invention provides a deep learning-driven method for extracting terrain features of an oil tea orchard, the method comprising: obtaining an RGB terrain perception image and terrain point cloud perception data of the oil tea orchard using a modal perception device; performing coordinate alignment on the RGB terrain perception image and terrain point cloud perception data and inputting the coordinates of the RGB terrain perception image and the terrain point cloud perception data into an occlusion area recognition module for joint identification of occlusion areas to obtain an occlusion area set; extracting a neighborhood occlusion area feature set corresponding to the occlusion area set, performing feature granularity analysis on the neighborhood occlusion area features corresponding to each occlusion area, and outputting a plurality of feature granularities corresponding to the occlusion area set; constructing a plurality of occlusion reconstruction learning modules, the plurality of occlusion reconstruction learning modules including a plurality of reconstruction granularities, establishing a mapping relationship between the occlusion area set and the plurality of occlusion reconstruction modules according to the plurality of feature granularities, adjusting the mapping relationship to the corresponding occlusion reconstruction module to perform feature reconstruction on the occlusion area set, and outputting a terrain extraction result of the oil tea orchard.
[0006] In a second aspect, the present invention provides a deep learning-driven oil tea orchard terrain feature extraction system, the system comprising: a perception acquisition unit, for using a modal perception device to acquire RGB terrain perception images and terrain point cloud perception data of the oil tea orchard; a joint identification unit, for performing coordinate alignment on the RGB terrain perception images and terrain point cloud perception data and inputting the coordinates of the RGB terrain perception images and terrain point cloud perception data into an occlusion area identification module for joint identification of occlusion areas to obtain an occlusion area set; a granularity analysis unit, for extracting a neighborhood occlusion area feature set corresponding to the occlusion area set, performing feature granularity analysis on the neighborhood occlusion area features corresponding to each occlusion area, and outputting a plurality of feature granularities corresponding to the occlusion area set; a feature reconstruction unit, for constructing a plurality of occlusion reconstruction learning modules, the plurality of occlusion reconstruction learning modules comprising a plurality of reconstruction granularities, establishing a mapping relationship between the occlusion area set and the plurality of occlusion reconstruction modules according to the plurality of feature granularities, adjusting the mapping relationship to the corresponding occlusion reconstruction module to perform feature reconstruction on the occlusion area set, and outputting the terrain extraction result of the oil tea orchard.
[0007] The beneficial effects of the present invention are: through deep learning technology, combined with RGB terrain perception images and terrain point cloud perception data, the occluded areas of the oil tea orchard are jointly identified and feature granularity analyzed, and the corresponding occlusion reconstruction learning module is selected according to the analysis results to perform feature reconstruction, thereby accurately extracting the terrain features of the oil tea orchard, effectively solving the problem that the extraction of terrain features of the oil tea orchard is affected by tree crown occlusion and complex terrain, and improving the accuracy of extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 Schematic diagram of the process of the deep learning-driven oil-tea orchard terrain feature extraction method provided by the present invention.
[0009] Figure 2 Schematic diagram of the structure of the deep learning-driven oil-tea orchard terrain feature extraction system provided by the present invention.
[0010] Explanation of the reference numerals: perception acquisition unit 11, joint recognition unit 12, granularity analysis unit 13, feature reconstruction unit 14. DETAILED DESCRIPTION
[0011] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0012] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the specified features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.
[0013] In the description of the present invention, the term "for example" is used to mean "used as an example, illustration or illustration". Any embodiment of the present invention described as "for example" is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is given to enable any person skilled in the art to implement and use the present invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other examples, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed herein.
[0014] Example 1:
[0015] like Figure 1 As shown, an embodiment of the present invention provides a deep learning-driven method for extracting terrain features from an oil-tea orchard, the method comprising:
[0016] S10: Use the modal perception device to obtain RGB terrain perception images and terrain point cloud perception data of the oil tea orchard.
[0017] For example, in this solution, modal sensing equipment is first used to acquire RGB terrain perception images and terrain point cloud perception data of the oil tea orchard. These modal sensing equipment typically includes advanced sensors such as high-resolution cameras and lidar. Specifically, high-resolution cameras can capture detailed image information of the oil tea orchard, including the color, texture, and shape characteristics of elements such as tree canopies, soil, and rocks, generating RGB terrain perception images. RGB terrain perception images are color images generated by capturing visual features such as the color and texture of the orchard's surface and objects on it using high-resolution cameras. This data is crucial for identifying areas obscured by tree canopies and inferring texture patterns in unobstructed areas. Simultaneously, lidar measures the three-dimensional coordinates of the orchard's terrain surface by emitting laser beams and receiving reflected signals, generating terrain point cloud perception data. This terrain point cloud perception data provides geometric features of the terrain, such as height, slope, and curvature, which are crucial for understanding terrain structure and inferring topographic details in obscured areas. For example, in an oil tea orchard, lidar can accurately capture the undulations of the ground beneath the tree canopies, while cameras can capture the color and texture information of the tree canopies. By fusing these two modal data, we can achieve a more comprehensive and accurate understanding of the topographical features of the oil tea orchard. This fused data acquisition method provides a foundation for subsequent terrain feature extraction and occluded area completion.
[0018] S20: performing coordinate alignment on the RGB terrain perception image and the terrain point cloud perception data and inputting the coordinates into an occlusion region recognition module to perform joint recognition of occlusion regions to obtain an occlusion region set.
[0019] Optionally, coordinate alignment is performed on the RGB terrain perception image and the terrain point cloud perception data. Coordinate alignment is key to ensuring accurate correspondence between pixels in the image and the 3D coordinates in the point cloud. This is achieved through image registration algorithms or the sensor's built-in calibration routine. After coordinate alignment, these data are input into the occluded region identification module. The region identification module is a system that integrates image processing and deep learning algorithms. It comprises an image occlusion detection submodule and a point cloud occlusion detection submodule. The image occlusion detection submodule uses a semantic segmentation network to analyze the RGB terrain perception image and identify areas occluded by objects such as tree canopies. These areas may appear as unusual variations in color, texture, or brightness in the image. The point cloud occlusion detection submodule analyzes the density distribution of the terrain point cloud perception data to identify areas where point cloud data is missing or unusually sparse due to object occlusion. The outputs of these two submodules are then combined for joint identification of occluded regions. Joint identification leverages the complementary nature of image and point cloud data to more accurately locate occluded areas in the oil tea orchard by comprehensively analyzing information from both modalities. The results of the joint recognition are integrated into an occlusion region set, which contains information about all identified occlusion regions, such as their location, size, and shape. This set provides key information for subsequent terrain feature extraction and occlusion region completion. For example, in an oil tea orchard, if an area has a particularly dense tree canopy, obscuring ground information in the RGB image and exhibiting a significantly reduced point density in the point cloud data, the occlusion region recognition module will mark this area as an occlusion region and add it to the occlusion region set.
[0020] S30: extracting a neighborhood occlusion region feature set corresponding to the occlusion region set, performing feature granularity analysis on the neighborhood occlusion region features corresponding to each occlusion region, and outputting a plurality of feature granularities corresponding to the occlusion region set.
[0021] In detail, extracting the neighborhood occlusion region feature set corresponding to the occlusion region set refers to extracting useful information from the surrounding environment of the occlusion region in order to better understand and reconstruct these occluded terrain features. Specifically, the neighborhood range of each occlusion region in the occlusion region set is determined. This neighborhood is an area with a certain size and shape centered on the occlusion region, which contains the image and point cloud data around the occlusion region. Afterwards, the image texture features and point cloud geometric features are extracted from this neighborhood to form a neighborhood occlusion region feature set. Image texture features include color gradient distribution, edge direction distribution, and texture consistency, which reflect the visual characteristics of the surface around the occlusion region. Point cloud geometric features include height change features, normal vector continuity, and curvature direction distribution, which reveal the geometric structure of the terrain around the occlusion region.
[0022] Performing feature granularity analysis on the features of the neighboring occluded regions corresponding to each occluded region involves classifying the features into different granularity levels based on their complexity and sophistication. Feature granularity is a relative concept, reflecting the ability of a feature to describe terrain details. For example, coarser-grained features may only capture the general outline of the terrain, while finer-grained features can reveal subtle variations and details. By analyzing factors such as the complexity, correlation, and redundancy of the features of the neighboring occluded regions, the feature granularity level corresponding to each occluded region can be determined. Subsequently, multiple feature granularities corresponding to the set of occluded regions are output. These feature granularity levels provide important guidance for subsequent reconstruction of the occluded regions. For example, for occluded regions with complex terrain variations, finer-grained features may be selected for reconstruction to ensure accuracy and sophistication of the reconstruction results. Conversely, for occluded regions with relatively gentle terrain variations, coarser-grained features can be selected for reconstruction to improve reconstruction efficiency and reduce computational cost.
[0023] For example, if an occluded area is located on a steep slope in a tea orchard, the characteristics of its neighborhood may include fine-grained features such as sharp changes in height and frequent changes in normal vector direction. In contrast, in a relatively flat area, the neighborhood characteristics of the occluded area may primarily manifest as coarse-grained features such as color gradients and texture consistency. Granular analysis of these features allows for the selection of appropriate reconstruction strategies and methods for subsequent reconstruction of the occluded area.
[0024] S40: Construct multiple occlusion reconstruction learning modules, each of which includes multiple reconstruction granularities. A mapping relationship between the occlusion area set and the multiple occlusion reconstruction modules is established according to the multiple feature granularities. The corresponding occlusion reconstruction module is transferred to perform feature reconstruction on the occlusion area set based on the mapping relationship, and the terrain extraction result of the oil tea orchard is output.
[0025] Specifically, constructing multiple occlusion reconstruction learning modules means providing diverse reconstruction strategies and methods for occlusion areas of different types, scales, and densities in oil tea orchards, such as large-scale continuous crown occlusion, local branch and leaf occlusion, etc. Multiple occlusion reconstruction learning modules are constructed according to different reconstruction granularities (i.e., the degree of refinement of terrain features). These modules can be neural networks based on deep learning. They are pre-trained and tuned to learn how to infer the occluded terrain information from the neighborhood features of the occluded area. Among them, each module is good at processing features at a specific granularity level, so that it can improve reconstruction efficiency and adaptability while ensuring reconstruction accuracy.
[0026] Based on the previously obtained feature granularity corresponding to each occluded area, a mapping relationship is established between the set of occluded areas and multiple occlusion reconstruction modules. The construction of the mapping relationship ensures that each occluded area can be assigned to the reconstruction module that best suits its characteristics for processing. For example, for occluded areas with fine-grained features (such as complex terrain on steep slopes), the system will select modules with higher resolution and finer reconstruction capabilities for reconstruction; while for occluded areas with coarse-grained features (such as simple terrain in flat areas), modules with lower computational complexity and higher reconstruction speed can be selected.
[0027] Once the mapping relationship is established, the corresponding occlusion reconstruction module is intelligently invoked based on the mapping relationship to reconstruct the features of the occluded area set. Leveraging the powerful capabilities of deep learning models, detailed information about the occluded terrain is inferred from the neighborhood features of the occluded area, achieving high-fidelity terrain restoration. During this process, each occlusion reconstruction module meticulously reconstructs and completes the occluded area at its specific granularity, resulting in more accurate and complete terrain feature extraction results.
[0028] For example, if an occluded area is caused by a dense canopy of trees, the neighborhood characteristics may manifest as height variations and texture anomalies on the ground beneath the canopy. Based on these characteristics, the system intelligently selects an occlusion reconstruction module that excels at processing fine-grained features. This module uses deep learning techniques to infer the detailed terrain of the ground beneath the canopy from the sparse point cloud and image information, thereby achieving high-fidelity restoration of the occluded area. Ultimately, the reconstruction results of all occluded areas are integrated to form a complete extraction of the tea orchard terrain.
[0029] In a preferred embodiment, the occlusion area recognition module includes an image occlusion detection submodule and a point cloud occlusion detection submodule; the image occlusion detection submodule uses a semantic segmentation network to identify occlusion areas of the RGB terrain perception image and marks the image occlusion area set; the point cloud occlusion detection submodule uses a point density algorithm to identify occlusion areas of the terrain point cloud perception data and marks the point cloud occlusion area set; the image occlusion area set and the point cloud occlusion area set are regionally fused and the occlusion area set is output.
[0030] In one specific embodiment, the occlusion region identification module includes an image occlusion detection submodule and a point cloud occlusion detection submodule. The image occlusion detection submodule uses a semantic segmentation network to analyze RGB terrain perception images. By identifying abnormal changes in image features such as color and texture, it accurately marks areas obscured by objects such as tree canopies, forming an image occlusion region set. For example, in an image, dense tree canopies can obscure the ground, causing that area to appear different from the surrounding environment. The semantic segmentation network can capture these features and mark the occluded areas. The point cloud occlusion detection submodule processes terrain point cloud perception data using a point density algorithm. Because occlusions can cause point cloud data to be missing or have reduced density, this submodule analyzes the point cloud density distribution, identifies these abnormal areas, and marks them as a point cloud occlusion region set. For example, in point cloud data, the point cloud density of ground areas obscured by tree canopies will be significantly lower than that of surrounding unobstructed areas. The point density algorithm can accurately identify this difference and mark the occluded areas.
[0031] Subsequently, the image occlusion area set and the point cloud occlusion area set are regionally fused. Specifically, by comprehensively analyzing the occlusion information in the image and point cloud data and utilizing the complementarity of the two, the position and range of the occlusion area can be more accurately determined. The fusion result is the occlusion area set, which contains information on all occluded areas in the oil tea orchard, and provides an important basis for subsequent terrain feature extraction and occlusion area completion. For example, in some areas, the image may be difficult to accurately identify occlusion due to similar lighting or color, but the point cloud data can clearly reflect the point density anomaly in the area. By fusing these two types of information, the occlusion area can be marked more accurately. Among them, the specific fusion method can be spatial position alignment fusion, feature information complementary fusion, probability model fusion or machine learning fusion, which is not specifically limited here.
[0032] In a preferred embodiment, the image occlusion detection submodule identifies occlusion areas on the RGB terrain perception image through a semantic segmentation network, including: the semantic segmentation network analyzes the RGB terrain perception image to obtain color distribution features, texture consistency features, and brightness difference features; analyzes the texture change indicators of continuous areas with respect to the color distribution features, the texture consistency features, and the brightness difference features, marks areas where the texture change index is less than a first preset change index threshold as image occlusion areas, and outputs an image occlusion area set after the traversal is completed.
[0033] Optionally, the image occlusion detection submodule uses a semantic segmentation network to identify occluded areas in RGB terrain-aware images. First, the semantic segmentation network analyzes the input RGB terrain-aware image and extracts color distribution features, texture consistency features, and brightness difference features from the image. These features reflect the color, texture, and brightness changes in different regions of the image and are important bases for identifying occluded areas. Subsequently, the submodule analyzes the texture change indicators of continuous regions with respect to these features. Specifically, the degree of change in color, texture, and brightness between adjacent pixels or regions is calculated, and these degrees of change are quantified as texture change indicators. By comparing these indicators with a preset first change indicator threshold, the submodule can identify areas with smaller texture changes. In practical applications in oil tea orchards, for example, the dense part of the tree canopy may block the ground, causing the occluded area to appear in the image with a color, texture, and brightness similar to the surrounding environment. The texture change indicators of these occluded areas are often less than the preset threshold, so the submodule marks these areas as image occlusion areas. Finally, the submodule traverses the entire image and integrates all areas marked as occluded to form an image occluded area set. This set contains information about all identified occluded areas in the image, providing a key basis for subsequent terrain feature extraction and occluded area completion.
[0034] In a preferred embodiment, the point cloud occlusion detection submodule identifies occlusion areas on the terrain point cloud perception data through a point density algorithm, including: analyzing the terrain point cloud perception data through a point density algorithm to output horizontal point density features and vertical point density features; analyzing the point density change index of the continuous area with respect to the horizontal point density features and the vertical point density features, marking the area where the point density change index is less than a second preset index threshold as a point cloud occlusion area, and outputting a point cloud occlusion area set after the traversal is completed.
[0035] Specifically, the point cloud occlusion detection submodule uses a point density algorithm to identify occluded areas in terrain point cloud data. This algorithm comprehensively analyzes the terrain point cloud data and calculates horizontal and vertical point density features. The horizontal point density feature reflects the horizontal distribution density of the point cloud, while the vertical point density feature reveals the vertical distribution of the point cloud. These features are crucial for identifying areas where point cloud data is missing or sparse due to occlusion. The point density algorithm analyzes the distribution density of points in point cloud data. Its core principle is to measure the density of the area by counting the number of points within a specific region. The algorithm first defines an analysis region, which can be regular (such as a grid or cube) or irregular, depending on the application scenario. It then counts the number of points within the region and divides the number of points by the area or volume of the region to obtain the point density. The point density value intuitively reflects the density of the point distribution. The algorithm can also adjust the size and shape of the analysis region as needed, allowing for flexible and intuitive density analysis and improving recognition accuracy.
[0036] The submodule analyzes the point density change index of the continuous area with respect to the point density feature. Specifically, it compares the point density differences between adjacent areas or points, and quantifies these differences as point density change indexes. By setting the second preset index threshold, the submodule can identify areas with smaller point density changes, which are often areas where point cloud data is missing or sparse due to occlusion. Taking the oil tea orchard as an example, if an area is blocked by a tree canopy, the point cloud data in that area will show a lower density in both the horizontal and vertical directions. When the point cloud occlusion detection submodule analyzes such an area, it will find that its point density change index is less than the preset threshold, thereby marking the area as a point cloud occlusion area. Finally, the submodule will traverse the entire point cloud data and integrate all areas marked as occluded to form a point cloud occlusion area set. This set contains information on all identified occluded areas in the point cloud data, providing an important basis for subsequent terrain feature extraction and occlusion area completion.
[0037] In a preferred embodiment, the feature granularity analysis of the neighborhood occlusion area features corresponding to each occlusion area includes: wherein the neighborhood occlusion area features include image texture features and point cloud geometric features of a preset neighborhood window, the image texture features include color gradient distribution, edge direction distribution and texture consistency, and the point cloud geometric features include height change features, normal vector continuity and curvature direction distribution; a pre-trained feature granularity classification module is used to input the neighborhood occlusion area features into the feature granularity analysis module for multi-channel convolution learning to output feature granularity labels, the feature granularity labels include a first granularity level, a second granularity level and a third granularity level; and a plurality of feature granularity labels corresponding to the occlusion area sets are output according to the feature granularity labels.
[0038] Preferably, the clear neighborhood occlusion area features include image texture features and point cloud geometric features of a preset neighborhood window. The preset neighborhood window refers to a local area of a fixed size or shape pre-set with the target point or target area as the center in image processing or point cloud analysis, which is used to extract neighborhood features around the target. Image texture features specifically cover color gradient distribution, which reflects the rate and direction of color change in the image; edge direction distribution, which describes the direction of the edge in the image; and texture consistency, which reflects the uniformity of the texture within the neighborhood. Point cloud geometric features include height variation features, that is, the undulation of the point cloud in the vertical direction; normal vector continuity, which indicates the degree of coherence of the normal vector of the point cloud surface; and curvature direction distribution, which reveals the direction and degree of curvature of the point cloud surface.
[0039] To further analyze feature granularity, a pre-trained feature granularity classification module was developed. This module uses multi-channel convolutional learning to input features from neighboring occluded regions. Multi-channel convolutional learning automatically extracts deep information from features. Through convolutional layers, image texture features and point cloud geometry are abstracted and transformed layer by layer to capture complex patterns and regularities within the features. After learning, the module outputs feature granularity labels, which include first, second, and third granularity levels, representing varying degrees of feature refinement.
[0040] For example, in an oil-tea orchard, for an area blocked by a tree canopy, the features of its neighboring blocked areas may show different granularities. If the terrain around the blocked area changes relatively gently, the color gradient distribution in the image texture features is relatively uniform, the edge direction distribution is relatively consistent, the texture consistency is high, and at the same time, the height change features in the point cloud geometry features are not obvious, the normal vector continuity is good, and the curvature direction distribution is relatively simple, then the feature granularity label of the area may belong to the first granularity level, indicating that the feature is relatively coarse-grained. On the contrary, if the terrain around the blocked area is complex, and both the image texture features and the point cloud geometry features show large changes, then the feature granularity label may belong to the third granularity level, indicating that the feature is relatively fine-grained.
[0041] Finally, the corresponding feature granularity label is output for each occluded area in the occluded area set according to the feature granularity label, thereby obtaining multiple feature granularity labels corresponding to the occluded area set. These labels provide an important basis for the subsequent selection of appropriate occlusion reconstruction learning modules according to the feature granularity.
[0042] In a preferred embodiment, a plurality of occlusion reconstruction learning modules are constructed, and the plurality of occlusion reconstruction learning modules include a first occlusion reconstruction learning module, a second occlusion reconstruction learning module and a third occlusion reconstruction learning module; an initialization pre-training network and a neighborhood occlusion area feature training sample are defined, wherein the initialization pre-training network includes a first-level convolutional network, a second-level convolutional network and a third-level convolutional network, wherein the convolution kernel of the first-level convolutional network is larger than that of the second-level convolutional network and larger than that of the third-level convolutional network; hierarchical transfer learning is performed according to the initialization pre-training network and the neighborhood occlusion area feature training sample to obtain the first occlusion reconstruction learning module, the second occlusion reconstruction learning module and the third occlusion reconstruction learning module.
[0043] Furthermore, to construct multiple occlusion reconstruction learning modules to accommodate reconstruction requirements for occluded regions of varying granularity, a system consisting of a first occlusion reconstruction learning module (G1 module), a second occlusion reconstruction learning module (G2 module), and a third occlusion reconstruction learning module (G3 module) was designed. First, an initialization pre-training network and training samples of neighborhood occlusion region features were defined. The initialization pre-training network consists of a three-level convolutional network. The first-level convolutional network has the largest convolution kernel, used to capture large-scale feature information, followed by the second-level convolutional network, and the third-level convolutional network has the smallest convolution kernel, focusing on extracting detailed features.
[0044] Hierarchical transfer learning is performed based on the initialization of the pre-trained network and training samples of features from neighboring occluded regions. The G1 module (coarse-grained reconstruction) employs an encoder-decoder architecture combined with a global attention mechanism. This design effectively reconstructs large-scale terrain undulations. For example, in oil tea orchards, for areas obscured by large tree canopies, the G1 module can capture the overall contours and changing trends of the terrain. The G2 module (medium-grained reconstruction) employs a method combining local patch fusion with a conditional GAN (generative adversarial network). This approach is suitable for texture-geometry co-completion of moderately occluded areas. For example, in oil tea orchards, for medium-scale terrain loss caused by tree canopy edges, the G2 module can generate reconstructions that both conform to the terrain geometry and preserve texture detail. The G3 module (fine-grained reconstruction) focuses on refining small missing areas, employing local convolutional completion networks such as PCN (Point Completion Network) or PartialConv. These techniques are capable of highly accurate completion and restoration of small areas of terrain loss caused by occlusion by small branches or leaves in oil tea orchards. Through hierarchical transfer learning, we obtained G1, G2 and G3 modules that are adapted to the reconstruction needs of occluded areas of different granularities. These modules together constitute the occlusion reconstruction learning module system, providing strong support for high-precision terrain extraction in complex terrain areas such as oil tea orchards.
[0045] In a preferred embodiment, hierarchical transfer learning is performed based on the initialized pre-trained network and the neighborhood occlusion area feature training samples, including: freezing the parameters of the second-level convolutional network and the third-level convolutional network, performing deep learning on the first-level convolutional network using the neighborhood occlusion area feature training samples, and outputting a first occlusion reconstruction learning module; freezing the parameters of the first-level convolutional network and the third-level convolutional network, performing deep learning on the second-level convolutional network using the neighborhood occlusion area feature training samples, and outputting a second occlusion reconstruction learning module; freezing the parameters of the first-level convolutional network and the second-level convolutional network, performing deep learning on the third-level convolutional network using the neighborhood occlusion area feature training samples, and outputting a third occlusion reconstruction learning module.
[0046] Specifically, to construct the first occlusion reconstruction learning module (G1 module), the parameters of the second and third convolutional networks are frozen to ensure they remain unchanged during the learning process. Then, deep learning is performed on the first convolutional network using training samples of neighborhood occlusion region features. Because the first convolutional network has larger convolution kernels, it is better suited to capturing large-scale features. Therefore, through deep learning, it can learn the large-scale undulations and overall structural features of the terrain, thereby outputting the G1 module suitable for coarse-grained reconstruction. For example, in a tea orchard, the G1 module can learn the overall contours of terrain obscured by large areas of tree canopy. Next, to construct the second occlusion reconstruction learning module (G2 module), the parameters of the first and third convolutional networks are frozen, leaving only the second convolutional network to learn. The convolution kernel of the second convolutional network is moderate, capable of capturing medium-scale features. By using training samples of neighborhood occlusion region features for deep learning, the second convolutional network can learn the medium-scale texture and geometric structure of the terrain, thereby outputting the G2 module suitable for medium-grained reconstruction. In tea oil orchards, the G2 module can perform texture-geometry collaborative completion for medium-scale terrain loss caused by canopy edge occlusion. Similarly, to construct the third occlusion reconstruction learning module (G3 module), the parameters of the first- and second-level convolutional networks are frozen, leaving only the third-level convolutional network to learn. The third-level convolutional network has smaller kernels and focuses on extracting detailed features. By deeply learning from training samples of features in neighboring occluded regions, the third-level convolutional network is able to learn small-scale terrain details, thereby outputting the G3 module suitable for fine-grained reconstruction. In tea oil orchards, the G3 module can perform high-precision completion and restoration of small-scale terrain loss caused by occlusion by small branches or leaves. Through layered transfer learning, the G1, G2, and G3 modules are derived to adapt to the reconstruction needs of occluded regions of different granularities, providing strong support for high-precision terrain extraction. Layered transfer learning fully utilizes the multi-level feature extraction capabilities of the pre-trained network. By freezing the parameters of the convolutional networks at different levels, the network at specific levels can be trained to meet the reconstruction requirements of different granularities. This method not only improves learning efficiency and avoids the huge computational overhead of training from scratch, but also enables fine-tuning for specific tasks (such as reconstruction of occluded areas of different granularities), thereby significantly improving the accuracy and effect of reconstruction.
[0047] The deep learning-driven method for extracting terrain features from a camellia orchard provided by the embodiments of the present invention has at least the following technical effects:
[0048] 1. The system combines RGB terrain-aware images with terrain point cloud data, aligns coordinates, and then feeds them into the occlusion region recognition module, enabling joint analysis of both image and point cloud data. This multimodal fusion approach more comprehensively captures the topographical characteristics of the oil tea orchard, effectively identifying occluded areas caused by factors such as tree canopy occlusion and terrain undulations, thereby improving the accuracy and completeness of occluded region recognition.
[0049] 2. A method for feature granularity analysis of the features of the neighborhood occluded areas corresponding to the occluded areas is proposed. By extracting image texture features and point cloud geometric features and performing multi-channel convolution learning on them, feature labels of different granularities are output. Then, multiple occlusion reconstruction learning modules are constructed based on these feature granularity labels. Each module corresponds to a different reconstruction granularity and can adaptively select the most suitable reconstruction module to reconstruct the features of the occluded areas. This adaptive reconstruction method significantly improves the accuracy and efficiency of terrain extraction, especially when dealing with complex occlusion situations.
[0050] 3. A layered transfer learning strategy is employed. By freezing the parameters of convolutional networks at different levels, the network at specific levels is trained to adapt to reconstruction requirements at different granularities. This approach not only fully utilizes the multi-level feature extraction capabilities of the pre-trained network, but also avoids the enormous computational overhead of training from scratch, significantly improving learning efficiency. Furthermore, by constructing a pre-trained network with initialization containing different convolutional kernel sizes, it is able to capture terrain features at different scales, further enhancing the accuracy and robustness of terrain extraction.
[0051] Example 2:
[0052] like Figure 2 As shown, based on the same inventive concept as the deep learning-driven oil-tea orchard terrain feature extraction method provided in Example 1, an embodiment of the present invention also provides a deep learning-driven oil-tea orchard terrain feature extraction system, the system comprising:
[0053] The perception acquisition unit 11 is used to acquire RGB terrain perception images and terrain point cloud perception data of the oil tea orchard using a modal perception device.
[0054] The joint identification unit 12 is used to coordinately align the RGB terrain perception image and the terrain point cloud perception data and input them into the occlusion area identification module to perform joint identification of occlusion areas to obtain an occlusion area set.
[0055] The granularity analysis unit 13 is configured to extract a neighborhood occlusion region feature set corresponding to the occlusion region set, perform feature granularity analysis on the neighborhood occlusion region features corresponding to each occlusion region, and output a plurality of feature granularities corresponding to the occlusion region set.
[0056] The feature reconstruction unit 14 is used to construct multiple occlusion reconstruction learning modules, which include multiple reconstruction granularities. A mapping relationship between the occlusion area set and the multiple occlusion reconstruction modules is established according to the multiple feature granularities. The corresponding occlusion reconstruction module is adjusted according to the mapping relationship to perform feature reconstruction on the occlusion area set, and the terrain extraction result of the oil tea orchard is output.
[0057] Furthermore, the joint identification unit 12 is further configured to perform the following steps:
[0058] The occlusion area recognition module includes an image occlusion detection submodule and a point cloud occlusion detection submodule; the image occlusion detection submodule uses a semantic segmentation network to identify occlusion areas of the RGB terrain perception image and marks an image occlusion area set; the point cloud occlusion detection submodule uses a point density algorithm to identify occlusion areas of the terrain point cloud perception data and marks a point cloud occlusion area set; the image occlusion area set and the point cloud occlusion area set are regionally fused and the occlusion area set is output.
[0059] Furthermore, the joint identification unit 12 is further configured to perform the following steps:
[0060] The semantic segmentation network analyzes the RGB terrain perception image to obtain color distribution features, texture consistency features, and brightness difference features; analyzes the texture change indicators of continuous regions with respect to the color distribution features, the texture consistency features, and the brightness difference features, marks the regions where the texture change indicators are less than a first preset change indicator threshold as image occlusion regions, and outputs a set of image occlusion regions after the traversal is completed.
[0061] Furthermore, the joint identification unit 12 is further configured to perform the following steps:
[0062] The terrain point cloud perception data is analyzed by a point density algorithm to output horizontal point density features and vertical point density features; the point density change index of the continuous area with respect to the horizontal point density features and the vertical point density features is analyzed, and the area where the point density change index is less than a second preset index threshold is marked as a point cloud occlusion area, and after the traversal is completed, a point cloud occlusion area set is output.
[0063] Furthermore, the particle size analysis unit 13 is further configured to perform the following steps:
[0064] The neighborhood occlusion area features include image texture features and point cloud geometric features of a preset neighborhood window, the image texture features include color gradient distribution, edge direction distribution and texture consistency, and the point cloud geometric features include height change features, normal vector continuity and curvature direction distribution; a pre-trained feature granularity classification module inputs the neighborhood occlusion area features into the feature granularity analysis module for multi-channel convolution learning to output feature granularity labels, the feature granularity labels include a first granularity level, a second granularity level and a third granularity level; and a plurality of feature granularity labels corresponding to the occlusion area sets are output according to the feature granularity labels.
[0065] Furthermore, the feature reconstruction unit 14 is further configured to perform the following steps:
[0066] Construct multiple occlusion reconstruction learning modules, which include a first occlusion reconstruction learning module, a second occlusion reconstruction learning module and a third occlusion reconstruction learning module; define an initialized pre-trained network and a neighborhood occlusion area feature training sample, wherein the initialized pre-trained network includes a first-level convolutional network, a second-level convolutional network and a third-level convolutional network, wherein the convolution kernel of the first-level convolutional network is larger than that of the second-level convolutional network and the third-level convolutional network; perform hierarchical transfer learning according to the initialized pre-trained network and the neighborhood occlusion area feature training sample to obtain the first occlusion reconstruction learning module, the second occlusion reconstruction learning module and the third occlusion reconstruction learning module.
[0067] Furthermore, the feature reconstruction unit 14 is further configured to perform the following steps:
[0068] Freeze the parameters of the second-level convolutional network and the third-level convolutional network, use the neighborhood occlusion area feature training samples to perform deep learning on the first-level convolutional network, and output the first occlusion reconstruction learning module; freeze the parameters of the first-level convolutional network and the third-level convolutional network, use the neighborhood occlusion area feature training samples to perform deep learning on the second-level convolutional network, and output the second occlusion reconstruction learning module; freeze the parameters of the first-level convolutional network and the second-level convolutional network, use the neighborhood occlusion area feature training samples to perform deep learning on the third-level convolutional network, and output the third occlusion reconstruction learning module.
[0069] Through the above detailed description of the deep learning driven oil tea orchard terrain feature extraction method in this specification, those skilled in the art can clearly understand the deep learning driven oil tea orchard terrain feature extraction system in this embodiment. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0070] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A deep learning-driven method for extracting terrain features from oil-tea orchards, characterized by: The method comprises: Use modal perception equipment to obtain RGB terrain perception images and terrain point cloud perception data of the oil tea orchard; Aligning the coordinates of the RGB terrain perception image and the terrain point cloud perception data and inputting them into an occlusion region recognition module for joint recognition of occlusion regions to obtain an occlusion region set; Extracting a neighborhood occlusion region feature set corresponding to the occlusion region set, performing feature granularity analysis on the neighborhood occlusion region features corresponding to each occlusion region, and outputting a plurality of feature granularities corresponding to the occlusion region set; Construct multiple occlusion reconstruction learning modules, each of which includes multiple reconstruction granularities. Establish a mapping relationship between the occlusion area set and the multiple occlusion reconstruction modules according to the multiple feature granularities. Use the mapping relationship to transfer the corresponding occlusion reconstruction module to perform feature reconstruction on the occlusion area set, and output the terrain extraction result of the oil tea orchard.
2. The method according to claim 1, wherein The occlusion area recognition module includes an image occlusion detection submodule and a point cloud occlusion detection submodule; The image occlusion detection submodule identifies occlusion areas on the RGB terrain perception image through a semantic segmentation network and marks a set of image occlusion areas. The point cloud occlusion detection submodule identifies occlusion areas on the terrain point cloud perception data through a point density algorithm and marks a set of point cloud occlusion areas. The image occlusion area set and the point cloud occlusion area set are regionally fused, and the occlusion area set is output.
3. The method according to claim 2, wherein The image occlusion detection submodule identifies occlusion areas on the RGB terrain perception image through a semantic segmentation network, including: The semantic segmentation network analyzes the RGB terrain perception image to obtain color distribution features, texture consistency features, and brightness difference features; Analyze the texture change index of the continuous area with respect to the color distribution feature, the texture consistency feature and the brightness difference feature, mark the area where the texture change index is less than the first preset change index threshold as the image occlusion area, and output the image occlusion area set after the traversal is completed.
4. The method according to claim 2, wherein The point cloud occlusion detection submodule identifies occlusion areas on the terrain point cloud perception data using a point density algorithm, including: Analyzing the terrain point cloud perception data using a point density algorithm to output horizontal point density features and vertical point density features; Analyze the point density change index of the continuous area with respect to the horizontal point density feature and the vertical point density feature, mark the area where the point density change index is less than the second preset index threshold as the point cloud occlusion area, and output the point cloud occlusion area set after the traversal is completed.
5. The method according to claim 1, wherein The feature granularity analysis is performed on the features of the neighboring occlusion areas corresponding to each occlusion area. include: The neighborhood occlusion area features include image texture features and point cloud geometric features of a preset neighborhood window. The image texture features include color gradient distribution, edge direction distribution, and texture consistency. The point cloud geometric features include height variation features, normal vector continuity, and curvature direction distribution. A pre-trained feature granularity classification module inputs the neighborhood occlusion area features into the feature granularity analysis module to perform multi-channel convolution learning and output feature granularity labels, wherein the feature granularity labels include a first granularity level, a second granularity level, and a third granularity level; A plurality of feature granularity labels respectively corresponding to the occlusion area sets are output according to the feature granularity labels.
6. The method according to claim 1, wherein Constructing a plurality of occlusion reconstruction learning modules, wherein the plurality of occlusion reconstruction learning modules include a first occlusion reconstruction learning module, a second occlusion reconstruction learning module, and a third occlusion reconstruction learning module; Defining an initialized pre-trained network and a neighborhood occlusion region feature training sample, wherein the initialized pre-trained network includes a first-level convolutional network, a second-level convolutional network, and a third-level convolutional network, wherein the convolution kernel of the first-level convolutional network is larger than that of the second-level convolutional network and larger than that of the third-level convolutional network; Hierarchical transfer learning is performed according to the initialized pre-trained network and the neighborhood occlusion region feature training samples to obtain the first occlusion reconstruction learning module, the second occlusion reconstruction learning module and the third occlusion reconstruction learning module.
7. The method according to claim 6, wherein Performing layered transfer learning based on the initialized pre-trained network and the neighborhood occlusion region feature training samples, including: Freeze the parameters of the second-level convolutional network and the third-level convolutional network, perform deep learning on the first-level convolutional network using the neighborhood occlusion region feature training samples, and output a first occlusion reconstruction learning module; Freeze the parameters of the first-level convolutional network and the third-level convolutional network, perform deep learning on the second-level convolutional network using the neighborhood occlusion region feature training samples, and output a second occlusion reconstruction learning module; Freeze the parameters of the first-level convolutional network and the second-level convolutional network, use the neighborhood occlusion area feature training samples to perform deep learning on the third-level convolutional network, and output a third occlusion reconstruction learning module.
8. Deep learning driven oil tea orchard terrain feature extraction system, characterized by: A system for implementing the deep learning-driven oil-tea camellia orchard terrain feature extraction method according to any one of claims 1 to 7, comprising: A perception acquisition unit, used for acquiring RGB terrain perception images and terrain point cloud perception data of the oil-tea orchard using a modal perception device; A joint recognition unit is used to coordinately align the RGB terrain perception image and the terrain point cloud perception data and input them into an occlusion region recognition module to perform joint recognition of occlusion regions to obtain an occlusion region set; a granularity analysis unit, configured to extract a neighborhood occlusion region feature set corresponding to the occlusion region set, perform feature granularity analysis on the neighborhood occlusion region features corresponding to each occlusion region, and output a plurality of feature granularities corresponding to the occlusion region set; A feature reconstruction unit is used to construct multiple occlusion reconstruction learning modules, wherein the multiple occlusion reconstruction learning modules include multiple reconstruction granularities, and establish a mapping relationship between the occlusion area set and the multiple occlusion reconstruction modules according to the multiple feature granularities. The mapping relationship is used to adjust the corresponding occlusion reconstruction module to perform feature reconstruction on the occlusion area set, and output the terrain extraction result of the oil tea orchard.